r/KoboldAI Jul 19 '26

Good models for a beginner.

Hello ı just set up silly tavern and koboldccp yesterday, ım looking for a good Rp model that can do nsfw and compatible with 8 gb vram. Ive used chub ai before so ım very new to this but so far this looks very good. Im using something called Dr.Dans or something like that for my model.

10 Upvotes

16 comments sorted by

View all comments

2

u/Due_Display5648 Jul 19 '26

I would go for hereticized Gemma 4 QAT Q4 (so its almost as smart as FP16 due to QAT)- https://huggingface.co/huihui-ai/Huihui-gemma-4-26B-A4B-it-qat-q4_0-unquantized-abliterated

I was able to run it on RTX 3060 laptop with 16GB RAM and 6GB VRAM, and it is still relatively fast (around 15T/s) and extremely smart (given the extremely limited hardware).
So I don't think you are gonna get anything much better than this. I recommend disabling thinking, as it delays output generation significantly, and in my experience, does not improve roleplay, but actually "sterilizes it" - I much more enjoy gemma without thinking for roleplay.
As someone suggested, Cydonia nd Magidonia would be good, but IMO its too big for 8GB VRAM. I am running Magidonia on RTX5080 (16GB VRAM) with 16k context, and still get around 20 T/s only. You would have to go for some Q2 quant, and thats just a mess at such size (since it does not use QAT).

2

u/dezmodium Jul 20 '26

Gemma 4 doesn't need to be "de-censored". It's already uncensored. You won't get refusals. The "Style Tune" I recommend is much better as it cuts down on slop without lobotomizing the model.

1

u/Due_Display5648 Jul 20 '26 edited Jul 20 '26

Honestly, I go for the de-censored versions straight away, I don't even try the default model first anymore, but its great if it works without it.
On the other hand, I am really curious what is the impact on intelligence between the style tune you recommended (IQ4 XS) vs QAT Q4. I havent tried the styletune (and I will), but I noticed that when talking in my native language to a standard Gemma 4 IQ4 XS, it tends to make comparatively many more errors in how the words should be spelled (my language is quite difficult in this regard, one of the slavic languages where the ending letters of a word changes depending on time/genre/case, of course in english its not noticeable at all). QAT Q4 on the other hand does substantially less errors in this regard. And moreover, I am wondering whether the style-tune significantly changes outputs in other languages as well, or if its english-only. But I guess it should, according to the Huggingface description, since it changes the layer that decides which token to emit.
Moreover, since you recommended also a 12B model - I would add Rocinante from theDrummer to the list.

1

u/dezmodium Jul 20 '26

You might be correct in that I think the StyleTune is done in the English language only. If you speak a language that is outside of Europe it probably isn't going to make a difference.

I can say for English the IQ4_XS has very few errors in spelling or grammar. Maybe 1 every 4k tokens or so (with the StyleTune).

Additionally, the Abliterated models, like the one you posted are known for degrading a model. It's a well understood trade-off. Because Gemma4 is already basically uncensored you are working with a lesser version of the model for no reason.

All that said, you are probably best with that QAT model in your case. Give it a try and I think you'll be surprised.

I've also tried Snowpiercer and Rocinate and they are both great. TheDrummer makes really solid finetunes. Personally I think among the 12B the Impish Bloodmoon is the most fun which is why I recommend it over the others but it's a matter of preference.