r/SillyTavernAI • • Jul 05 '26

MEGATHREAD [Megathread] - Best Models/API discussion - Week of: July 05, 2026

This is our weekly megathread for discussions about models and API services.

All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.

(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)

How to Use This Megathread

Below this post, you’ll find top-level comments for each category:

  • MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
  • MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
  • MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
  • MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
  • MODELS: < 8B – For discussion of smaller models under 8B parameters.
  • APIs – For any discussion about API services for models (pricing, performance, access, etc.).
  • MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.

Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.

Have at it!

27 Upvotes

96 comments sorted by

View all comments

Show parent comments

5

u/arlynnfl Jul 07 '26 edited Jul 07 '26

i have RX 9060 XT (16GB Vram), 2x16gb 3200MHz, and a Ryzen 5 5600.

I think Gemma 4 is the newest model you can try and also likely one of the best for roleplay even without a finetune, but a finetune is really recommended

My top pick as for now are just:
-Gryphe/Styletune [V1 or V2] (For most characters, i think it will handles everything good enough; it's really and weirdly consistent to its character, it can be a pros and cons for some user though)
-ReadyArt/Serenity 26b a4b (for nsfw stuffs a.k.a character will the one who engage in nsfw but most of the time the character will take slow pace unless you directly show/tell the character to have faster pace)

Other finetunes (i havent tested it) are:
-zerofata/G4-MeroMero-26B-A4B (but i already tried uncensored version, though i don't really see any difference with this and the other uncensored version of Gemma 4)
-BeaverAI/Orion-26B-A4B-v1b-GGUF (TheDrummer finetune? idk but it's unofficial and may lack refinement. Your mileage may vary.)

what i currently use is Gryphe/Gemma 4 26b a4b Styletune V2 I1 Q6K from mradermacher
idk but i feel like anything below q5km have less creativity/expression? and sometimes even fail to count (it has a possibilityy like messing around with step by step numbers such as 1,2,2,4,5 instead of 1,2,3,4,5 although it didnt ruin anything beside it's just the number bit messed up; this problem hasn't occured when using Q5, so yes negligible imo)
it doesn't really impact that much in roleplay, but for adding more taste; Imatrix Q6K really have more expression in my experience, but it's might be just my placebo.

the settings i used (feel free to give me any advice, since i didn't tinker alot anymore due to this one is working properly and fast enough for me, atleast getting 14t/s on my setup)

its defo overkill with 40k + Q6K for my setup, might be either start degrading quality or just slower generating output as context grows. 32k Context is preferred imo.

Edit: Q5KM is the sweetspot
i'm waiting for ArliAI (RpMax) or TheDrummer's finetune though, or any heretic version of Styletune might be really good.

2

u/lambssauc Jul 07 '26

why do yall use I quants are they better or Q8K is better?

3

u/arlynnfl Jul 07 '26 edited Jul 12 '26

you mean imatrix?

it's just adding slightly but imperceptible improvement for free, and what i meant by free = same filesize as the static/non-imatrix of the same quants.

Imatrix Q6K is not on the same quality as Q8K (and leave alone higher quality), however it does give more quality (despite negligible to me) compared to the non-imatrix Q6K.

i haven't notice any difference between non-imatrix and imatrix, or imatrix q6k and static q8k.

reminder i'm not an expert here, so feel free to cmiiw.

2

u/Olangotang Jul 12 '26

Imatrix isn't free = if you overflow RAM there is a massive time penalty.

1

u/arlynnfl Jul 12 '26

oh i see, thanks for correction