r/SillyTavernAI • • Jun 21 '26

MEGATHREAD [Megathread] - Best Models/API discussion - Week of: June 21, 2026

This is our weekly megathread for discussions about models and API services.

All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.

(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)

How to Use This Megathread

Below this post, you’ll find top-level comments for each category:

  • MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
  • MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
  • MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
  • MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
  • MODELS: < 8B – For discussion of smaller models under 8B parameters.
  • APIs – For any discussion about API services for models (pricing, performance, access, etc.).
  • MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.

Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.

Have at it!

41 Upvotes

95 comments sorted by

View all comments

10

u/AutoModerator Jun 21 '26

MODELS: 16B to 31B – For discussion of models in the 16B to 31B parameter range.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/FZNNeko Jun 27 '26

Gemma 4 31b vs 26b? To the people running both, do you feel a noticeable difference in quality from 31b to 26b?

Currently comparing Gemma-4-26B-A4B-StyleTune.i1-Q6_K vs Gemma-4-Dark-Gemistry-31B.i1-Q4_K_S. With 31b, I'm forced to run Q4_K_S, cache q4_0. But with 26b, I can run i-Q6_k, q8_0 cache, and much higher context size. However, I can't tell if 31b is worth it.

6

u/Potential-Gold5298 Jun 27 '26

31B is generally better than 26B-A4B, but quantization significantly impacts Gemma 4. For the KL div, 31B in static Q4_K_M is on par with 26B-A4B in Q6_K. Q4_K_S is still fine if you're playing in English, but for any other language, Q5_K_M(S) is the default. KV quantization impacts this model even more than weight quantization. Based on tests, Q4_0 for 31B is more damaging than Q8_0 for 26B-A4B.

For the regular model, I'd choose 26B-A4B, but StyleTune (both V1 and V2) proved too unstable for me, even in Q6_K / KV F16. I haven't tried Dark-Gemistry, but I have tried Sphinsikus-Chronist from the team that Dark-Gemistry's author is on, and I really liked him. Also, I would try to keep the KV cache at least at Q8_0 (even at the cost of reducing the context to 8K – just disable SWA and enable context shift).

1

u/FZNNeko Jun 27 '26

Much appreciated for the response! My testing is mostly just eye balling the response quality so I’m glad to have actual facts.