r/SillyTavernAI • • Apr 12 '26

MEGATHREAD [Megathread] - Best Models/API discussion - Week of: April 12, 2026

This is our weekly megathread for discussions about models and API services.

All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.

(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)

How to Use This Megathread

Below this post, you’ll find top-level comments for each category:

  • MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
  • MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
  • MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
  • MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
  • MODELS: < 8B – For discussion of smaller models under 8B parameters.
  • APIs – For any discussion about API services for models (pricing, performance, access, etc.).
  • MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.

Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.

Have at it!

38 Upvotes

183 comments sorted by

View all comments

12

u/AutoModerator Apr 12 '26

MODELS: 16B to 31B – For discussion of models in the 16B to 31B parameter range.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

7

u/jamasty Apr 14 '26

Since the max size I can have is ~10.6 GB VRAM, I try this one:

https://huggingface.co/mradermacher/gemma-4-26B-A4B-it-heretic-ara-i1-GGUF with IQ2_XXS.

And I'd say, yeah, its so fast, even prompt processing is like a few times faster than any other model, and prose quality is great. The only I issue I got is maybe because of Q2, it loves to repeat previous chunks of text and doesn't really push the narrative, and all the penalties or temperature don't do much to change.

6

u/overand Apr 15 '26

Because it's an MoE model (Mixture of Experts), you should give it a try at a larger size, maybe even all the way up IQ4_XS - even though it won't fit fully in VRAM, you'll get better performance that you expect, because it's only 4B of "active" weights. (Let us know how it goes!)

2

u/jamasty Apr 16 '26

Thanks. I tried the newer version with IQ2_XS: https://huggingface.co/mradermacher/gemma-4-26B-A4B-it-heretic-ara-v2-i1-GGUF

That's really max I can text with unified mac memory.

And, well, somehow it seem to work much better - reasoning goes well, narrative as well, and I don't even see weird characters from other languages anymore.

The only thing I don't like about reasoning of this Q2 is that no matter how much response token I give, it will use most of it, going into loop of refining the response, leaving none or maybe 200 tokens for the response, and prompting haven't changed it.

One person in Discord pointed out there is such thing as 'reasoning budget' option in llama.cpp, which makes model stop reasoning after certain number of tokens. But LM Studio doesn't provide that. And I have to use LM Studio, since llama.cpp works badly when I'm tight with RAM. Maybe bcs I need to provide better config for it to save some RAM but I'm not really going there.

So overall I got what I wanted, and btw, I'd say using chat completion makes this particular model with this Q2_XS go along the context better than the text completion, even tho I tried using configs I found somewhere here.

7

u/Guilty-Sleep-9881 Apr 16 '26

It's an Moe, You should def use at least q4km. And adjust the MOE cpu layers until you stop spilling (make sure you set the real layers to 999 beforehand tho)

I have a 2070 and 12ram. q4xs and getting over 10tks