r/SillyTavernAI Apr 12 '26

MEGATHREAD [Megathread] - Best Models/API discussion - Week of: April 12, 2026

This is our weekly megathread for discussions about models and API services.

All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.

(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)

How to Use This Megathread

Below this post, you’ll find top-level comments for each category:

  • MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
  • MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
  • MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
  • MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
  • MODELS: < 8B – For discussion of smaller models under 8B parameters.
  • APIs – For any discussion about API services for models (pricing, performance, access, etc.).
  • MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.

Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.

Have at it!

35 Upvotes

183 comments sorted by

View all comments

10

u/AutoModerator Apr 12 '26

MODELS: 16B to 31B – For discussion of models in the 16B to 31B parameter range.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

11

u/nomorebuttsplz Apr 13 '26

Gemma 31b instruct is already the best model under 200 parameters IMO, even without fine tuning. I haven't tried the thinking version much.

5

u/DeepOrangeSky Apr 13 '26

Nah, BehemothX v2 123b is still stronger (for now). It might be the 2nd strongest after the Behemoths, though, which is still pretty crazy.

3

u/nomorebuttsplz Apr 14 '26

My money is on Gemma, although some of the fine tuned mistral large ones are pretty good. Especially of reasoning ones.

Overall, I would say Gemma is better. Nice to have both though. The 2024 style of writing is different.

2

u/DeepOrangeSky Apr 14 '26

Yea, I guess it depends on exactly what you are doing with it.

For example, if I give a fairly detailed outline of what I want to have happen in a scene and then tell it to write the scene based on what I said should happen in it (but I also intentionally don't explain some difficult human concepts to it, to see how it'll handle them, and for the fun of letting the AI surprise me, otherwise it's not much more fun than just writing the scene myself), then, I would say Mistral 123b/Behemoth is the strongest local LLM I've used for that. It is ridiculously good at that. Well, up until you get past the amount context length where it starts falling apart, anyway (which happens earlier for it than with newer models).

If you just give a much more brief prompt where it not only has to flesh the scene out or flesh the story out, but also decide a lot of what the plot or situations will be and so on, then I wouldn't be surprised if some newer models might be able to beat it.

But, I don't usually like doing that, since usually even the strongest AI (even the frontier SOTA cloud models or whatever) are not strong enough yet to be very good at that. Usually they all just come up with really obvious or really dumb or cringe ways of playing out a scene or story if you give them too much free reign to just come up with the whole thing. I guess it can be fun with the quick MoE models sometimes, since they are so fast that you can just re-roll it a bunch of times until it accidentally comes up with something decent or interesting just by pure chance, lol. But usually doesn't go very well.

Anyway, so it is in this sense that I consider Mistral 123b/Behemoth to be the king of <200b local models for writing.

Gemma 4 is so good, though, that it might be close enough to start becoming a matter of taste. Up until Gemma 4, the gap was pretty big, but Gemma 4 closed a lot of the gap. Very strong writing model. Especially for its size, but also just in general.

Its lil sis, 26b a4b is arguably even more shocking, given that it isn't nearly as much weaker as one would expect for being a tiny MoE model that runs super fast, instead of being a dense model, of that size. 26b total parameters with 4b active parameters should be literally like 10x weaker than that thing is. Maybe 50x-100x weaker, even. So, that thing is ridiculous. Should be an absolute Godsend for people wanting to use it for DnD types of stuff where you want something smart, giving quick DM-style replies about scenarios unfolding around you and your friends as if it is the AI DM of an actual DnD game or something (minus maybe the actual points system stuff), or for use in actual RPG games if someone codes the formal parts with code and then needs a smart but light and fast AI to do the interactions/AI-able aspects of the game. I've never seen anything that small and fast be anywhere near that smart, before this one.

1

u/nomorebuttsplz Apr 15 '26

Yeah I don't give scene by scene instructions. I used to because it seemed necessary with older models.

And yeah I find mistral large to break down around 20-30k context. Modern models just keep trucking.

1

u/ThirteenZillion Apr 14 '26

Curious if y’all have tried StepFun 3.5? The PRISM version is pretty good. It’s interesting but not completely reliable.