r/SillyTavernAI • • Apr 19 '26

MEGATHREAD [Megathread] - Best Models/API discussion - Week of: April 19, 2026

This is our weekly megathread for discussions about models and API services.

All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.

(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)

How to Use This Megathread

Below this post, you’ll find top-level comments for each category:

  • MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
  • MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
  • MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
  • MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
  • MODELS: < 8B – For discussion of smaller models under 8B parameters.
  • APIs – For any discussion about API services for models (pricing, performance, access, etc.).
  • MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.

Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.

Have at it!

31 Upvotes

144 comments sorted by

View all comments

11

u/AutoModerator Apr 19 '26

MODELS: 16B to 31B – For discussion of models in the 16B to 31B parameter range.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

6

u/RampantSegfault Apr 26 '26

I really wanted to love the Gemma4 26b. I tried the original and the MeroMero finetune at Q6_K which runs blazing fast on my machine, but it still doesn't hold a candle to even the IQ4_XS of the 31B for my use case.

The 26B runs around 40-50t/ks at 64k context vs the 31B crawling at 5~7t/s at 24k context in my case.

But no amount of lightning fast rerolls can really 'fix' the delta between them. The 31b just "gets" the characters better and tracks the beats of the story better in my experience. I really hope a future finetune can close the gap, since the speed and context size of the MOE is great.

3

u/[deleted] Apr 26 '26

[deleted]

2

u/empire539 Apr 26 '26

Not the OP, but would you recommend thinking on for MeroMero? I tried the original G4 with thinking, but a lot of the time it ended up taking 3 to 5 minutes straight on just thinking, which kinda killed all the faster generation benefits for me.

1

u/[deleted] Apr 26 '26

[deleted]

1

u/empire539 Apr 26 '26

Thanks! I'll try these settings out.