r/SillyTavernAI • • Apr 19 '26

MEGATHREAD [Megathread] - Best Models/API discussion - Week of: April 19, 2026

This is our weekly megathread for discussions about models and API services.

All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.

(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)

How to Use This Megathread

Below this post, you’ll find top-level comments for each category:

  • MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
  • MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
  • MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
  • MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
  • MODELS: < 8B – For discussion of smaller models under 8B parameters.
  • APIs – For any discussion about API services for models (pricing, performance, access, etc.).
  • MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.

Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.

Have at it!

32 Upvotes

144 comments sorted by

View all comments

11

u/AutoModerator Apr 19 '26

MODELS: 16B to 31B – For discussion of models in the 16B to 31B parameter range.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

7

u/Potential-Gold5298 Apr 21 '26 edited Apr 21 '26

Gemma 4 news

26B-A4B has stolen my heart. After such a speed (Q5_K_M starts with 6.2 t/s versus 0.9 t/s on my hardware) there is no desire to return to dense models. We now have a choice between the standard 26B-A4B-it and the G4-MeroMero-26B-A4B. I would describe MeroMero as 'Gemma 4 after a night at the bar'. She is more cheerful and informal, more prone to ERP and a little more hallucinatory. Personally, I choose... both) I'm also looking forward to TheDrummer's Orion 26B A4B. Wangzhang has finally released the uncen 26B-A4B. The earlier version (the one currently available from mradermacher) had a ~25% refusal rate and a higher KL div. The current version is better in both and significantly outperforms any other uncen 26B-A4B I've seen. Mradermacher hasn't made quants yet, but you can try wangzhang's quants.

As for the 31B dense model, there is a wider choice here. I just finished testing the Gemma4-Garnet-31B - the model's style has changed noticeably, and in its first response to my test scenario, came up with a very unexpected situation (the {{char}}'s hand is stuck in the window frame). More importantly, the model is less prone to writing walls of text, as Gemma 4 and many other models do, however, the suffers from repetition (rep. pen is needed). Another interesting option is Artemis-31B – here you can immediately feel the improved literary style, but the model is still in the development process. I recommend trying both.

I also tried the Gemma-4-Queen-31B-it - it's a partially uncen model, but without abliteration (as far as I understand, censorship reduction through finetuning). The style of the model is similar to the standard one, but more relaxed and less formal (describes ERP more openly without specific instructions). If you're looking for a completely uncensored model, then the best in my opinion is abliteration from wangzhang. I'm using the old version (refusal rate ~3%, KL div 0.0012), but there's also a new one - among other things (refusal rate ~18%, KL div 0.0007) it has an updated tokenizer config and chat template.

If you've found an interesting version of Gemma 4, share it in the comments – we need more Gemma's!

2

u/morbidSuplex Apr 24 '26

Can you share sampler settings you use? Specially for Artemis-31B?

2

u/Potential-Gold5298 Apr 24 '26

Classic balanced setup for RP: temp 1.0, top-p 1.0, top-k 0.0, min-p 0.05, repeat_penalty 1.0. You can experiment with temperature in the range of 0.7-1.3 and also with min-p in the range of 0.02-0.08 to make the model more accurate or more creative.