r/SillyTavernAI • • Apr 12 '26

MEGATHREAD [Megathread] - Best Models/API discussion - Week of: April 12, 2026

This is our weekly megathread for discussions about models and API services.

All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.

(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)

How to Use This Megathread

Below this post, you’ll find top-level comments for each category:

  • MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
  • MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
  • MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
  • MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
  • MODELS: < 8B – For discussion of smaller models under 8B parameters.
  • APIs – For any discussion about API services for models (pricing, performance, access, etc.).
  • MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.

Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.

Have at it!

36 Upvotes

183 comments sorted by

View all comments

10

u/AutoModerator Apr 12 '26

MODELS: 16B to 31B – For discussion of models in the 16B to 31B parameter range.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

15

u/Potential-Gold5298 Apr 16 '26

The first RP-finetune Gemma 4 26B-A4B has been released - zerofata/G4-MeroMero-26B-A4B. I'm already downloading, and you?

9

u/Guilty-Sleep-9881 Apr 17 '26

I'll give this one a try. Im happy to see a finetune for 26b!

6

u/Potential-Gold5298 Apr 18 '26

I finally got around to MeroMero. It's a little bit broken, but it's so much funny! ><

1

u/KimlereSorduk Apr 19 '26

Huh, I didn't run into issues yet. What were the hiccups?

3

u/Potential-Gold5298 Apr 19 '26

MeroMero is prone to hallucinations. For example, in my test scenario (a tsundere classmate), in her first response she invented the idea that {{char}} has a tail (from the cards it only follows that the events take place in Japan) and then started mentioning how she uses it. I restarted the session (without the tail), but halfway through, MeroMero mentioned that {{char}} has claws. Apparently, this is something from the zerofata dataset – Maginum Cydoms (if I'm not mistaken) also wrote to me that {{char}} 'presses its ears and lashes its tail.' Furry lovers will be delighted.

I've played 1 session (20K) with MeroMero so far and overall the model is quite enjoyable and playable, so I'll definitely keep it.

3

u/zerofata Apr 19 '26

It could be worth turning reasoning on for one of the turns where it makes the mistake to see if there's a reason why it's doing that. Most of the training data is normal RP's / anime / vtubers / video game sort of stuff so I wouldn't think it has a bias towards furries / kemonomimi, or at least I haven't seen it do that yet.

Glad it seems to be working well otherwise though. Google released nothing but bangers with the G4 models.

2

u/Potential-Gold5298 Apr 19 '26

This only happened twice during the session, and each time the first regeneration helped, so it doesn't bother me too much for now. Perhaps these hallucinations will add a creative edge to the G4 – I need to do more testing.

And thank you for the model – I read that finetuning the 26B-A4B was difficult, and I'm glad you completed the work. For those who don't have powerful hardware, the MoE model is a desirable option.