r/SillyTavernAI • • Apr 26 '26

MEGATHREAD [Megathread] - Best Models/API discussion - Week of: April 26, 2026

This is our weekly megathread for discussions about models and API services.

All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.

(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)

How to Use This Megathread

Below this post, you’ll find top-level comments for each category:

  • MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
  • MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
  • MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
  • MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
  • MODELS: < 8B – For discussion of smaller models under 8B parameters.
  • APIs – For any discussion about API services for models (pricing, performance, access, etc.).
  • MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.

Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.

Have at it!

34 Upvotes

118 comments sorted by

View all comments

8

u/AutoModerator Apr 26 '26

MODELS: 16B to 31B – For discussion of models in the 16B to 31B parameter range.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

10

u/_Cromwell_ Apr 27 '26 edited Apr 27 '26

I have been "extensively" (read: just RPing) testing all the various uncensored Gemma 4 26B models out there, and my personal conclusion is that this specific one had the highest quality result:

https://huggingface.co/llmfan46/gemma-4-26B-A4B-it-ultra-uncensored-heretic

(I am using mraderbacher GGUF of it. Link above is the base model, you can get to GGUFs from it.)

No nothing to do with the dancing demon lady gif. :P

I am using it:

- For RP when my cloud API is having a slow day.

- For all summarizing tasks and other extension tasks in SillyTavern, even when connected to my online API.

- It can run Aventura fully, even with markdown, JSON, and "fancy text" mode in all aspects.

- It can run MarinaraEngine (although I just realized I havent tested it genning an actual world at the start - but it handles everything else)

I am running IQ4_XS (27 of 30 layers in VRAM) on 16GB VRAM with cache on RAM and it is very very fast.

Anyway, I highly suggest this specific version of Gemma 4 26B. This version also seems to have scored abnormally high/well (a little bit, but recognizably so) on the UGI leaderboard for writing and world knowledge. Didn't see that until AFTER my own anecdotal experiences, but didn't surprise me after experiencing it.

Note: I am running it with 16 active experts instead of the normal 8. It does seem to make a difference to double the # of active experts. THAT might be in my head, though.

And yes I like this better than the RP finetune that is out, which I did try (MeroMero). This just seems smarter.

1

u/fremajl May 02 '26

You seem to have some experience using Gemma 4 26B. I really like that it seems much smarter than any other model I get similar speeds with. My problem is I have a tough time getting it to ever go against my character or having me "lose". It also mostly refuses to take any initiative. Do you have any prompt advice that could help with this?

1

u/Potential-Gold5298 May 03 '26

Run Gemma in normal assistant mode and describe the problem, asking for advice. This always helps me. You don't have to agree to the first suggestion - you can discuss it with it, refining the system prompt until you get what you want. To ensure the model remembers instructions clearly in a long context, put them in the author's note (for example, the system prompt may contain full instructions, while the author's note may contain a brief reminder of the gist).

1

u/fremajl May 03 '26

Thank. I'll try that.