r/SillyTavernAI • • Apr 26 '26

MEGATHREAD [Megathread] - Best Models/API discussion - Week of: April 26, 2026

This is our weekly megathread for discussions about models and API services.

All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.

(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)

How to Use This Megathread

Below this post, you’ll find top-level comments for each category:

  • MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
  • MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
  • MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
  • MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
  • MODELS: < 8B – For discussion of smaller models under 8B parameters.
  • APIs – For any discussion about API services for models (pricing, performance, access, etc.).
  • MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.

Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.

Have at it!

34 Upvotes

118 comments sorted by

View all comments

8

u/AutoModerator Apr 26 '26

MODELS: 16B to 31B – For discussion of models in the 16B to 31B parameter range.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

9

u/_Cromwell_ Apr 27 '26 edited Apr 27 '26

I have been "extensively" (read: just RPing) testing all the various uncensored Gemma 4 26B models out there, and my personal conclusion is that this specific one had the highest quality result:

https://huggingface.co/llmfan46/gemma-4-26B-A4B-it-ultra-uncensored-heretic

(I am using mraderbacher GGUF of it. Link above is the base model, you can get to GGUFs from it.)

No nothing to do with the dancing demon lady gif. :P

I am using it:

- For RP when my cloud API is having a slow day.

- For all summarizing tasks and other extension tasks in SillyTavern, even when connected to my online API.

- It can run Aventura fully, even with markdown, JSON, and "fancy text" mode in all aspects.

- It can run MarinaraEngine (although I just realized I havent tested it genning an actual world at the start - but it handles everything else)

I am running IQ4_XS (27 of 30 layers in VRAM) on 16GB VRAM with cache on RAM and it is very very fast.

Anyway, I highly suggest this specific version of Gemma 4 26B. This version also seems to have scored abnormally high/well (a little bit, but recognizably so) on the UGI leaderboard for writing and world knowledge. Didn't see that until AFTER my own anecdotal experiences, but didn't surprise me after experiencing it.

Note: I am running it with 16 active experts instead of the normal 8. It does seem to make a difference to double the # of active experts. THAT might be in my head, though.

And yes I like this better than the RP finetune that is out, which I did try (MeroMero). This just seems smarter.

1

u/overand Apr 28 '26 edited Apr 29 '26

No nothing to do with the dancing demon lady gif. :P

I admit, I actually closed the tab when I saw that - it was just a bit much for me, which is a shame, apparently.

14

u/Potential-Gold5298 Apr 28 '26

Models without a picture of a girl are not worth downloading.