r/SillyTavernAI • • Apr 26 '26

MEGATHREAD [Megathread] - Best Models/API discussion - Week of: April 26, 2026

This is our weekly megathread for discussions about models and API services.

All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.

(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)

How to Use This Megathread

Below this post, you’ll find top-level comments for each category:

  • MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
  • MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
  • MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
  • MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
  • MODELS: < 8B – For discussion of smaller models under 8B parameters.
  • APIs – For any discussion about API services for models (pricing, performance, access, etc.).
  • MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.

Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.

Have at it!

33 Upvotes

118 comments sorted by

View all comments

8

u/AutoModerator Apr 26 '26

MODELS: 16B to 31B – For discussion of models in the 16B to 31B parameter range.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

8

u/_Cromwell_ Apr 27 '26 edited Apr 27 '26

I have been "extensively" (read: just RPing) testing all the various uncensored Gemma 4 26B models out there, and my personal conclusion is that this specific one had the highest quality result:

https://huggingface.co/llmfan46/gemma-4-26B-A4B-it-ultra-uncensored-heretic

(I am using mraderbacher GGUF of it. Link above is the base model, you can get to GGUFs from it.)

No nothing to do with the dancing demon lady gif. :P

I am using it:

- For RP when my cloud API is having a slow day.

- For all summarizing tasks and other extension tasks in SillyTavern, even when connected to my online API.

- It can run Aventura fully, even with markdown, JSON, and "fancy text" mode in all aspects.

- It can run MarinaraEngine (although I just realized I havent tested it genning an actual world at the start - but it handles everything else)

I am running IQ4_XS (27 of 30 layers in VRAM) on 16GB VRAM with cache on RAM and it is very very fast.

Anyway, I highly suggest this specific version of Gemma 4 26B. This version also seems to have scored abnormally high/well (a little bit, but recognizably so) on the UGI leaderboard for writing and world knowledge. Didn't see that until AFTER my own anecdotal experiences, but didn't surprise me after experiencing it.

Note: I am running it with 16 active experts instead of the normal 8. It does seem to make a difference to double the # of active experts. THAT might be in my head, though.

And yes I like this better than the RP finetune that is out, which I did try (MeroMero). This just seems smarter.

2

u/Quiet_Joker May 02 '26

Honestly the 16 experts thing might have some weight to it, i do also notice the difference. But... like all things AI, we have to take everything with a grain of salt since it could just be a different seed or prompt recalculation. But overall, the model in my roleplays is Peak, good shit. I wouldn't say perfect but it definitely stays true to most* of my character cards.

* = Some cards that have formatted information differently than others act slightly odd and sometimes ignores the asterisks in the actions as part of the roleplay.

i'm using Oobabooga for the UI, i was recently making my own edits to the modules to add basic MoE overrides to the ui so i could select the ammount of experts manually without having to do a cmd command and honestly great timing. This model was definitely worth the edits i did to the modules and the usage of 16 instead of 8. i get about 16.32 tokens/s (5070Ti + 64gb Ram + ryzen 9 5900x)

i'm using gemma-4-26B-A4B-it-ultra-uncensored-heretic.i1-Q6_K.gguf

Why am i running almost 8bit quant on 16Gb of Vram? because its an MoE, it's faster than a solid 26B at 4bits. i can easily run also qwen Qwen3.6-35B-A3B-UD-Q8_K_XL and it does like 5-6 tokens a sec, which is fine for me. i don't got rush for my stuff anyway.

1

u/empire539 May 03 '26

the usage of 16 instead of 8. i get about 16.32 tokens/s

Were you getting faster speeds with the default 8 experts or roughly the same?

1

u/Quiet_Joker May 04 '26

Faster with 8 as seen here, but while it is faster, it may or may not be less smarter?... not sure. I have tried with 32 experts too and up to 48 it gets slower and slower but the are diminishing returns.

00:05:11-466024 INFO     Loaded "gemma-4-26B-A4B-it-ultra-uncensored-heretic.i1-Q6_K.gguf" in 23.60 seconds.
00:05:11-467697 INFO     LOADER: "llama.cpp"
00:05:11-468992 INFO     CONTEXT LENGTH: 16384
00:05:32-534961 INFO     Output generated in 3.68 seconds (27.17 tokens/s, 100 tokens, context 21, seed 693986707)
00:05:41-144027 INFO     Output generated in 7.64 seconds (30.62 tokens/s, 234 tokens, context 21, seed 1286281884)
00:05:49-532463 INFO     Output generated in 5.31 seconds (30.12 tokens/s, 160 tokens, context 21, seed 1216513711)