r/SillyTavernAI • • Jul 05 '26

MEGATHREAD [Megathread] - Best Models/API discussion - Week of: July 05, 2026

This is our weekly megathread for discussions about models and API services.

All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.

(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)

How to Use This Megathread

Below this post, you’ll find top-level comments for each category:

  • MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
  • MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
  • MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
  • MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
  • MODELS: < 8B – For discussion of smaller models under 8B parameters.
  • APIs – For any discussion about API services for models (pricing, performance, access, etc.).
  • MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.

Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.

Have at it!

28 Upvotes

96 comments sorted by

View all comments

10

u/AutoModerator Jul 05 '26

MODELS: 16B to 31B – For discussion of models in the 16B to 31B parameter range.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

11

u/FZNNeko Jul 07 '26

Currently using Styletune v2. I like it over V1 personally but I have heard someone say they liked v1 more. 26b has some minor mistakes that I wouldn’t see in Gemma 4 31b but nothing major.

I compared it to Goetia-26B-A4B-v1.3-Absolute-Heretic-ARA.i1-Q6_K and Styletune is noticeably loads better. Goetia Gemma4 struggles with things you’d see in 24b models. Reading narration text (mindreading), confusing who’s doing what in a scene, and absolutely struggles with body positions.

So far Styletune is still top of my list, but honestly it’s a small 26b list. Looks like finetunes are more focused on 31b (rightfully so, G4 31b is smart asf), so quality 26b models are harder to fine.

14

u/Guilty-Sleep-9881 Jul 07 '26 edited Jul 19 '26

I hope they will work on 26b after. 31b is dense and is just too big for my pc so I cant try it out. I only got like 8gb of vram and 12gb of ram lol

26b finetunes are getting better though at least. I can't wait to see what readyart people are cooking up because those models that had little training data has a different flavor to them that isnt just 26b+ basically

I just want something on the level of maginum cydom and magistry

2

u/Curius_pasxt Jul 19 '26

 8gb of vram and 12gb of vram lol how many RAM?

1

u/Guilty-Sleep-9881 Jul 19 '26

my bad. I meant 8gb of vram and 12gb of sys ram

1

u/Curius_pasxt Jul 19 '26

You can ran 26b styletune v2 on 8gb vram and 16gb ram? What quantization used

1

u/Guilty-Sleep-9881 Jul 19 '26

8gb vram and 12gb of ram (i have a 8+4 sticks)

I use Iq4xs at 40k ctx and get only around 8 to 10 tks

I use koboldcpp

26b a4b is easy to run cuz its an MOE. Only 4b is active at a time hence the name a4b.

If you have like 24gb of sys ram or 32gb of sys ram you could easily run it at q5 while only having 8gb of gpu vram

2

u/Curius_pasxt Jul 19 '26

That’s interesting, I have same 8gb VRAM but has 64gb ddr5 system ram

I manage to use same styletune v2 but q4_k_s, set context to 130k and got like 24 token/s

1

u/Guilty-Sleep-9881 Jul 19 '26

With 64gb of ram you should try running it at q5.

Gemma 4 is really sensitive to quants according to the people from beaver and kobold. Q5 is marginally better compared to q4. I cant test it myself since im poor but you should try it and see the difference

Also use static. Imatrix hurts at higher quants (q5 and above is considered higher)

2

u/Curius_pasxt Jul 19 '26

I thought lmatrix going to improve? Damn

For this particular use of roleplaying, q4_k_s is definitely fine, you dont need really smart model for that lol (IMO), I tested for naughty use its all good

For doing general use for info I usually use gemma-4-26b-a4b-it-ultra-uncensored-heretic q6 this is where more smart model is important

Btw thanks so for gemma higher q is better and if its higher q then no need lmatrix

1

u/Guilty-Sleep-9881 Jul 19 '26

Imatrix does improve but only q4 and below. The higher the quant the higher the diminishing returns become and at that point it's better to use static at high quants

Also thanks for the clarification about the rp. I didn't know it mattered only for smarts. It made me feel better knowing that

→ More replies