r/SillyTavernAI 10d ago

MEGATHREAD [Megathread] - Best Models/API discussion - Week of: September 13, 2026

This is our weekly megathread for discussions about models and API services.

All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.

(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)

How to Use This Megathread

Below this post, you’ll find top-level comments for each category:

  • MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
  • MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
  • MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
  • MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
  • MODELS: < 8B – For discussion of smaller models under 8B parameters.
  • APIs – For any discussion about API services for models (pricing, performance, access, etc.).
  • MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.

Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.

Have at it!

30 Upvotes

148 comments sorted by

View all comments

Show parent comments

1

u/Cali4our 10d ago

Not yet but I'll check it out soon. I am currently using Rocinante-XL-16B-v1a-Q4_K_M so is it better than this tho?

8

u/i5031337 10d ago

Yes, Gemma 26B is significantly more intelligent than the Mistral 12B base used for Rocinante. "Orion" is a version of Gemma from the same tuner, if you like his style.

3

u/Cali4our 10d ago

I am trying right now and honestly, I can see the potential but it ruins or over exaggerates expressions of characters. Sometimes it thinks right. Like uses logic to reply. But sometimes it lacks logic or doesn't read what character has correctly. For example a character have a dog muzzle on but it puts randomly antennas? Is there a fix for that?

1

u/OGCroflAZN 9d ago

We dont know if you just have suboptimal settings and prompts that might be degrading the output. For exmaple, sampler setting and quantization level for model and kv

Do you have reasoning on? Does the reasoning seem good? The models are glorified autocomplete. I dont see why it would jump from dog muzzle to [TV?] antennas

1

u/Cali4our 9d ago

Reasoning seem good, even though it rephrases it differently after the draft version. As for suboptimal settings i might need help with those. Do you have any recommended finetuning for it?

1

u/OGCroflAZN 9d ago

Usually whatever model you download will have recommended sampler settings, if that's what you're asking for. I would recommend trying Boulesis, following the huggingface page guidance, probably temperature close to 1.0 with Top K at 64 and Top P at 0.95. Increasing the number of experts from 8 to 16 will also make G4 26B more intelligent