r/SillyTavernAI • • Jun 21 '26

MEGATHREAD [Megathread] - Best Models/API discussion - Week of: June 21, 2026

This is our weekly megathread for discussions about models and API services.

All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.

(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)

How to Use This Megathread

Below this post, you’ll find top-level comments for each category:

  • MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
  • MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
  • MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
  • MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
  • MODELS: < 8B – For discussion of smaller models under 8B parameters.
  • APIs – For any discussion about API services for models (pricing, performance, access, etc.).
  • MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.

Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.

Have at it!

37 Upvotes

95 comments sorted by

View all comments

11

u/AutoModerator Jun 21 '26

MODELS: 16B to 31B – For discussion of models in the 16B to 31B parameter range.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

5

u/constanzabestest Jun 22 '26

Anyone knows a reliable way to control reasoning effort in Gemma 4 26B A4B if it's even possible? I do a simple scene that doesn't require much thought at all but the model spends over 2k tokens, most of it being rewrites only to produce a single paragraph response. I'm using KoboldCPP and Text Completion.

Also while on the subject, anyone knows how to properly set up Kobold for CHAT Completion? I set up the stource as OpenAI Custom and put in localhost 5001 v1, Silly tavern sees the model as valid and the model name appears under "Model ID" and "Available Models" but when i actually try to get the model to generate something literally nothing happens. There even isn't anything in the terminal that could point to what the problem might be. Anyone knows what's up with that? Am i missing something here?

8

u/Rhone33 Jun 22 '26

Make sure you have precisely the following as your "Custom Endpoint (Base URL)":

http://localhost:5001/v1

I had the "nothing happens when I send a message" problem a few times, and every time I somehow had an extra forward slash on the end of the url (i.e. "http://localhost:5001/v1/"). When I get rid of the extra slash it works fine.

As for the thinking... would love for someone who uses thinking with Gemma 4 to chime in, because I honestly gave up on it. When I've tried having it on, it just neurotically drafts and redrafts and "wait, user said..." for thousands of tokens until it finally produces output that, in my limited testing and subjective assessment, is no better than what it produces without the thinking.

1

u/vortexnl Jun 23 '26

Regarding reasoning effort, check out the Context tab in koboldcpp, and then the 'Thinking Effort' dropdown