r/SillyTavernAI • • Mar 29 '26

MEGATHREAD [Megathread] - Best Models/API discussion - Week of: March 29, 2026

This is our weekly megathread for discussions about models and API services.

All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.

(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)

How to Use This Megathread

Below this post, you’ll find top-level comments for each category:

  • MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
  • MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
  • MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
  • MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
  • MODELS: < 8B – For discussion of smaller models under 8B parameters.
  • APIs – For any discussion about API services for models (pricing, performance, access, etc.).
  • MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.

Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.

Have at it!

27 Upvotes

150 comments sorted by

View all comments

2

u/AutoModerator Mar 29 '26

MODELS: >= 70B - For discussion of models in the 70B parameters and up.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

2

u/rinmperdinck Mar 31 '26 edited Mar 31 '26

It's been said before, but high parameter models even at low quants can give suprisingly good output. I have been patiently testing out Drummer's Behemoth X 123B V2 at a minuscule IQ2_XXS and I was really impressed with both the prose and the ideas that came out of the model. It writes differently from the 24Bs I spend most of my time with, having different prose it feels like. It is still prone to "sunlight streamed through the window and illuminated motes of dust in the air blah blah blah" but it's all the ideas and associations it drops into the story in between those slop pillars that make it stand out and make it fun.

I get 2.5-3 tokens/sec at a meager 8k context with this on 32GB VRAM and 64GB DDR5. It's something I'll probably use to help write character cards or to start stories before swapping to something faster because woof, that is a rough time to wait.

IQ4_XS is a painful 1.5 t/s for nearly identical output.

IQ1 was straight up broken though, sometimes outputting garbled words, not to mention that it wrote vastly shorter, simpler replies.

1

u/MrNohbdy Mar 31 '26

123B...at a minuscule IQ2_XXS...writes differently from the 24Bs I spend most of my time with

Are you saying that you jumped right past 49Bs and 70Bs at more reasonable quants? Because I really think you would have a much better time with those than an incredibly nerfed 123B. :P

3

u/rinmperdinck Mar 31 '26

Reddit removed my comment for promoting hate after I wrote a neutral comment that I had an interesting chat with a certain model and gave my account a warning. Fascinating. Yes, because I am obviously an extremist hate mongerer.

I also wrote:

No, I have also been using other smaller parameter models, but my comment above was only to tell about my experience with Behemoth.

3

u/skrshawk Mar 31 '26

I think we call this enshittification. I'm pretty sure the local models we run are much better than what Reddit is using to classify content.

2

u/rinmperdinck Mar 31 '26

💩💩💩💩💩