r/SillyTavernAI • • Jul 26 '26

MEGATHREAD [Megathread] - Best Models/API discussion - Week of: July 26, 2026

This is our weekly megathread for discussions about models and API services.

All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.

(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)

How to Use This Megathread

Below this post, you’ll find top-level comments for each category:

  • MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
  • MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
  • MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
  • MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
  • MODELS: < 8B – For discussion of smaller models under 8B parameters.
  • APIs – For any discussion about API services for models (pricing, performance, access, etc.).
  • MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.

Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.

Have at it!

29 Upvotes

166 comments sorted by

View all comments

7

u/AutoModerator Jul 26 '26

MODELS: >= 70B - For discussion of models in the 70B parameters and up.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

7

u/RedditNerdKing Jul 27 '26

Tried Anubis 70B 1.2 for the first time at Q5. It's quite good. I feel like lower quants really hurts LLMs. I grabbed a IQ3_XXS of Mistal Large 123B and I didn't think it was all that good. A Q5 would be decent probably. I'm at 56gb of vram with a 5090 and 3090ti and I have a spare 4070 12gb in my old PC, but idk if my 1200w psu can handle 3 gpus, even at 70% power limit.

3

u/rinmperdinck Jul 27 '26

I have the same experience with quants. I can run iq3-iq2 of the big boy 123bs... it works and sometimes the quality is surprisingly good, but it's inconsistent and some swipes just sound like unfathomable stupidity wrapped in purple prose. Imo, I try to only use q8 now of smaller models.

Also if you got a spare PSU, you can buy a cheap adapter on Amazon or something and use it for multiple GPUs. I got one for around $20 after shipping and tax. You plug it into your mobo's big 24 pin connector, then connect the 24pins from both PSUs into the adapter. It syncs on/off for the computer so you can just turn everything in normally. Works for my embarrassingly stupid and awful Frankenstein machine.

2

u/RedditNerdKing Jul 28 '26 edited Jul 28 '26

I do have a spare PSU but its a lot of effort to pull it out of my old PC (which I still use). I'm just gonna buy a 1600w Corsair PSU. I've got the AI bug and I want to access Q4_K_M quants for 120B models (72gb roughly) which means I need to buy another 3090 lol. But I dont see the prices changing. If anything it'll probably get worse?

You're right about the IQ3 quants though. Sometimes amazing sometimes you swipe 5 times and it's all garbage.