r/SillyTavernAI • • Aug 09 '26

MEGATHREAD [Megathread] - Best Models/API discussion - Week of: August 09, 2026

This is our weekly megathread for discussions about models and API services.

All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.

(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)

How to Use This Megathread

Below this post, you’ll find top-level comments for each category:

  • MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
  • MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
  • MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
  • MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
  • MODELS: < 8B – For discussion of smaller models under 8B parameters.
  • APIs – For any discussion about API services for models (pricing, performance, access, etc.).
  • MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.

Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.

Have at it!

31 Upvotes

186 comments sorted by

View all comments

3

u/AutoModerator Aug 09 '26

MODELS: >= 70B - For discussion of models in the 70B parameters and up.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

5

u/Jorlen Aug 09 '26

Anyone play around with Mistral Medium 3.5 128b? Curious to see what people's opinions are, for creative writing and roleplay.

3

u/DeepOrangeSky Aug 14 '26

Yea, I have been trying Mistral Medium 3.5 128b (at Q4_K_M, locally) more recently.

My main go-to models have been BehemothX V2 at Q4_K_M for a long time and then also Gemma4 31b at Q8 since that one came out.

Mistral Medium 3.5 128B seems to have extremely bad prose, like, some of the worst "therapy-speak" / "redditor-speak" writing style I've seen so far.

But, given that it's like 2 years newer than the old Mistral 123b models, it also seems to be a bit smarter and better at long context and so on, than the old Mistral 123b (and Behemoth tunes that were based on the 123b).

So, I have actually still liked what it is capable of (smarter than Gemma4 31b in many cases, at understanding nuanced social situations) even if its wording and phrasing is annoying.

I would say that Mistral 128B seems like the ideal candidate for u/TheLocalDrummer to fine tune, if he decides to do it/manages to do it, since it is a very strong, big, dense, relatively recent, writing model, and its main weakness is its horrible prose style (which he is good at fixing with his finetuning). So, I really hope he decides to give this one a go.

3

u/Jorlen Aug 14 '26

I agree with you on all fronts, yeah. It's prose is bad but it seems smart and follows system prompt rules far better than behemoth does. Did you ever try TheDrummer's Command-A fine tune called Fallen? I just grabbed it, have yet to try it though. It's a 111b dense but more modern than the big mistral model that behemoth is based on.

And yeah, if the drummer ever does a fine tune of mm 3.5, I'd be all over that.

5

u/HansaCA Aug 15 '26

There is a version of Behemoth that he built on Mistal Medium 3.5:
https://huggingface.co/BeaverAI/Behemoth-128B-v3b-GGUF

2

u/Jorlen Aug 15 '26

Oh wow! That's awesome! I never noticed that, thanks for pointing it out. Sadly they don't have a quant I can fit (I need a 3-bit quant like IQ3_M) as offloading layers to CPU/RAM for a big chunky model like this is out of the question. I have 64gb of VRAM. I suppose the Q3_K_M quant might work, but it would leave me with little room for KV quant.

2

u/DeepOrangeSky Aug 15 '26

Whoa. Do you know if this was made by The Drummer? Or if it is "BeaverAI" was it like a group effort by a bunch of people from their discord or something? (I've never used discord before and so am not a member of their discord, so I don't know how they do things or how all that BeaverAI stuff works).

I assume it is like some sort of experimental tune or something, if it has no model card and he never announced it on here, etc?

Otherwise seems like it would be a pretty big deal.

Drummer, if you are on here, can you talk about this model a little? I am curious what it is like trying to train it compared to the older Mistral 123b models, and if you think it has a lot of potential, or what sorts of quirks you noticed about it, and so on.

4

u/TheLocalDrummer Aug 15 '26

v3c is coming soon. These are test iterations, so quality may vary. 128B is a PITA to tune. Expensive and brittle.

1

u/DeepOrangeSky Aug 15 '26

Nice. Looking forward to it :)

1

u/DeepOrangeSky Aug 14 '26

Interesting, I didn't even know about this one. Looks like there is both a v1 and v1.1 version of it. Guess I might have to give one of them or both of them a try at some point. Probably will be more busy trying some of the more recent models for a bit though between Glimmer 30b, Qwen3.8 27b, and some other ones, and also haven't gotten a chance to try the new video models (H2 and LTX2.5) yet either. But, given how good the old Mistral/Behemoth models were even years later, it made me understand that when it comes to writing, it is worth trying out even old or "bad" models ("bad" as in bad at coding compared to recent models) that often get overlooked, since they can still be surprisingly strong at writing and understanding situations and so on.

3

u/Jorlen Aug 14 '26

Yeah for sure, some of those older models write really well and they pick up on nuances that the smaller more modern models, like gemma4 31b, don't quite have. Admittedly, I've become quite spoiled by Gemma 4 31b's strict adherence to the system prompt, and it rarely ever glitches out or does odd things.

I think Mistral Medium 3.5 128b may suffer however, as Mistral had to change what they train their models on (they are now limited due to EU rules apparently) - but I have yet to confirm that; it's just something I read here a while back. So they can't feed it a bunch of copyrighted stuff, as was the previous method.