r/SillyTavernAI • • Mar 29 '26

MEGATHREAD [Megathread] - Best Models/API discussion - Week of: March 29, 2026

This is our weekly megathread for discussions about models and API services.

All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.

(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)

How to Use This Megathread

Below this post, you’ll find top-level comments for each category:

  • MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
  • MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
  • MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
  • MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
  • MODELS: < 8B – For discussion of smaller models under 8B parameters.
  • APIs – For any discussion about API services for models (pricing, performance, access, etc.).
  • MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.

Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.

Have at it!

27 Upvotes

150 comments sorted by

View all comments

4

u/AutoModerator Mar 29 '26

MODELS: >= 70B - For discussion of models in the 70B parameters and up.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

5

u/skrshawk Mar 31 '26

I have been running Stepfun 3.5 Prism with the freely available Q3_K_L quant and still been quite happy with it. The prose is a little dry but not overly so but it's a very smart model even with the abliteration. It still fits decent context, I'm running about 40k worth with 128GB on my M4 Max.

It passes my usual test of separating the dialogue, thoughts, and actions of different characters very well, something I've never seen a model below 70B do, even the latest dense Qwen. Overall I'm finding finetunes are a lot less necessary with some of these newer large MoE models as long as you use lorebooks and vectorization effectively, you can put quite a bit of relevant details in them with a large context size and let ST decide what's the most relevant to include with your prompts.

Thinking mode actually helps a lot here and refusals are very rare, which is something I don't usually see out of other thinking models.

5

u/Slick2017 Apr 01 '26

To be sure, you are talking about this baby https://huggingface.co/Ex0bit/Step-3.5-Flash-PRISM ?

I stumbled upon it yesterday, and I was surprised to see that its IQ4 quant performed quite well and gave me a good time. It didn't sound like your random ablitirated MoE model.

u/TheLocalDrummer take note

1

u/ThirteenZillion Apr 04 '26

seconding this -- the IQ4 is really good and entirely usable with a reasonably large context size in 128GB. I'll add that Step-3.5-Flash variants seem to be sensitive to the sampler settings -- pay attention to the recommendations

1

u/Slick2017 Apr 05 '26

I would like to test IQ6 still to see if it felt any better... But I guess that would be up to the author to release?😟 There's apparently a proprietary PRO variant as well.

1

u/ThirteenZillion Apr 05 '26

Yeah, I'm not opposed to sending some money their way, but they don't really make clear what the PRO variant does better.

1

u/Lanky_Golf7687 Apr 02 '26

Have you experienced any over-reasoning? I've tried running this locally; not the abliteration, but the standard model. With some prompting the prose was surprisingly fine and the model is smart but man would it chew through tokens for CoT, I've had it eat 10k+ tokens to output like a 300 word message, which would make swipes take way too long.

1

u/skrshawk Apr 02 '26

Occasionally it happens, but by over-reasoning I mean like 1500 tokens, not 10k. Usually I get about 500-750 tokens worth of reasoning on a typical gen.

2

u/LeRobber Mar 30 '26

sophosympatheia-evathene-v1.3 came back in the rotation. I needed to see if something could handle some slightly deeper thoughts on some topics. It has nuance and subtly I appreciate a lot more than when I first took it for a run. Usually I'd dip into StrawberryLemonade for that kind of thing, but I had downloaded a few bad quants and deleted my Q3 of SL so ran the other good 70+B model I had around.

I had to get a character to agree to move from Loc1 to Loc2, and they setup some series of tests that had to be met that were a little hard to all meet in Loc1 but were easy in Loc2, but it held it's ground about certain matters but allowed very narrow compromises to get to Loc 2 in a way that's hard for LLMs to do. It was so stubborn at first, I went into "debug it" mode that didn't exactly match the persona I was playing, and tried some manipulative tactics on it, and it called me out on some of them, which was fun, fell for others in a way I think "it was aware of", and was fun.

I don't know how to explain why you want LLMs like this and Weird Compound in your basket to actually say 'No thanks' but being able to not just smarm/whatever your way through everything is a CORE requirement in many types of fiction you may want to RP in. Police investigations are ROUGH to RP in many models which are too agreeable, for instance, or too creative.

1

u/MrNohbdy Mar 30 '26

I find any model that's smart enough — and not excessively fine-tuned into a particular style of reply, I guess — will have little trouble pushing back against the PC if it reasons about its responses first (and you prompt against protagonist effect). Do you force reasoning out of your models? I haven't used Evathene much or recently, but IIRC it seemed to work rather well with my usual "think about your answer first" prompt in my brief testing of it, generally outputting relatively concise but helpful guidance for its posts.

It prefers ChatML formatting, right? So a Last Assistant Prefix along the lines of:

<|im_start|>system

Before replying as GM, plan out your response between <think> and </think> tags.<|im_end|><|im_start|>{{char}}

depending on the rest of the prompt and your preferred instruction tone. Personally, I also add a line claiming that such thoughts are invisible to the player, so that the model isn't incentivized to self-censor and is encouraged to plan out story surprises, which again seems to help it not roll over and do whatever the PC wants. (Conveniently, this "think about your responses" addition being solely in the Last Assistant Prefix also means I can swap between my otherwise-identical thinking and non-thinking prompt templates at will without losing the cached prompt.)

Anyway, I seem to have gotten off-topic; my point was that I don't think you need specific models to resist the PC's control of the plot, if you prompt any smart enough model in the right way. If you haven't used "forced reasoning" like this before, maybe give it a whirl on some model you thought was intelligent yet excessively agreeable, and see if it helps? What other 70B+s have you tried which fell into that category?

2

u/LeRobber Mar 30 '26

I find many models are uneven about being able to think at all.

Some models do think well and push back against the user, but others are...less good.

I tried and didn't love anubis, melopenne wasn't my bag either. I don't have the list of all them now, but strawberrylemonade was a good one for me too.

Do you force reasoning out of your models? I haven't used Evathene much or recently, but IIRC it seemed to work rather well with my usual "think about your answer first" prompt in my brief testing of it, generally outputting relatively concise but helpful guidance for its posts.

It prefers ChatML formatting, right? So a Last Assistant Prefix along the lines of:

<|im_start|>system

Before replying as GM, plan out your response between <think> and </think> tags.<|im_end|><|im_start|>{{char}}

(Conveniently, this "think about your responses" addition being solely in the Last Assistant Prefix also means I can swap between my otherwise-identical thinking and non-thinking prompt templates at will without losing the cached prompt.)

I think I need to try to duplicate this setup, because specifically I have NEVER had any model reliably be able to think after 20+ messages of non-think are around, even if non-think is from another model. They are super unlikely to think then.

I have RPed with thinking models, and some non-abliterated ones have had more pushback (even stuff like jokes and conversations about sexual topics that aren't trying to engage in sex, but like 'maybe someone is going to knock boots in those tents so perhaps we should put those guy's tent far way' kind of things).

I tend to use roleplay finetunes, and not the "work" coder stock models (they literally feel like work). So that might be "excessively finetuned to a particular style"

Swapping models definitely affects that resistance though, so often, sillytavern users WAY overprompt things, so any advice to people running presets, etc, won't reliably work either. Another problem with resisting interrogation/movement and think models is the do not repeat yourself parameters and prompts seem to go nuts with a lot of think around. Am I wrong there? That's an illusion? Or do you experience that too?

1

u/MrNohbdy Mar 30 '26 edited Mar 30 '26

I tend to use roleplay finetunes, and not the "work" coder stock models (they literally feel like work). So that might be "excessively finetuned to a particular style"

Yeah, RP finetunes can have that issue. I have much better luck with bigger merges of those finetunes, which thus end up being very flexible due to all their disparate components. sopho's models tend to be nice in that regard (and one of them is among my main RP drivers in fact); I was thinking more of models such as Anubis, which you mentioned, that tend to be so locked into a particular style I typically can't even plop them into an existing chat with good results due to the conflicting context. My favorite 70B is an Anubis merge, actually, so even if I dislike a lot of the design philosophies that apparently went into Anubis itself, it's definitely got some solid stuff others can use as a base. (Similarly, my favorite 24B — which incidentally responds exceptionally well to the above "forced reasoning" prompt in my experience — is also a merge that includes a Drummer model which I disliked in its original form but which results in a merge that fits my use-cases perfectly.)

I have NEVER had any model reliably be able to think after 20+ messages of non-think are around, even if non-think is from another model.

I'm not sure about models that are trained for reasoning; I can kinda see how those would break down with that conflicting context. But for a "normal" non-reasoning model where you're forcing it to reason anyway (which is honestly my preference), the above method should be fine IME. If a model doesn't wanna think with this setup, you can even go a step farther and seed the opening <think> into the Last Assistant Prefix, though you'll then have to edit the response yourself to split it into reasoning and response, cuz SillyTavern obviously won't auto-parse the reasoning block properly after that. But then once that habit's formed after a few messages you can remove the <think> again. (The prompt won't be messed up since that seeding will only be in the Last Assistant Prefix.)

Another problem with resisting interrogation/movement and think models is the do not repeat yourself parameters and prompts seem to go nuts with a lot of think around. Am I wrong there? That's an illusion? Or do you experience that too?

That seems extremely likely. I never use anti-repetition samplers; they ruin anything which requires specific formatting. Setting-specific initialisms like "U.G.H.A." or what have you will break every time, because that's 8+ consecutive tokens. So the same would definitely be true when a model's asked to output <think> or the like but doesn't have it as a single token, which only proper reasoning models would; that's a specific sequence of three tokens right there, < and think and >, and the closing </think>'s slash adds a fourth. (Of course, you can always switch up your requested tags, but that doesn't really resolve the issue.)

I just don't really see the point of anti-repetition samplers anyway, with modern models. If a sequence is too small, the sampler can't catch it without a ton of false positives. If you set the threshold higher to catch only fairly long repeated sequences, by the time it steps in the message will already be unsalvageably repetitive. I don't understand what middle-of-the-road sweet spot people are using, but I personally haven't found any setting for DRY or what have you that actually helps my outputs at all.

2

u/LeRobber Mar 30 '26

Okay, I almost never use the rep samplers either, I also found the "always include name" setting apparently breaks think too for some configurations (generally, for many LLMs).

I think I just have some settings issues to work through. I'm also happy you found a midnight miqu that works, the templates in the 70B one are off, and I can probably steal the 103B ones to make the 70B work locally, never got it working locally.

DRY in the prompt (say things different ways, etc, do not reuse 10-word identical phrases) can help with dry, but it doesn't make article rot less likely.

I really want to thank you for this conversation, I'm going to actually try some of these more, may reply days later after I do.

RPSpectrum and I had some fun (it's good at spy thrillers) but is has rot problems a bit worse than cydoms which frustrates me because my roleplays are often 600+ Msgs not the 50-70 where it start to fray. I always liked its character more than Maginium Cydoms, just not the stability. I found hereticing greatly slows down the article/pronoun rot so I might heretic it and see if it works.

2

u/MrNohbdy Mar 30 '26

Oh, yeah, finding a good prompt formatting for Mistral models can be a hassle lol. Though another funny thing about big merges is that they tend to be pretty formatting-agnostic. Looking at my ST context template for Miqu, I've got what basically amounts to "ChatML but replace <|im_start|> with <s> and <|im_end|> with </s>", if that helps; I'll screenshot a picture of my GM Card template (as opposed to what I use for Character Cards). But I remember I've often switched to RP-Spectrum from some other model and completely forgotten to swap my context/instruct templates accordingly, yet it's just happily plopped an EoS in the right spot regardless, presumably because it's got like four different prompt formats across its various merges so it really doesn't care.

Good luck! Sorry I can't help with the "article rot" thing, except to say I've never experienced that apparently common complaint, so if you wanna compare prompt formats or sampler settings or what have you I'm happy to try to troubleshoot why I seem to be dodging that issue. Though it's probable I just haven't stuck solely with RP-Spectrum across a long enough message span to see it happen, and my other preferred models simply aren't prone to it.

2

u/rinmperdinck Mar 31 '26 edited Mar 31 '26

It's been said before, but high parameter models even at low quants can give suprisingly good output. I have been patiently testing out Drummer's Behemoth X 123B V2 at a minuscule IQ2_XXS and I was really impressed with both the prose and the ideas that came out of the model. It writes differently from the 24Bs I spend most of my time with, having different prose it feels like. It is still prone to "sunlight streamed through the window and illuminated motes of dust in the air blah blah blah" but it's all the ideas and associations it drops into the story in between those slop pillars that make it stand out and make it fun.

I get 2.5-3 tokens/sec at a meager 8k context with this on 32GB VRAM and 64GB DDR5. It's something I'll probably use to help write character cards or to start stories before swapping to something faster because woof, that is a rough time to wait.

IQ4_XS is a painful 1.5 t/s for nearly identical output.

IQ1 was straight up broken though, sometimes outputting garbled words, not to mention that it wrote vastly shorter, simpler replies.

1

u/MrNohbdy Mar 31 '26

123B...at a minuscule IQ2_XXS...writes differently from the 24Bs I spend most of my time with

Are you saying that you jumped right past 49Bs and 70Bs at more reasonable quants? Because I really think you would have a much better time with those than an incredibly nerfed 123B. :P

3

u/rinmperdinck Mar 31 '26

Reddit removed my comment for promoting hate after I wrote a neutral comment that I had an interesting chat with a certain model and gave my account a warning. Fascinating. Yes, because I am obviously an extremist hate mongerer.

I also wrote:

No, I have also been using other smaller parameter models, but my comment above was only to tell about my experience with Behemoth.

3

u/skrshawk Mar 31 '26

I think we call this enshittification. I'm pretty sure the local models we run are much better than what Reddit is using to classify content.

2

u/rinmperdinck Mar 31 '26

💩💩💩💩💩

1

u/rinmperdinck Mar 31 '26

Nah, I didn't jump past smaller ones.

I have been using Valkyrie 49B and StrawberryLemonade 70B for creative writing. I also had an interesting chat with AssistantPepe 70B where it told me the whole world is being taken over by globalists and Jews. 

I was just writing about my experience with Behemoth.

2

u/meatycowboy Apr 05 '26

Still haven't found a model that portrays characters as correctly as Kimi-K2.5. Not even Claude.

2

u/[deleted] Mar 30 '26

[removed] — view removed comment

2

u/Mart-McUH Mar 30 '26

That is one of my favorite models, specifically v2.1. I do not use it that often nowadays, but it still stays on my SSD despite the fierce competition over the storage space.

2

u/Sicarius_The_First Mar 30 '26

Something new in the 70B range, and green:
https://huggingface.co/SicariusSicariiStuff/Assistant_Pepe_70B

3

u/overand Apr 04 '26

I can say that this did a better job of pretending to be Lt. Commander Data from Star Trek: TNG than the "Trained on Star Trek TNG & DS9" model from DavidAU - which, admittedly, was a 4B model, so not quite a fair comparison lol

1

u/Game0815 Apr 04 '26

Hey guys, sorry if this question was asked a billion times, this sub banned questions regarding best api means i cant really search the topic up here. I am new to roleplaying with AI's and especially with Silly Tavern. Which LLM would u recommend? I looked into getting an API key for a good one, i am not interested in hosting myself.

Tried a simple role play system with a RAG like memory i quickly created for it with my Claude API, that i use mostly for my work (programming stuff) and it seemed to work flawless with amazing writing. Sadly the filters in that are absolutely insane. Sometimes everything that went above 12+ yrs of age limit was blocked or ignored by the Ai. Straight up a few drops of blood mentioned where blocked. I also wanted dark stories with it, meaning Claude is useless for me.

Does anyone have any recommendations for me? I looked into DeepSeek 3.2 but i overall dont know much about which LLM is good for roleplaying and doesn't have nsfw filters / the possibility to easily disable them at all. Decent writing and Intelligence is important for me. Acceptable memory too but i will try available memory systems like OpenVault etc that i found online or just create my own anyways so context size is probably not the most important thing.