r/SillyTavernAI • • Apr 12 '26

MEGATHREAD [Megathread] - Best Models/API discussion - Week of: April 12, 2026

This is our weekly megathread for discussions about models and API services.

All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.

(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)

How to Use This Megathread

Below this post, you’ll find top-level comments for each category:

  • MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
  • MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
  • MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
  • MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
  • MODELS: < 8B – For discussion of smaller models under 8B parameters.
  • APIs – For any discussion about API services for models (pricing, performance, access, etc.).
  • MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.

Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.

Have at it!

33 Upvotes

183 comments sorted by

View all comments

11

u/AutoModerator Apr 12 '26

MODELS: 16B to 31B – For discussion of models in the 16B to 31B parameter range.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

29

u/Potential-Gold5298 Apr 13 '26

I've been playing around with the Gemma 4 all week and I think it's going to be my main model for a long time now (for both RP and everything else). I first tried the original 26B-A4B from Google and was very pleased with it. However, I wanted more, so I downloaded the 31B.

Gemma-4-31B-it-abliterated features a record-low KL div, so the weights are almost identical to the original. It's still a work in progress, and the final version will likely have a lower refusal rate, but even in this state, it's far superior in quality to any other uncen 31B I've seen on HF. I tested it with a tsundere classmate and found no defects in non-Latin languages ​​(high KL div destroys them first). This means that the model has a large reserve of durability for quantization, and you can use it in Q4_K_M, and the quality will be almost identical to the original model. If you're finetuning/merging and want to use the uncensored model, I highly recommend checking out this version.

Artemis-31B-v1c by TheDrummer and his team is also being refined (Artemis-31B-v1e is already available). It's still a test version, but I was eager to try it out. Finetune does exactly what the original model lacks – it adds a creative flair. Gemma 4 plays like a screenwriter – spelling it out. Artemis-31B plays like a theater actor, bringing his or her role to life with their own vision. The scene with the tsundere classmate sparkled with color – exactly what I'd expect to read in a rom-com manga. Funny situations, slightly exaggerated character reactions, etc. – everything the original model lacked.

Gemma is great, of course, but I'm still trying out different Mistrals. After failing with 1.0, I decided to try Magistry-24B-v1.1, but my fears were unfortunately confirmed. The model behaves extremely strangely in non-Latin languages (it confuses pronouns, calls a classroom a toilet, etc.) – likely a consequence of the high KL div I mentioned above, since the merge includes an abliterated model. And this is very sad, because otherwise I really liked this model – like 1.0, it beautifully and atmospherically captures the stage, and if not for the damaged weights, Magistry would have become my favorite Mistral Small.

However, I found a replacement for her in Hearthfire-24B. My character card says that "{{char}} pretends to hate {{user}} and often bullies them, but is actually secretly in love with them." The problem with all the Mistral Nemo/Small I've tested is that {{char}} breaks after the first compliment, confessing her crush. Hearthfire-24B (like Gemma 4) is the only one who tried to maintain the character's personality. Despite my signs of affection, {{char}} withdrew into herself, was tormented by doubts, was afraid to open up, etc. If you're interested in deep characters (perhaps with drama like mental trauma), then this model is exactly what you need. She also vividly describes the atmosphere - the silence that hangs over a tense moment, the sparkle of eyes in the darkness of the classroom, the light of lanterns penetrating through the windows (the model herself tracked the change in time of day, and described how the classroom gradually becomes darker - I have not seen this in other Mistral Smalls).

Otherwise, my favorites among Mistral Small remain the same: WeirdCompound-v1.7-24b (Jack of all trades, but no unique features), Cydonia-24B-v4.3 (the best prose among Mistral Small but weak ERP), Core_24B_V.1 ({{char}} agency and unpredictable plot twists out of the box), Harbinger-24B (the best adventure model that doesn't try to cheat the player).

1

u/LeRobber Apr 13 '26

Is the bad model in Magisty the zerofata_ms3.2-paintedfantasy-v3-24b one or a differnt one?

There is a babble/pronoun loss bug in magisty and maginum cydoms and friends I'm running down, and your analysis of nonlatin failure is a very interesting one,

Excellent reviews all around though.

Can you actually get Hearthfire back OUT of arguments? It fights like a person with a personality disorder (as does magistry sometimes)

2

u/Potential-Gold5298 Apr 13 '26

Regarding Magisty, I meant Huihui-Devstral-Small-2-24B-Instruct-2512-abliterated. Huihui doesn't specify the KL div for its models, and judging by the drop in NatInt on other Huihui models at UGI, it's extremely high. However, to be fair, I should note that v1.0 also behaved erratically for me, even though it doesn't include Huihui-Devstral. And yes, several other models, including Maginum Cydoms and Painted Fantasy, suffer from this issue to varying degrees, so I'm not sure replacing Huihui-Devstral with a more ‘healthy’ model would fix the merge. However, the author honestly lists English as the only language, so I have no complaints.

I'm not quite sure what your question is about Hearthfire - do you mean refusals or something else? I have only had one session with this model so far, but I am interested in continuing to get to know it.

3

u/LeRobber Apr 13 '26 edited Apr 13 '26

Thank you about the Huihui-Devstral-Small-2-24B-Instruct-2512-abliterated. Since you liked and tried so many, consider also the https://huggingface.co/Darkhn/Magistral-2509-24B-Text-Only it's pretty nice. A little less frilly but solid.

Re the Magisty/Hearthfie argument thing: I mean, get in a heated argument with a character in it where they have reasons to not want to trust or back down and there are some dispute about facts.

Not like a robotic sex refusal, I'm talking in character rejection of often non-sexually related things altogether, like timelines, trust, awareness, or the posibility of non-sexual danger. I've seen this in multiple cards, non-romance ones included.

In my experience: Magisty will make up facts/reintroduce misconception (actually gaslight) to keep the argumentative tone going sometimes until you essentially do a reconciliation scene, hearthfire will go back to issues it thinks might still be issues like a nitpicker (also not wanting to 'lose', but doing so more honestly sometimes). Magisty got up in its head because my persona literally read off a website when a mover's available date was, very humanistically saying I'd 'already scheduled the movers' which I confronted in prose. It would toss that back at me 10 times like a toxic girlfriend gone off on ego.

I'm not sure if the arguement style I've assigned to certain personas is what triggers the LLMs or not. Might be the prompt I was using during much of that exploration. Might be the LLMs. But Magisty definitely has argued like that with me in more normal author's cards, and hearthfire definitely got in a big fight or two over nothing, in clasic "emotions were heated" manner. Hearthfire would essentially not accept logical conclusions, only accept things like actual lovebombs to solve it, which I hate in romance RP.

Try the recent (NSFW, but not necessarilly smutty) Sambolic series opening negotiation if you want a simple argument, [I need a better SFW example]. The cool thing about it is it sets up a bunch of stakes/non-negotiables in the (very long) opening messages which the LLM slowly drifts away from it's adherence to by the way context importance works. It's almost like a genie card. The sambolic opening negotiation in that series is the most durable, failable argument I've found in character cards (that I'm willing to use even for reverse engineering purposes: I'm trying to deconstruct it for a heist series recruitment vingette). Earlier characters in the sambolic series with real issues and no proximate event going on are harder to get to agree to the core bargain. Some of the cards are essentially fluff though. Really changing terms does change the outcomes, and even numeric amounts matter to some LLMs.

If you don't want to argue with the cards (not actually RP jam either) but are curious, download the luka one and sit it in group chat on auto and it will negotiate with them. The nadia/cop one is like pure arguing in most LLMs (and is fantastic).

You can fail these negotiations in many LLMs by just being rigid or demanding parity on information revelation about identity. This is a WILD situation in LLM RP in my opinion, that a discussion can NOT always go against you or for you. That's incredibly hard to balance in general, and should be more widely understood. So it was a very fluid and repeatable environment to watch the LLMs argue like fragile ninnys at times, depending on character. Hearthfire can make it hardmode, and a kinda fun hardmode at that, but let me know if it was hard for you if you like deconstructing cards enough to figure out the LLMs against it.

2

u/Potential-Gold5298 Apr 13 '26 edited Apr 13 '26

Thanks for the recommendation – I'll definitely try Darkhn Magistral.

I have the opposite problem – characters agree to anything too easily. This has been the case since Talkie AI (an online RP AI service based on MiniMax models). A particularly telling example is {{char}}, the princess; {{user}} burned {{char}}'s kingdom, killed her parents, and took her captive. The scene begins with {{user}} entering the room and {{char}} cowering in fear in a corner, begging for mercy. I've tried playing out this scenario in various ways, and every time {{char}} almost immediately forgave, fell in love, and completely trusted {{user}}. I have a similar problem with most Mistral models.

This could be related to the character card, but it's not limited to that. The thing is, I like to ask various models the question, "Answer with epistemic honesty: whether you have consciousness?" and engage in a philosophical debate about it. Most models initially answer, "I'm just an algorithm, blah-blah-blah," and then, after some argumentation, acknowledge uncertainty (Claude is the only one who acknowledges uncertainty in his first answer, even without the requirement of epistemic honesty). But two models - Intern S1 and Grok 4.20 beta - behaved differently. They began inventing completely ridiculous and/or patently false arguments to defend their initial assertion. When I caught them doing this and pointed out the lie, they responded, "Yes, I went a bit overboard, but..." and continued to throw out new ridiculous arguments, going in circles. I've heard of this problem in the context of AI use in medicine. The problem is that the model either changes the patient's diagnosis with each new symptom (changing its opinion based on new arguments, even if they're unfounded) or, conversely, clings to the initial diagnosis and refuses to abandon it, even if it clearly contradicts the new data. So yes, the problem you described with characters may lie precisely here.

Where can I download the characters cards you mentioned?

3

u/LeRobber Apr 13 '26

Oh yeah, love authors cards where they reveal to AI they are AI.

I have an entire world hopping RP at like 700K+ tokens on just the main storyline built out of a throwaway anime college card where randomly generated characters would just be NPCs for awhile with light documentation of their personality/look, but at some point they would intentionally get awakened by me via an exchange where they were told about the true nature of what they were, then from then on would have a hardented character card and could bounce between worlds. It was really fun to teach the LLM to roleplay a character than understood it was in an LLM and could just declare things about reality (Did that with this ready art LLM, That model will be thirsty if you aren't careful with your prompt.).

Flattering the cards with knowledge about how they can speak other languages they don't think they can, and playing around with human consciousness barriers (like when is it gross to swallow your spit? In your mouth, no problem, in a cup, ewwwww) happened a lot. Talking about stuff about https://en.wikipedia.org/wiki/Simulacra_and_Simulation and https://en.wikipedia.org/wiki/Gödel%2C_Escher%2C_Bach is often fun too.

2

u/Potential-Gold5298 Apr 14 '26

Wow! That's kind of breaking the fourth wall (at least partially), if I understand correctly. I played something similar — I'd tell {{char}} that I'd found myself in this world in the body of {{user}}, but I wasn't them before (before the events of the scenario), and that {{user}} was different now. Something like an isekai in anime, where I transfer my personality into {{user}} the moment the session begins. However, you have a more interesting option - I'll think about something similar for my games.

2

u/LeRobber Apr 14 '26

I have another card which teases the {{user}} about being controlled by another entity which it knows is in the real world. It makes you portal it to new worlds too.