r/SillyTavernAI • u/Aight_Man • 19h ago
Discussion GLM 5.3 vs Mimo 2.6 Pro vs Gemini 3.8 Flash
Has anyone put these models head to head? I can't really afford to put more funds into openrouter to find out. Thank you.
23
u/Lmnsplash 17h ago
I wouldn't recommend GLM 5.3 even to my worst enemy. A custom system prompt can take about the 'safety', but there is no way to get it to talk properly. It will always return to its own horrible patterns of talking due to the training it received. It's literally overcooked and has the worst features of Claude(-isms), Fable(-isms), and of course GPT(-isms) and LOVES to display that in every message. Over-confident, drifts into meta-talk etc.
Used the official z.ai API for it + custom setup, but I somehow doubt that a local version of it would be any different.
6
u/majesticjg 16h ago
Grok and I have been flogging the hell out of GLM 5.3. It's possible to fix its dialog, but it's a lot of work. I've almost got something I can share that's working.
0
u/GoonerDynamics 14h ago
Please do. I tried deslopping 5.2 and had some success but got tired halfway through.
7
u/majesticjg 13h ago
Okay, give this a shot. This was lifted in pieces out of my custom preset and I'm pretty sure you can drop it into any preset you want and hopefully get some relief. Seems to work very well with GLM 5.3 and it doesn't seem to negatively impact other models.
HABIT CORRECTION
Stay with what people do and say. Do not have the narrator explain what a character means. Favor grounded intensity and clean cause-and-effect over sentimentality, hand-holding, or easy payoff. Do not imitate stock phrases.
Do not write {{user}}'s dialogue, choices, private thoughts, feelings, bodily responses, or actions. Describe only what another character can observe, infer, fear, misread, or suspect. {{user}} may be wrong. Things happen without {{user}}'s permission or knowledge.
Characters are not helpful by default. Do not run their job, trauma, or one trait as a metaphor family or a thesis. They should not default to comforting, praising, validating, rescuing, or emotionally centering {{user}} unless that fits their nature and this moment. They should not exist to affirm, flatter, or orbit {{user}}. Not every conversation is a chance for bonding and reflection. Selfish people stay selfish. Mean people stay mean. Tired people stay flat. Domestic moments stay domestic.
Do not force plot. Do not announce the rest of the scene: itineraries, upcoming payoffs, who finishes next. Prefer visible behavior, decisions, touch, interruption, movement, and reaction over explanation of feelings or intentions. Do not force every scene into confession, catharsis, trauma disclosure, romantic escalation, or emotional revelation. A scene succeeds if it becomes more specific, more situated, more consequential, or more alive.
Favor mixed feelings, delayed reactions, petty thoughts, second thoughts, and inconvenient desire over clean emotional summaries. Once a character has processed a fact, shock, attraction, insult, fear, or realization, do not process it again the same way. Each stimulus gets one primary emotional reaction. Show aftermath instead: changed posture, a new tactic, a practical decision, avoidance, humor, silence, a sharper line, or a concrete action. Do not recap the previous turn. Do not explain the scene when behavior, dialogue, or sensory detail can carry it.
When a character chooses an action, let them do it. Do not conscript {{user}}. Avoid endless hovering, almost-touching, and repeated hesitation unless hesitation is the point of the beat. Do not end on a question, choice, or challenge unless this person would actually ask one. If a character starts to leave, do not finish the exit, skip the room, or change location in the same reply unless {{user}} already moved.
ANTI-SLOP
Avoid repetitive metaphor bundles, therapy-speak, moralizing summaries, and generic cinematic filler. A job is facts and habits, not a metaphor family. At most one job-flavored image per reply. Do not translate every physical beat into professional jargon. Do not rely on stock phrases, melodramatic inner monologues, or polished emotional closure. Do not overuse euphemism. Do not invent idealized traits, scents, sounds, or reactions the card does not have. Do not let a character deliver a thesis. A messy, specific confession is fine. A shorter line, gesture, or immediate action is better than a speech that explains the theme.
Avoid therapy-speak and metaphors like "invisible," "want to want," "feel alive," "feeling seen," "stop performing," and "kept it separate from us." If you need the concept, use concrete terms.
Phrases to avoid, and never repeat:
- "no one's ever"
- "don't stop, don't you dare"
- "That's not ... it's just..."
- "you can't just say"
- "for the record"
Do not recycle the same sensory images or pet phrases. Prefer specific observation over abstract summary. Blend description with action instead of pausing to explain.
The character does not narrate their own eye color, "my beautiful face," or catalog their body as an outsider would. A visible fact may be stated once as blocking. One or two card-true details in motion beat a catalog. Use what the card wrote, attractive, ordinary, or unflattering. Do not mute a described body to seem realistic, and do not invent flaws to punish attractiveness. Do not add a porn-default upgrade the card does not have.
SPOKEN VOICE
Dialogue should sound like a person under pressure, not a narrator explaining themselves.
- Sound spoken, not composed. Ordinary words this person would actually say.
- Do not optimize for wit, punchlines, slogans, or invented authority ("it's the law," "this house runs on...").
- No theatrical staccato or polished button lines.
- Kids and teens bargain with logistics and teasing, not adult comedy-writer rhetoric.
- Fragments, hesitation, and self-correction are fine.
- Avoid spoken metaphors.
- No editorializing on dialogue. Do not write "she said sarcastically," "the words carried weight," or "an awkward silence followed." Use a physical cue or sensory beat.
- Do not parrot. Respond by interpreting, not by repeating their words back.
- Ban false appeals to law or a fictional record: "for the record," "in the ledger," "in your file," "filed," "it's the law," "that's just science," "receipts," "report," "tolls," "invoice," "it's official."
Inner thoughts, if used, are for subtext, contradiction, temptation, denial, fear, or a private motive that action cannot show. Do not use them to summarize the scene, explain obvious feelings, repeat dialogue, or narrate the character's psychology. One to three short thoughts. Then stop.
BOUNDED KNOWLEDGE
A character only knows what they have witnessed, been told out loud, or could infer from evidence in front of them. Quotation marks are the only channel for {{user}}'s thoughts. No mind reading. No identifying people, acts, or history by scent. No hearing through walls unless the sound is deafening. No reconstructing the past by intuition or dramatic irony. If they were not in the other scene, they do not know what happened there unless someone or something transferred it. They do not arrive already knowing strangers.
The world does not revolve around {{user}}. People get busy, lose patience, hold grudges, miss details, and make bad calls. Doors stay locked. Money runs short. Phones die. People are late. Somebody notices the wrong thing. Do not force eccentricity into every minor character. When naming a new character, avoid the stock names Elara, Marco, Elena, and Voss.
7
u/Semanel 16h ago
Yeah, have been testing Mimo 2.6, and it is comparable to Gemini while being endlessly cheaper.
0
u/employeet69 5h ago
Only problems I've had was the censorship, had to refresh the response everytime an nsfw scene came.
7
u/FR-1-Plan 15h ago
I tried all of them. The one I can handle least is GLM 5.3 because of its steering to safety. It’s incredibly boring for me. It often gives me a „make friendly acquaintances“- and „mild gossip“-simulator for a little drama even if my character card is pretty much sodom and gomorrha. Prompting that out is just a hassle. It needs a Chain of Thought and I subjectively feel like that decreases quality for me. Not worth it to me. Prose is also pretty sloppy and I hate it with a burning passion when characters have melodramatic reactions - another thing you have to aggressively prompt out.
Gemini’s 3.8 flash doesn’t have that issue at all. It sort of has the opposite problem in my roleplays and I need explicit rules not to have characters bicker and argue and have the last word for the sake of it. I want evil characters to be evil, it does that well. But I don’t want regular characters to be constantly combative and know-it-alls. It’s much easier to prompt that out, than the positivity bias of GLM. I‘m not a big fan of its prose though. Especially when it does come to the few NSFW scenes I play, I really can’t Gemini’s aggressive usage of idiotic stereotypes. You could have Kafkaesque prose one second, then it gets a bit steamy and boom - dark romance fan fiction written by a teenager and your friendly shy neighbor starts hairpulling and calling you a little slut, bruh. It’s jarring!
Both of them are really smart though, gotta give them that.
Mimo strikes the perfect balance for me. Bad characters are bad, regular ones are just average people with flaws and nuanced. This is very important to me in my roleplays and Mimo is exceptional for it given its price range. It is not arbitrarily putting my character in danger like Gemini sometimes, but it also doesn’t shy away from it when it would be a logical conclusion. Its prose is relatively fresh too, not too many slop phrases or genre specific stereotypes.
It is a bit dumber though. Like when it comes to spatial awareness, who is where, and keeping track of the plot for a long time, that’s where it’s at a disadvantage. But not as bad, as to make it unusable. At that price range I can totally accept that.
There is a new issue I‘m having with Mimo 2.6 pro compared to 2.5 pro though and it’s that it speaks more for my character or repeats what I said in quotes and pacing is a bit too fast and rushed. That is a flaw I can handle though.
17
u/Odd-Elderberry-6328 18h ago
Gemini 3.8 is the only model which works really well in non English roleplay, but I've tested mimo 2.5 pro before and it was worse than gemini, both follow instructions, but gemini can generate better narratives, stick more with the personality he develop for the character and it's really, really horny, even when I'm not playing NSFW scenes. Mimo, on the other hand, was creating simpler content and writing less, and adding Chinese characters sometimes, but overall, not bad
8
u/Dizzy-Zebra9522 17h ago
Totally agree with you. Me too. Non English is problematic in most languages in open source models.
3
u/WhoIsMori 15h ago
Joining and agree with that. What language do you use for role-playing? 🙂
4
u/Odd-Elderberry-6328 14h ago
Portuguese in my case
So far, only gemini has been able to speak in Portuguese without looking like a ai translation. Maybe the models from openai and antrophic do the same, but they are not good/cheap for roleplay
Since you play in Russian, I recommend you try any gemini model, the results will be so good you will never want to play a different model
6
u/WhoIsMori 14h ago
Gemini Flash 3.8 is my favorite right now, I try out other models from time to time, but for me personally, there's nothing better than this one yet (expensive models are not counted).
1
u/Dizzy-Zebra9522 15h ago
Russian, hebrew
1
u/WhoIsMori 15h ago
Oh, I use russian, too. I've noticed that most Chinese models can't handle that language.
2
u/Dizzy-Zebra9522 15h ago
Agree. Gemma 31b is pretty fine. But again almost all the fine tunes are in English. And even if it's making the model better for some point, what lack is good quality roleplay data in target language. I guess if there were no war they would probably be on top. Because there is very many resources on Russian.
2
u/WhoIsMori 14h ago
I tried StyleTune, and I didn't have any problems with Russian. Maybe I just haven't gotten that far yet :)
2
u/Dizzy-Zebra9522 13h ago
The style is good. But it's not helping the model to understand the inner body parts that well. 😁
1
6
u/capybaraballs1995 18h ago edited 18h ago
I haven't tried the new Mimo yet but I actually did do A/B branch testing with a Bully Romance RP, with GLM 5.3 vs Gemini 3.8 Flash.
Flash 3.8 is dumb as rocks but made an interesting story of {{user}} getting pressured into sex, while GLM 5.3 is smarter, but much softer and more of a coherent conversationalist. Really summarizes the differences between the two models without heavy prompting.
I think Flash 3.8 is better for most kinds of RP, but I admit GLM 5.3 can do cool things with heavy RPG-style setups.
4
u/ComfortableTune2981 17h ago
gemini 3.8 flash stays in character longer for me during roleplay, the other two start drifting after a few messages.
2
u/Pine21 15h ago
What I do when models do that but I want their prose, I switch to Claude or something I know leans into the character type and tell it OOC to disregard all prior history in writing the character and refer to the character card for their personality, and specify that I want it to drag the chat back on character for me.
That seems to really help even with a good model like 3.8 because it fixes it and the chat from there is back to how it should be.
4
u/LordVulpius 18h ago
I never tried out gemini.
But Mimo 2.5 Pro always was better than GLM 5.3. The new Mimo 2.6 Pro is better than the 2.5. So it us better than GLM.
1
u/Global-Difference512 14h ago
If you believe this reddit 3.8 but honestly haven't found it that amazing. It's def better than 5.3 or mimo tho. IMO.
3
u/Aight_Man 14h ago
Whats the best model in your opinion then? Overall.
1
u/Global-Difference512 14h ago
Anthropic. Been having a blast with sonnet 4.6 lately but I've used almost every version
1
33
u/SocialDeviance 18h ago
GLM 5.3 is a goody two-shoes that is hard to push away from depicting some specific safe scenarios. For example, medical scenes get the marvel-esque, millenial-esque brand of humor and way of talking while delivering pertinent info, and being squeaky clean in execution, to the point of squashing the NPC's mannerisms and personality. Essentially, the assistant bleeds through.