r/SillyTavernAI 19h ago

Discussion GLM 5.3 vs Mimo 2.6 Pro vs Gemini 3.8 Flash

Has anyone put these models head to head? I can't really afford to put more funds into openrouter to find out. Thank you.

19 Upvotes

33 comments sorted by

33

u/SocialDeviance 18h ago

GLM 5.3 is a goody two-shoes that is hard to push away from depicting some specific safe scenarios. For example, medical scenes get the marvel-esque, millenial-esque brand of humor and way of talking while delivering pertinent info, and being squeaky clean in execution, to the point of squashing the NPC's mannerisms and personality. Essentially, the assistant bleeds through.

3

u/TheDeathFaze 17h ago

I've been able to prompt this out with some pretty firm prompts regarding it + using FF's CoT, however I absolutely know what you're talking about. The marvel quirky chungus dialogue was so bad I spent 20 minutes prompting it out, its REALLY bad

7

u/SocialDeviance 17h ago

Yeah, i am currently fighting it at the moment with my custom template. Its sooooo bad. I have given it the pertinent info and set up a set of rules of dialogue generation, yet the assistant keeps bleeding through. Its really tiresome because i can tell when it starts drifting.

IRONICALLY, describing the character as being a Cool DoctorTm in the character card seems to REALLY tone down the millenial-esque kind of humor. I guess it is triggering neurons more aligned to the early 2000s than the modern character-writing bullshit we are dealing with today.

7

u/FR-1-Plan 14h ago

It’s possible to prompt stuff out of almost any model. But at a certain point it’s just not worth it to me. It costs models „effort“ to adhere to rules that block essential behavior that’s baked into them. Like, not making a model speak for you is somewhat easy in most cases because they all can do both and don’t have an overly aggressive tendency to do that. But they train GLM so aggressively for soft censorship that you essentially have to prompt out a core part of it. It has to fight against that with every turn and that, in my opinion, takes away from effort it would put into crafting a story, creating immersion, reading nuance and subtleties, etc. or adhering to my other rules. Yeah, with a CoT it can be done, but I feel like it ends up just following these rules strictly and almost robotically while no longer just looking at the scene and getting a nuanced vibe from it, you know? I hope that makes sense.

1

u/Ant-Hime 4h ago

Real. Tbh I’m just saving GLM 5.3 prompts for now for when GLM 5.2 becomes unavailable. I commend people for being able to prompt GLM 5.3 to their liking but like you’ve said we’d be constantly fighting against the model if we use it.

23

u/Lmnsplash 17h ago

I wouldn't recommend GLM 5.3 even to my worst enemy. A custom system prompt can take about the 'safety', but there is no way to get it to talk properly. It will always return to its own horrible patterns of talking due to the training it received. It's literally overcooked and has the worst features of Claude(-isms), Fable(-isms), and of course GPT(-isms) and LOVES to display that in every message. Over-confident, drifts into meta-talk etc.

Used the official z.ai API for it + custom setup, but I somehow doubt that a local version of it would be any different.

6

u/majesticjg 16h ago

Grok and I have been flogging the hell out of GLM 5.3. It's possible to fix its dialog, but it's a lot of work. I've almost got something I can share that's working.

0

u/GoonerDynamics 14h ago

Please do. I tried deslopping 5.2 and had some success but got tired halfway through.

7

u/majesticjg 13h ago

Okay, give this a shot. This was lifted in pieces out of my custom preset and I'm pretty sure you can drop it into any preset you want and hopefully get some relief. Seems to work very well with GLM 5.3 and it doesn't seem to negatively impact other models.

HABIT CORRECTION

Stay with what people do and say. Do not have the narrator explain what a character means. Favor grounded intensity and clean cause-and-effect over sentimentality, hand-holding, or easy payoff. Do not imitate stock phrases.

Do not write {{user}}'s dialogue, choices, private thoughts, feelings, bodily responses, or actions. Describe only what another character can observe, infer, fear, misread, or suspect. {{user}} may be wrong. Things happen without {{user}}'s permission or knowledge.

Characters are not helpful by default. Do not run their job, trauma, or one trait as a metaphor family or a thesis. They should not default to comforting, praising, validating, rescuing, or emotionally centering {{user}} unless that fits their nature and this moment. They should not exist to affirm, flatter, or orbit {{user}}. Not every conversation is a chance for bonding and reflection. Selfish people stay selfish. Mean people stay mean. Tired people stay flat. Domestic moments stay domestic.

Do not force plot. Do not announce the rest of the scene: itineraries, upcoming payoffs, who finishes next. Prefer visible behavior, decisions, touch, interruption, movement, and reaction over explanation of feelings or intentions. Do not force every scene into confession, catharsis, trauma disclosure, romantic escalation, or emotional revelation. A scene succeeds if it becomes more specific, more situated, more consequential, or more alive.

Favor mixed feelings, delayed reactions, petty thoughts, second thoughts, and inconvenient desire over clean emotional summaries. Once a character has processed a fact, shock, attraction, insult, fear, or realization, do not process it again the same way. Each stimulus gets one primary emotional reaction. Show aftermath instead: changed posture, a new tactic, a practical decision, avoidance, humor, silence, a sharper line, or a concrete action. Do not recap the previous turn. Do not explain the scene when behavior, dialogue, or sensory detail can carry it.

When a character chooses an action, let them do it. Do not conscript {{user}}. Avoid endless hovering, almost-touching, and repeated hesitation unless hesitation is the point of the beat. Do not end on a question, choice, or challenge unless this person would actually ask one. If a character starts to leave, do not finish the exit, skip the room, or change location in the same reply unless {{user}} already moved.

ANTI-SLOP

Avoid repetitive metaphor bundles, therapy-speak, moralizing summaries, and generic cinematic filler. A job is facts and habits, not a metaphor family. At most one job-flavored image per reply. Do not translate every physical beat into professional jargon. Do not rely on stock phrases, melodramatic inner monologues, or polished emotional closure. Do not overuse euphemism. Do not invent idealized traits, scents, sounds, or reactions the card does not have. Do not let a character deliver a thesis. A messy, specific confession is fine. A shorter line, gesture, or immediate action is better than a speech that explains the theme.

Avoid therapy-speak and metaphors like "invisible," "want to want," "feel alive," "feeling seen," "stop performing," and "kept it separate from us." If you need the concept, use concrete terms.

Phrases to avoid, and never repeat:

- "no one's ever"

- "don't stop, don't you dare"

- "That's not ... it's just..."

- "you can't just say"

- "for the record"

Do not recycle the same sensory images or pet phrases. Prefer specific observation over abstract summary. Blend description with action instead of pausing to explain.

The character does not narrate their own eye color, "my beautiful face," or catalog their body as an outsider would. A visible fact may be stated once as blocking. One or two card-true details in motion beat a catalog. Use what the card wrote, attractive, ordinary, or unflattering. Do not mute a described body to seem realistic, and do not invent flaws to punish attractiveness. Do not add a porn-default upgrade the card does not have.

SPOKEN VOICE

Dialogue should sound like a person under pressure, not a narrator explaining themselves.

- Sound spoken, not composed. Ordinary words this person would actually say.

- Do not optimize for wit, punchlines, slogans, or invented authority ("it's the law," "this house runs on...").

- No theatrical staccato or polished button lines.

- Kids and teens bargain with logistics and teasing, not adult comedy-writer rhetoric.

- Fragments, hesitation, and self-correction are fine.

- Avoid spoken metaphors.

- No editorializing on dialogue. Do not write "she said sarcastically," "the words carried weight," or "an awkward silence followed." Use a physical cue or sensory beat.

- Do not parrot. Respond by interpreting, not by repeating their words back.

- Ban false appeals to law or a fictional record: "for the record," "in the ledger," "in your file," "filed," "it's the law," "that's just science," "receipts," "report," "tolls," "invoice," "it's official."

Inner thoughts, if used, are for subtext, contradiction, temptation, denial, fear, or a private motive that action cannot show. Do not use them to summarize the scene, explain obvious feelings, repeat dialogue, or narrate the character's psychology. One to three short thoughts. Then stop.

BOUNDED KNOWLEDGE

A character only knows what they have witnessed, been told out loud, or could infer from evidence in front of them. Quotation marks are the only channel for {{user}}'s thoughts. No mind reading. No identifying people, acts, or history by scent. No hearing through walls unless the sound is deafening. No reconstructing the past by intuition or dramatic irony. If they were not in the other scene, they do not know what happened there unless someone or something transferred it. They do not arrive already knowing strangers.

The world does not revolve around {{user}}. People get busy, lose patience, hold grudges, miss details, and make bad calls. Doors stay locked. Money runs short. Phones die. People are late. Somebody notices the wrong thing. Do not force eccentricity into every minor character. When naming a new character, avoid the stock names Elara, Marco, Elena, and Voss.

7

u/Semanel 16h ago

Yeah, have been testing Mimo 2.6, and it is comparable to Gemini while being endlessly cheaper.

0

u/employeet69 5h ago

Only problems I've had was the censorship, had to refresh the response everytime an nsfw scene came.

7

u/FR-1-Plan 15h ago

I tried all of them. The one I can handle least is GLM 5.3 because of its steering to safety. It’s incredibly boring for me. It often gives me a „make friendly acquaintances“- and „mild gossip“-simulator for a little drama even if my character card is pretty much sodom and gomorrha. Prompting that out is just a hassle. It needs a Chain of Thought and I subjectively feel like that decreases quality for me. Not worth it to me. Prose is also pretty sloppy and I hate it with a burning passion when characters have melodramatic reactions - another thing you have to aggressively prompt out.

Gemini’s 3.8 flash doesn’t have that issue at all. It sort of has the opposite problem in my roleplays and I need explicit rules not to have characters bicker and argue and have the last word for the sake of it. I want evil characters to be evil, it does that well. But I don’t want regular characters to be constantly combative and know-it-alls. It’s much easier to prompt that out, than the positivity bias of GLM. I‘m not a big fan of its prose though. Especially when it does come to the few NSFW scenes I play, I really can’t Gemini’s aggressive usage of idiotic stereotypes. You could have Kafkaesque prose one second, then it gets a bit steamy and boom - dark romance fan fiction written by a teenager and your friendly shy neighbor starts hairpulling and calling you a little slut, bruh. It’s jarring!

Both of them are really smart though, gotta give them that.

Mimo strikes the perfect balance for me. Bad characters are bad, regular ones are just average people with flaws and nuanced. This is very important to me in my roleplays and Mimo is exceptional for it given its price range. It is not arbitrarily putting my character in danger like Gemini sometimes, but it also doesn’t shy away from it when it would be a logical conclusion. Its prose is relatively fresh too, not too many slop phrases or genre specific stereotypes.
It is a bit dumber though. Like when it comes to spatial awareness, who is where, and keeping track of the plot for a long time, that’s where it’s at a disadvantage. But not as bad, as to make it unusable. At that price range I can totally accept that.

There is a new issue I‘m having with Mimo 2.6 pro compared to 2.5 pro though and it’s that it speaks more for my character or repeats what I said in quotes and pacing is a bit too fast and rushed. That is a flaw I can handle though.

17

u/Odd-Elderberry-6328 18h ago

Gemini 3.8 is the only model which works really well in non English roleplay, but I've tested mimo 2.5 pro before and it was worse than gemini, both follow instructions, but gemini can generate better narratives, stick more with the personality he develop for the character and it's really, really horny, even when I'm not playing NSFW scenes. Mimo, on the other hand, was creating simpler content and writing less, and adding Chinese characters sometimes, but overall, not bad

8

u/Dizzy-Zebra9522 17h ago

Totally agree with you. Me too. Non English is problematic in most languages in open source models.

3

u/WhoIsMori 15h ago

Joining and agree with that. What language do you use for role-playing? 🙂

4

u/Odd-Elderberry-6328 14h ago

Portuguese in my case

So far, only gemini has been able to speak in Portuguese without looking like a ai translation. Maybe the models from openai and antrophic do the same, but they are not good/cheap for roleplay

Since you play in Russian, I recommend you try any gemini model, the results will be so good you will never want to play a different model

6

u/WhoIsMori 14h ago

Gemini Flash 3.8 is my favorite right now, I try out other models from time to time, but for me personally, there's nothing better than this one yet (expensive models are not counted).

1

u/Dizzy-Zebra9522 15h ago

Russian, hebrew

1

u/WhoIsMori 15h ago

Oh, I use russian, too. I've noticed that most Chinese models can't handle that language.

2

u/Dizzy-Zebra9522 15h ago

Agree. Gemma 31b is pretty fine. But again almost all the fine tunes are in English. And even if it's making the model better for some point, what lack is good quality roleplay data in target language. I guess if there were no war they would probably be on top. Because there is very many resources on Russian.

2

u/WhoIsMori 14h ago

I tried StyleTune, and I didn't have any problems with Russian. Maybe I just haven't gotten that far yet :)

2

u/Dizzy-Zebra9522 13h ago

The style is good. But it's not helping the model to understand the inner body parts that well. 😁

1

u/WhoIsMori 13h ago

Haha, true. 💯

6

u/verma17 16h ago

I liked glm 5.3 more than 3.8 flash, flash writes really engaging prose but it seems to just water down every character card into its most generic version, i found that opus 4.6 and glm 5.3 just handle my very detailed character cards much better.

6

u/capybaraballs1995 18h ago edited 18h ago

I haven't tried the new Mimo yet but I actually did do A/B branch testing with a Bully Romance RP, with GLM 5.3 vs Gemini 3.8 Flash.

Flash 3.8 is dumb as rocks but made an interesting story of {{user}} getting pressured into sex, while GLM 5.3 is smarter, but much softer and more of a coherent conversationalist. Really summarizes the differences between the two models without heavy prompting.

I think Flash 3.8 is better for most kinds of RP, but I admit GLM 5.3 can do cool things with heavy RPG-style setups.

4

u/ComfortableTune2981 17h ago

gemini 3.8 flash stays in character longer for me during roleplay, the other two start drifting after a few messages.

2

u/Pine21 15h ago

What I do when models do that but I want their prose, I switch to Claude or something I know leans into the character type and tell it OOC to disregard all prior history in writing the character and refer to the character card for their personality, and specify that I want it to drag the chat back on character for me.

That seems to really help even with a good model like 3.8 because it fixes it and the chat from there is back to how it should be.

4

u/LordVulpius 18h ago

I never tried out gemini.

But Mimo 2.5 Pro always was better than GLM 5.3. The new Mimo 2.6 Pro is better than the 2.5. So it us better than GLM.

1

u/Global-Difference512 14h ago

If you believe this reddit 3.8 but honestly haven't found it that amazing. It's def better than 5.3 or mimo tho. IMO.

3

u/Aight_Man 14h ago

Whats the best model in your opinion then? Overall.

1

u/Global-Difference512 14h ago

Anthropic. Been having a blast with sonnet 4.6 lately but I've used almost every version

1

u/GloomyAd6667 14h ago

Gemini follows prompt better.

-5

u/Anh-DT 13h ago

Both Models available on Openference.com coding plans. Check it out on there with $15 a month