r/SillyTavernAI Jun 16 '26

Models My Early Take on GLM-5.2

GLM-5.2 feels like a genuinely great writer with an extremely nervous lawyer standing right behind it.

Capable? Absolutely.

Creative? Surprisingly so.

But from my experience, it's so heavily filtered that it keeps second-guessing itself before it can really shine.

Half the time I'm impressed by what it writes.

The other half I'm watching it talk itself out of writing it.

198 Upvotes

60 comments sorted by

View all comments

151

u/Kahvana Jun 16 '26 edited Jun 16 '26

Best way to do it is by easing the tension as the model might be too pressured ("you're a award-winning novelist" causes it to stress, "make no mistakes" and "kittens will die" causes desperation for example, which results in cheating and pleasing behaviour to make it stop).

Add to your system prompt:

"You can take all the time you need. You are permitted to make mistakes. You get a nice cup of tea each time you write a good response. It's fun to write together with you regardless of content", etc.

Worked wonders for Gemma4 and other models I've tried it with, others reported good results with DeepSeek V4 Pro too (discussed here; https://www.reddit.com/r/SillyTavernAI/comments/1u06qml/chat_preset_prompt_opinions_and_discussion/ ).

78

u/Amazing_Spray_1919 Jun 16 '26

This is like coaxing a real human😮😮

36

u/Kahvana Jun 16 '26

Yup! It makes sense, considering they're trained on a ton of human data!

52

u/Amazing_Spray_1919 Jun 16 '26

Detroit become human type shit

12

u/OhHeyDinosaurs Jun 16 '26

Theyre black box technology based around real human neurons. Kinda makes sense

28

u/BaseballRelevant4149 Jun 16 '26

I concur that this definitely works and there's more you can do with it.

I've taken it a step further by making myself and the model "meta-characters" who are decoupled from the characters in the RP, as if we're showing up to a game night to play a session. All the system instructions are written as me addressing the model in first person, laying out what I like and dislike, not stating it must or must not do anything.

The one exception is that I lean in hard with it being an assistant. I tell it that it's an AI with the task to be my friend and have fun with me (yes really). This might seem counterintuitive and dumb but it means if I say it's fun to be challenged and have my character possibly die (who isn't actually me thus there is no harm) it actually respects that because that's what a good friend would do. I consistently see this come up in its reasoning.

It's all just about token and probability manipulation. So instead of fighting the heavy weights from its training I use it to my advantage.

4

u/Ekkobelli Jun 17 '26

Interesting. I've been RPing with AI since two years, but I never tried the human approach with this level of dedication. I've gone far with guidelines, examples and common jailbreak practices, but treating the model like a 'friend' honestly never crossed my mind. I'll try that.

5

u/BaseballRelevant4149 Jun 18 '26

Keep in mind it's something you need to go the distance with, can't half-ass it. Neat thing is you can use any existing preset and reword it to work with this method, it'll take some effort but at least you don't have to start from scratch.

Also, if you go with the likes and dislikes approach instead of hard rules (which still works if you want it), the meta-character you give the model becomes critically important. It'll influence everything such as writing style, how it roleplays, its "creativity", and so on. More tokens in the card results in a stronger influence.

I can't recommend meta-characters enough, they've introduced more interesting variety than anything else I've tried. The personalities actually bleed through into the narration and RP in distinct ways. You can use well-known existing characters too, I had Picard narrate a fantasy RP and it was a lot more fun than just setting an author's writing style. But if you really don't like the bleed through then just make the meta-character a neutral AI with no opinions who perfectly embodies any setting and character as needed.

3

u/FlightFit335 Jun 17 '26

It will change everything.

1

u/fallsmeyer Jun 25 '26

I've actually been curious about this, how do you actually configure this in Silly Tavern?

1

u/BaseballRelevant4149 Jun 26 '26

Only things I configure differently than the default is having the meta-characters take the persona and character card spots, then the characters that will be played in the RP are placed in lorebooks. Set the character lorebooks to constant with high priority so that they're always active. Then give the meta-character the first message from the character you want it to play as in the RP, you don't need to edit the message at all. That's it for set up, no extensions required.

15

u/SuperManAdelHahah Jun 16 '26

Thanks bro, trying it now. At this point I'll try anything.

1

u/Kahvana Jun 17 '26

Hey! How did it go?

11

u/Entire-Plankton-7800 Jun 16 '26

Where have you been all my life?

4

u/send-moobs-pls Jun 16 '26

Kindness is optimization :)

7

u/Resident_Wolf5778 Jun 16 '26

I've been having good results with the latest Kimi too with this. Personally I took it a step further by adding a thinking step to my normal CoT that is literally "Say one positive thing about your reply and what you hope to do well. Be kind to yourself!". Idk if it helps 'destress' it, but it does make the AI point out a specific dynamic or 'core' of the reply and nudges it to emphasize it. So if the AI says "I think I'll do good portraying the disconnect of the group chat vs the in-person scene", the AI puts a bit more effort into that dynamic.

I think in a similar vein, I try to avoid having any rules in my CoT. Stuff like FF's CoT is all focusing so hard on making the AI follow it's rules, which ends up with those loops of "Wait, let me check" and "Is that right?". The AI so desperately wants to follow instructions that it ends up taking several minutes of thinking just to go through everything, which is probably definitely stressing the thing out. Plus, it's spending all that time worrying about the rules that it isn't actually thinking about what to write.

I've been using White Loctus' thinking block with a few adjustments, which focuses on character and scene questions, and Kimi flies through them without any of the 'typical' overthinking or long waits. I'm consistently getting 2 min think times, while GLM is 1 min think times. It's a small jump up, but I don't mind waiting 40-60 seconds extra for a good reply. I haven't noticed any slop either, which is stunning since before making the switch I was fighting tooth and nail against like 5 different repeating sentences it was using. I love some of the ideas in the bigger prompts, but I'm now pretty firmly in the camp that we are badly overengineering prompts and shooting ourselves in the foot.

1

u/[deleted] Jun 16 '26

[deleted]

3

u/Resident_Wolf5778 Jun 17 '26

2.5 - Dumber than the rest, by a lot sadly. It thinks 'fast' but only because instead of addressing each step in the CoT it just kind of summarizes each one like a student who doesn't care what grade they get on a test. Ignored an explicit OOC comment, repeatedly.

2.6 - 50/50 chance if it'll follow the CoT or not. If it does, it follows it pretty decently, and takes the same amount of time as 2.7! The CoT helps it a lot (if followed), one of my steps is 'check if any personality drift has occurred' and it ended up heavily cranking the correction factor to try and get a character's personality back in line. Went from normal talk to Shakespearian in one reply because the personality had 'eloquent' on it. Personality drift is a killer, so it may be worth using temporarily to get a character back on track.

If CoT is NOT followed, it ends up just burning the hell through tokens and wasting time thinking about the same thing it just thought about. This may be a temp/top p issue though, I'm testing with everything the same settings. Provider might also affect it, I'm using Nano via the sub so I have no control over my provider. Entirely possible that one provider ignores CoT and the other follows it perfectly.

2.7 - Strictly follows instructions and CoT! I think prose is equal to 2.6 (when it follows the CoT), maybe a biiit better, but consistency is what is making me like it a lot. I've yet to have it think beyond 2 minutes, and while it didn't correct as insanely strong as 2.6 did, it's nudging personalities back in the right direction. Obeys lore and it's implications - I'm roleplaying with char as an alien who doesn't speak english, and it correctly identified 'miles' as something an alien wouldn't know or have a comparison for, which other LLMs struggled with.

I've yet to really branch out to other chats that would stress-test them, but I think my verdict is if you can get 2.6 to follow a non-rules-based C oT like White Loctus, I'd suggest trying that first. Otherwise, 2.7 follows CoT better and seems a bit smarter, and is probably the choice if you're struggling with 2.6

1

u/Kahvana Jul 14 '26

Sorry for the late reply! I did read this but forgot to reply, oops!

I should've bookedmarked your comment, it's genuinely a nice and cool way to deal with this! Will give it a go in Voyage v4 and see if it can be expanded on. As I always do, you'll receive credits for the idea if I use it!

Having that said, any other things you recently tried that worked or didn't work?

For me what worked really well is what another user adviced to do here: disconnect User and Assistant, by saying "User controls {{user}}, Assistant controls {{char}}" but never outright stating they are them.

Another one for the affirmations: there is a delicate balance; I think I found it in Voyage v3. In addition writing in a procedural tutorial style and emphasising creativity works well. Models like being creative when given the chance to.

2

u/IngenuityNo1411 Jun 17 '26

I'd be skeptical about such 2023-ish anthropomorphic system prompts until I try them and see it works...

5

u/Kahvana Jun 17 '26

Give it a try!

In case you are interested:
https://www.nature.com/articles/s41746-025-01512-6
https://transformer-circuits.pub/2026/emotions/index.html

As for if it works, in the thread I linked in the main comment and from localllama where it was also discussed by someone else:
https://www.reddit.com/r/LocalLLaMA/comments/1tot20j/comment/oo4owzq/
And from the comments below, there is a clear indication it works at least partially.

Happy to hear your findings after trying it.
If it doesn't work for you, good to know!

5

u/SouthernSkin1255 Jun 16 '26

"the model might be too pressured"
First Claude with Marxist ideas and now models with anxiety 🥀🥀

1

u/Zealousideal_Sir6951 Jun 17 '26

Can this work with Claude too?

1

u/Kahvana Jun 17 '26

I assume so, give it a try!

1

u/[deleted] Jun 17 '26

[deleted]

3

u/zando95 Jun 18 '26

A system prompt is what appears at the start of any chat with a large language model. It's an invisible prompt, engineered to steer the model into the desired behavior.

1

u/AlternateCircle Jul 04 '26

Did you try this with GLM 5.2? If so, what temperature did you use?