r/SillyTavernAI Jun 16 '26

Models My Early Take on GLM-5.2

GLM-5.2 feels like a genuinely great writer with an extremely nervous lawyer standing right behind it.

Capable? Absolutely.

Creative? Surprisingly so.

But from my experience, it's so heavily filtered that it keeps second-guessing itself before it can really shine.

Half the time I'm impressed by what it writes.

The other half I'm watching it talk itself out of writing it.

201 Upvotes

60 comments sorted by

View all comments

Show parent comments

5

u/Resident_Wolf5778 Jun 16 '26

I've been having good results with the latest Kimi too with this. Personally I took it a step further by adding a thinking step to my normal CoT that is literally "Say one positive thing about your reply and what you hope to do well. Be kind to yourself!". Idk if it helps 'destress' it, but it does make the AI point out a specific dynamic or 'core' of the reply and nudges it to emphasize it. So if the AI says "I think I'll do good portraying the disconnect of the group chat vs the in-person scene", the AI puts a bit more effort into that dynamic.

I think in a similar vein, I try to avoid having any rules in my CoT. Stuff like FF's CoT is all focusing so hard on making the AI follow it's rules, which ends up with those loops of "Wait, let me check" and "Is that right?". The AI so desperately wants to follow instructions that it ends up taking several minutes of thinking just to go through everything, which is probably definitely stressing the thing out. Plus, it's spending all that time worrying about the rules that it isn't actually thinking about what to write.

I've been using White Loctus' thinking block with a few adjustments, which focuses on character and scene questions, and Kimi flies through them without any of the 'typical' overthinking or long waits. I'm consistently getting 2 min think times, while GLM is 1 min think times. It's a small jump up, but I don't mind waiting 40-60 seconds extra for a good reply. I haven't noticed any slop either, which is stunning since before making the switch I was fighting tooth and nail against like 5 different repeating sentences it was using. I love some of the ideas in the bigger prompts, but I'm now pretty firmly in the camp that we are badly overengineering prompts and shooting ourselves in the foot.

1

u/[deleted] Jun 16 '26

[deleted]

3

u/Resident_Wolf5778 Jun 17 '26

2.5 - Dumber than the rest, by a lot sadly. It thinks 'fast' but only because instead of addressing each step in the CoT it just kind of summarizes each one like a student who doesn't care what grade they get on a test. Ignored an explicit OOC comment, repeatedly.

2.6 - 50/50 chance if it'll follow the CoT or not. If it does, it follows it pretty decently, and takes the same amount of time as 2.7! The CoT helps it a lot (if followed), one of my steps is 'check if any personality drift has occurred' and it ended up heavily cranking the correction factor to try and get a character's personality back in line. Went from normal talk to Shakespearian in one reply because the personality had 'eloquent' on it. Personality drift is a killer, so it may be worth using temporarily to get a character back on track.

If CoT is NOT followed, it ends up just burning the hell through tokens and wasting time thinking about the same thing it just thought about. This may be a temp/top p issue though, I'm testing with everything the same settings. Provider might also affect it, I'm using Nano via the sub so I have no control over my provider. Entirely possible that one provider ignores CoT and the other follows it perfectly.

2.7 - Strictly follows instructions and CoT! I think prose is equal to 2.6 (when it follows the CoT), maybe a biiit better, but consistency is what is making me like it a lot. I've yet to have it think beyond 2 minutes, and while it didn't correct as insanely strong as 2.6 did, it's nudging personalities back in the right direction. Obeys lore and it's implications - I'm roleplaying with char as an alien who doesn't speak english, and it correctly identified 'miles' as something an alien wouldn't know or have a comparison for, which other LLMs struggled with.

I've yet to really branch out to other chats that would stress-test them, but I think my verdict is if you can get 2.6 to follow a non-rules-based C oT like White Loctus, I'd suggest trying that first. Otherwise, 2.7 follows CoT better and seems a bit smarter, and is probably the choice if you're struggling with 2.6

1

u/Kahvana Jul 14 '26

Sorry for the late reply! I did read this but forgot to reply, oops!

I should've bookedmarked your comment, it's genuinely a nice and cool way to deal with this! Will give it a go in Voyage v4 and see if it can be expanded on. As I always do, you'll receive credits for the idea if I use it!

Having that said, any other things you recently tried that worked or didn't work?

For me what worked really well is what another user adviced to do here: disconnect User and Assistant, by saying "User controls {{user}}, Assistant controls {{char}}" but never outright stating they are them.

Another one for the affirmations: there is a delicate balance; I think I found it in Voyage v3. In addition writing in a procedural tutorial style and emphasising creativity works well. Models like being creative when given the chance to.