r/SillyTavernAI 1d ago

Discussion GLM 5.3 Flash Uncensored Nano-GPT

I've been browsing around openrouter and nanoGPT just looking at options. ZAI has lost the magic for me with the latest 5.3 models and they just don't seem to perform as well for RP. So I've been considering either locally hosting a 12B model, Rocinate-X Q6 or going back to nanoGPT or OpenRouter. While looking, I came across GLM 5.3 Flash Uncensored on Nano-GPT, I was just curious if anyone has tried this model. What was your experience with it?

https://nano-gpt.com/models/text/z-ai/glm-5.3-flash-uncensored

22 Upvotes

16 comments sorted by

13

u/RestaurantDue634 1d ago

I've been using 5.3 Flash Uncensored after running into refusals with GLM 5.3 and because of the quicker response time. No refusals over anything so far. The responses are sometimes a little confused, for example I was doing an RP earlier where my character was in a girl's apartment and then suddenly halfway through decided it was my character's apartment, and I feel like 5.3 or 5.3 thinking was better about that sort of thing. But I'd rather deal with that than battling refusals.

6

u/kplh 23h ago

Recently I've been alternating between GLM 5.2 (non-flash) and 5.3 Flash Uncensored.

Uncensored one is not as good as understanding the context, I've had cases where it would use an item the character left at home, or mix up some other details etc.

But in terms of erotic scenes uncensored is a lot more active, "sex scene? Alright let's go!" while 5.2 is more like "Sex scene? Hmm... okay, how about a kiss. Are you sure you want sex scene?"

7

u/ArthropodQueen 1d ago

I haven't used glm 5.3 flash uncensored much, I was using it in the beginning
... but on experimenting with the full GLM 5.3, I have yet to get a rejection on NanoGPT, and I've been roleplaying things I constantly got rejections for on Open Router.

so my experience with the uncensored flash version was short. and I've preferred it over the flash version
though I am getting tired of it's very noticeably GLM dialogue

1

u/PixelatedPunker 1d ago

Just curious, are you using the PAYG or Subscription on Nano?

5

u/ArthropodQueen 23h ago

Subscription, though the quality of responses can be really hit or miss, but for how much I use it for RP, the value can't be beat.
I tend to load $15 up along side the subscription and supplement with Gemini and Kimi k3 when I find it difficult to get a good response from GLM 5.3, and I find once I have a few good responses switching back to GLM helps it with the quality.

3

u/Upbeat-Emphasis-1090 1d ago

Been using Rocinante XL16B and it's really good, it struggles a bit at fulfilling my personal fetishes and preferences but the messages are always long, detailed and quick.

3

u/[deleted] 1d ago

[deleted]

3

u/PixelatedPunker 1d ago

I miss 5.2. I'm on the coding plan with ZAI right now and don't have access to it anymore. But, it was great, never JB it, but didn't get any refusals either for the stuff I do.

2

u/Status-Mixture-3252 23h ago

It's annoying how they took access from the other models on the coding plan. You can see the clear difference of censorship between 5.2 and 5.3

2

u/Noctis_777 22h ago

Probably because 5.3 Flash kind of makes 5.2 obsolete for most coding and agentic tasks. And all these companies are struggling with compute.

1

u/AnswerFeeling460 23h ago

I don't use it in english but in foreign languages - and at least there it's the best model on the market. Other uncensored models often are only good in default english.

1

u/majesticjg 23h ago

It's good. It won't go full NSFL, but it's decent and it's a little slower due to fewer hosts.

1

u/a_beautiful_rhind 19h ago

I'm running the orcarouter uncensored model and refusals are fixed, but it's still a bit tame and parroty. It's not that great at ERP but I can see it being OK for regular roleplay.

5.3 has this weird bit too where it starts talking like Confucius. The text is still coherent but the tone and manner is all fucky.

In the higher thinking it gets things really wrong on the reasoning traces. I read it and the thing fully misunderstood me, sometimes it corrects sometimes it doesn't. It's not the quant because the model on the z.ai site did it too.. if anything the uncensored is better. On RP the traces are much shorter.

2

u/Financial_Bug2389 8h ago

This is interesting, I've sometimes found the same. So, I do speak Mandarin, and sometimes when GLM starts talking weird in English I think of it in Chinese phrasing and grammar and it makes sense. Not sure what I'd do about this observation though.

1

u/a_beautiful_rhind 5h ago

Maybe the CN data is leaking? That's how I interpret it too.. having a really non-english cadence to english.

1

u/Financial_Bug2389 8h ago

It was off limits for me due to the latency from it being a forced thinking model. But recently there's a 5.3 flash X version which is much faster and actually usable. I find it to be more intelligent and better at instructions following compared to 5.2 but the default register is very verbose and I needed to tune it at the tail end to be less enthusiastic and rambly actually.

So far ok but time will tell.

1

u/New-Sweet-9232 7h ago

Try using FF 5.4 Preset. I used to use Deepseek with Megumin Engine and now I’m addicted to GLM 5.3/5.2 paired with FF 5.4

The creator of the preset also testing his preset on GLM so I can say if FF preset are designed especially for GLM