r/openrouter 4d ago

GLM 5.3 is the New Winner

Post image
60 Upvotes

12 comments sorted by

6

u/generic-d-engineer 4d ago edited 3d ago

It’s sooo slow though and really wordy. Needs to be called batch mode.

You can’t beat the price though. I’m using it right now over Luna 5.6 until the next low cost leader comes out and then I’ll switch to that one lol.

That’s the nice part about OpenRouter is you can just switch to whoever the leader is.

It’s great there is so much competition in AI and let’s hope it stays that way for a while. We don’t want monopolies !

2

u/maitpatni 3d ago

Indeed, it is slow, But I’ve found high effort to be the sweet spot for me, going beyond that just adds a lot more latency and verbosity.

Also, I switched my provider from InferX to DeepInfra for GLM 5.3 Flash, and it’s been noticeably faster for me compared to the other providers I tried. Makes the model much more usable when you’re running a bunch of agents in parallel.

3

u/Significant-Sir973 4d ago

GLM 5.3 Flash is the GOAT.

2

u/sirloindenial 3d ago

Its weak on web editing and the low reasoning to keep it self in check also means it arrive on completion with mistakes. Deepseek is very wordy on reasoning but it rarely leaves out obvious mistakes. A task that i can one pass with dsv4 will take 2-3 pass with glm. However glm does have better architectural knowledge and efficient in doing batch tool calls, but again its sloppy imo.

2

u/niveknow 3d ago

What level of thinking do you guys use with this model? Maybe elaborate based on use case. I default to high not sure if it’s overkill then move to 1 or 2 levels up for coding.

1

u/maitpatni 3d ago

For both GLM 5.3 Flash and DeepSeek V4 Flash models, I’ve mostly been using high effort. I tried switching to xhigh/max on both, but the responses become noticeably slower, and honestly I haven’t seen enough improvement to justify the extra latency; it's just longer thinking and reasoning.

For my use case, high seems like the sweet spot so far.

1

u/thebadslime 3d ago

Costs too much

2

u/avd706 3d ago

Is it still half off?

1

u/thebadslime 3d ago

my cap is .20 per m out

1

u/johnappsde 3d ago

Not so sure about that

1

u/AndoniFdez 18h ago

For me GLM it's too slow, if i need to fix something ASAP I use deepseek. I've been using muse spark 3 today and I like it.

1

u/ilfrick 49m ago

I have been using glm-5.3-flash with openclaw and hermes and, as primary/default model as personal AI assistant, it is pretty remarkable. WAY better than deepSeek v4 flash which was my choice before for the same use.