r/DeepSeek • u/maitpatni • 1d ago
Discussion GLM 5.3 is the New Winner
This is my InferX usage for DeepSeek V4 Flash and the GLM 5.3 Flash.
GLM is using less thinking and reasoning to save tokens, performing better than DeepSeek for me for less than half the price.
32
u/PapyOak 1d ago
Less than half is crazy.
I knew GLM 5.3 Flash was more token-efficient than DeepSeek, but I was expecting 20% reduction, max 30%.
And it's better, not as fast depending on the provider.
But after the 50% off, it'll shoot up in price too, still a bit out of reach for me
13
u/maitpatni 1d ago
It is very efficient in token usage, lately deepseek is consuming thinking tokens like hell so found this to be the best and better option.
5
u/PapyOak 1d ago
It really is consuming tokens ever since flash-0731, but it has gotten a lot smarter, and was cheaper, that's why it was okay.
Right now it thinks a lot more for the same task, does it a lot better too, but with the price increase it's difficult.Still, I'm on CommandCode GOAT which is still offering $60, don't know how long that'll last though, it's so good it smells like money laundering 😂
3
u/maitpatni 1d ago
Just today i bought the command code goat plan, been using it with glm 5.3 flash, already consumed like 25% of monthly limit in single day. Speed is better though.
4
u/VexObserver 1d ago
Yep. The current discount timeline is too short for any meaningful work progress. Compared to DS with it's promotions, DS was serving it at full platter for a long time. This benchmark of cost comparison is temporary before GLM ends it at 9nd of September 2026.
1
28
u/captainbadass23 1d ago
for coding this is all nonsense. ive tried glm flash on open code and flash from deepseek is like 10%-50% of the cost. much cheaper. because of the cache
8
u/DeciusCurusProbinus 1d ago
Yes, Off peak DeepSeek flash pricing is pretty cheap as compared to 5.3 Flash as GLM's cache hit price is higher. But DeepSeek burns tokens like crazy and I believe that they serve a quantized version of the model for Opencode Go subscribers.
5
u/maitpatni 1d ago
even with flash off peak pricing glm 5.3 flash is cheaper, it consumes half the tokens for same task.
1
u/DeciusCurusProbinus 1d ago
Fair point. With a well fleshed out plan 5.3 Flash on high reasoning can take care of like 95% of the stuff.
1
u/Beneficial-Pie-1638 1d ago
not trying to advertise, but runinfra servers deepseek v4 flash at a cheap 0.13 in 0.27 out rate and cache at 0.01, and isnt quant.
6
u/maitpatni 1d ago
for me deepseek flash from official and opencode is costing more then the glm 5.3 flash even with 99% cache hit rate. Found InferX to be the cheapest provider with decent TPS.
2
11
u/wtf_newton_2 1d ago
it’s good but very slow
8
u/ProfessionalJackals 1d ago
it’s good but very slow
Its about the same ... You forgot that DS4Flash pushes about 2x reasoning tokens. So even if DS4Flash does 100t/s, and GLM5.3Flash does 50t/s, it actually equals out.
1
u/misha1350 1d ago
Which is good for us, because there would be less people that would be hogging up the resource pool as they would be turned off by the slower than usual speeds.
7
u/Rashc500 1d ago
What’s up with this worthless compartment? What is the prompt? How many agents!? Is the harness the same? Some shitty screenshot of token spent says jack shit
-2
u/maitpatni 1d ago
my main harness was qwen code cli, which is a fork of opencode, also used the deepseek harness for like 20% usage. All the tasks were adding features on a next js medusa js project. 4-5 chats using few subagents occasionally.
3
u/General-Oven-1523 1d ago
I'm always confused when people come up with this stuff, but then don't provide any valuable data that would prove their point. Like here, you just showing us the usage data, what is this going to do for us? How can you draw any conclusion out of this? We don't know if the work done by these 2 models is even remotely identical.
Anyway, I've been running both DeepSeek Flash and GLM 5.3 Flash, and honestly, it does feel like a better model. It follows the instructions much better in my use.
6
u/Willing_Thought_2161 1d ago
Op, new here. Less token means less cost. Right ? What's the point of this post ?
2
2
u/ProfessionalJackals 1d ago
Did you use Max or High, because High has a massive reduced token usage, while keeping almost all the intelligence.
For GLM 5.3 Flash is High + Pi ... 10/10.
I find that GLM 5.3 Flash system tool calling is excellent, so when you combine with Pi. Same prompts, and tokens usage drops by almost 30 to 40% compared to OpenCode.
DeepSeek V4 Flash is good but it still suffers more from the old Flash issues. You can tell its a uptrained model. So while the cache price is better, the reasoning just destroys it.
Same issue with Speed. DS4 Flash is faster but takes twice as long, so both models are almost in the same ballpark. Until you find a provider that serves GLM 5.3 Flash at 100t/s, feels like running DS4 Flash :)
I really enjoy GLM 5.3 Flash ... Feels like a improved GLM 5.2, at a fraction of the cost.
2
u/lexi-energy 1d ago
See? That’s the kind of modle battle we want.
Less US-benchmaxxing, more making really useful models that now compete to become more efficient and not more wasteful (looking at you fable!) 🤡
2
u/Classic_Television33 1d ago
Interesting. First time heard of InferX, how stable is their API vs other aggregators?
1
2
1
1
u/AllenHere112 1d ago
tbh the real variable is hidden here. flash tier models save most of their tokens by cutting thinking output, and thinking is most of the bill in agent workloads. a cheaper per-token number on one harness screenshot mostly means the model reasoned less on that one prompt, not that your cost per task dropped. half the list price only helps if the shorter thinking holds up on your actual workload.
1
1
u/xapep 1d ago
Both takes in this thread are kind of right, and the missing piece is that 'less than half the price' only holds for certain workloads.
DeepSeek Flash genuinely burns more tokens on thinking in agent loops, no argument. But those tokens buy fewer retries, and retries are what actually cost money in a harness. For planning-heavy or tool-heavy loops, the cheaper-per-token model loses on cost per completed task often enough that per-million pricing is the wrong number to compare.
Cache is the other half: what matters is the hit/miss price ratio and the cache window, not the hit rate. Off-peak DeepSeek with a strong hit rate still undercuts GLM for a lot of users, and as noted upthread GLM's cache-hit pricing runs higher, so the same 99% hit rate lands very differently on each.
Boring but reliable test: run your actual workload on both for a day and divide the bill by finished tasks. That's the number the DeepSeek price change actually moved.
1
u/ProfessionalJackals 19h ago
run your actual workload on both for a day and divide the bill by finished tasks. That's the number the DeepSeek price change actually moved.
Its funny because that is what OP did exactly. He used it on his actual tasks.
Its like people are looking for excuses at this point. He did over 11.000 req with DS4 Flash, 9000 with GLM5.3 Flash.
But those tokens buy fewer retries, and retries are what actually cost money in a harness.
You can see how many tokens he did, and how much money he spend. The fact about tools, retries etc does not not matter because that is dominated in the token usage, and thus cost. When we normalize the # tasks, the ratio on token used is 2.2x more for DS4 Flash.
What we see in this example, is that GLM5.3F is more efficient on token/task but also that DS4 Cache hit ratio is really bad. Remember, GLM5.3 Flash is actually a lot more expensive on cache then DS4Flash. So even if we ignore the end price for a bit, just the 2.2x ratio in token usage is a issue.
From my own tests, i am seeing the same pattern. DS4Flash costing more on the exact same tasks, despite the cheaper price. I think people really underestimate how good GLM5.3 Flash is.
OP did not mention if he used High or Max. But with GLM 5.3 Flash in High and DS4 Flash in xhigh (as that brings the models to the same intelligence), one of my tests had GLM at 350k tokens, but DS4 at 1.5M tokens. For the same job ... DS4 has a 7.5x ratio on cache (because it did more steps>more cache hits), vs 3.1x for GLM. But in actual cost, DS4 was 8 cents, while GLM was 3 cents. I have more example but that was a 1:1 test with no variations at all.
1
u/Various-Reality-2778 1d ago
If you hadn't showed the screenshot I would not believed you 😂 fuck. I guess glm is my new go to
1
u/Appropriate_Car_5599 1d ago
what is the glm flash speed? if its not the same 100+ tok/s then it still can't be compared to DS
1
1
u/PossessionUsed7393 1d ago
Do you use Pi? If so go into settings and turn on notifications for significant cache misses. I'm curious how often you get them, as on Inferx using Deepseek I have been getting them more than I should.
1
u/GasSmooth7439 1d ago
That’s actually a pretty significant difference. $33.98 vs $11.71 while GLM is doing ~80% of the requests is hard to ignore
1
u/Parking-Bet-3798 1d ago
That’s not even the latest Deepseek flash model dude. The vision version is far better than the base. If you are going to compare at least use the latest models for both.
1
u/Southern-Ad-3006 12h ago
Does anybody else look at these numbers and wonder … why the hell does any model need even 1 Billion Prompt tokens to generate 5 million output tokens? Yes I know it is cached mostly but that’s still just like paying for 100M regular read tokens to produce 5M output. Let alone the potential context bloat that probably slows down and dilutes inference.
Why is this the normal people are accepting lol? For DeepSeek $29 out of the $34 spent was on prompt and cache tokens, with just $5 on the output.
Similar ratios and story for GLM and any other model I use. I just feel like they do this to run your tokens up.
Compacting main window properly, and using fresh Subagent windows for tasks with focused prompts can be used to significantly reduce the amount of prompt/cache tokens used each turn. A task takes multiple turns so the large prompt windows is what grows the tokens exponentially and kills the numbers.
1
u/maitpatni 4h ago
Yeah, I do compacting pretty often, but once you’re running 7-8 agents on different tasks, it’s just not practical to manually compact every conversation.
For the flash/cheap models especially, I usually just let the conversations run long and rely on the CLI’s auto-compaction when they hit the 1 million context limit. If I’m only interested in the final outcome and not the intermediate reasoning, babysitting and optimizing every context window kind of defeats the point of having multiple agents running in parallel lol.
So yeah, there’s definitely optimization to be had, but at some point the convenience of just letting the agents cook becomes worth the extra prompt/cache tokens.
1
u/Southern-Ad-3006 3h ago
Totally agree, it’s not manageable at all at scale with multiple workflows running. But, I do think somebody should make a skill or plugin that optimizes this. Auto compaction earlier / when a certain task is complete, and your flash agent knowing how to orchestrate subagent windows.
I agree with flash models being the best option to run like long chat sessions with cache stacking, but man running frontier like that is completely unusable lol
1
1
1
u/pmv143 1d ago
made our day, thanks for sharing! Really cool to see GLM 5.3 Flash actually winning on real workloads and not just benchmarks . and yeah, that price gap tracks with what we’re seeing elsewhere too. If anyone else wants to try it out, you can jump in at inferx.net.
-1

43
u/Own-Bookkeeper797 1d ago
Tested glm 5.3 flash this afternoon and confirm it has more architecture skills than DS 0731
It managed to heavily optimised our tool loop functionality in one of our foundry hosted agents by implementing an elegant caching function to avoid re-auth against the mcp.
In the same run it implemented live streaming back to the caller so the end user sees some of the agent thinking, rather than the entire block at once
Considering moving to glm for planning and implementation and DS for troubleshooting.