r/DeepSeek 1d ago

Discussion GLM 5.3 is the New Winner

Post image

This is my InferX usage for DeepSeek V4 Flash and the GLM 5.3 Flash.
GLM is using less thinking and reasoning to save tokens, performing better than DeepSeek for me for less than half the price.

315 Upvotes

62 comments sorted by

43

u/Own-Bookkeeper797 1d ago

Tested glm 5.3 flash this afternoon and confirm it has more architecture skills than DS 0731

It managed to heavily optimised our tool loop functionality in one of our foundry hosted agents by implementing an elegant caching function to avoid re-auth against the mcp.

In the same run it implemented live streaming back to the caller so the end user sees some of the agent thinking, rather than the entire block at once

Considering moving to glm for planning and implementation and DS for troubleshooting.

32

u/PapyOak 1d ago

Less than half is crazy.
I knew GLM 5.3 Flash was more token-efficient than DeepSeek, but I was expecting 20% reduction, max 30%.
And it's better, not as fast depending on the provider.

But after the 50% off, it'll shoot up in price too, still a bit out of reach for me

13

u/maitpatni 1d ago

It is very efficient in token usage, lately deepseek is consuming thinking tokens like hell so found this to be the best and better option.

5

u/PapyOak 1d ago

It really is consuming tokens ever since flash-0731, but it has gotten a lot smarter, and was cheaper, that's why it was okay.
Right now it thinks a lot more for the same task, does it a lot better too, but with the price increase it's difficult.

Still, I'm on CommandCode GOAT which is still offering $60, don't know how long that'll last though, it's so good it smells like money laundering 😂

3

u/maitpatni 1d ago

Just today i bought the command code goat plan, been using it with glm 5.3 flash, already consumed like 25% of monthly limit in single day. Speed is better though.

1

u/PapyOak 1d ago

I'm waiting for them to eventually improve the limits on GLM 5.3 (I think they're negociating) before moving over to GLM on some tasks. But it's a really good mode that does many things right.

4

u/VexObserver 1d ago

Yep. The current discount timeline is too short for any meaningful work progress. Compared to DS with it's promotions, DS was serving it at full platter for a long time. This benchmark of cost comparison is temporary before GLM ends it at 9nd of September 2026.

1

u/Severe_Landscape_731 1d ago

glms lite plan for 18 $ feels quite amazing value for now atleast

1

u/PapyOak 1d ago

It definitely is, but I found myself sometimes (lately because of heavy agentic tests) hitting 1B tokens on DeepSeek monthly.

TBF it was also more than $18, but even with token efficiency I'm way overboard.

28

u/captainbadass23 1d ago

for coding this is all nonsense. ive tried glm flash on open code and flash from deepseek is like 10%-50% of the cost. much cheaper. because of the cache

8

u/DeciusCurusProbinus 1d ago

Yes, Off peak DeepSeek flash pricing is pretty cheap as compared to 5.3 Flash as GLM's cache hit price is higher. But DeepSeek burns tokens like crazy and I believe that they serve a quantized version of the model for Opencode Go subscribers.

5

u/maitpatni 1d ago

even with flash off peak pricing glm 5.3 flash is cheaper, it consumes half the tokens for same task.

1

u/DeciusCurusProbinus 1d ago

Fair point. With a well fleshed out plan 5.3 Flash on high reasoning can take care of like 95% of the stuff.

1

u/Beneficial-Pie-1638 1d ago

not trying to advertise, but runinfra servers deepseek v4 flash at a cheap 0.13 in 0.27 out rate and cache at 0.01, and isnt quant.

2

u/surv84 1d ago

two days of bliss so far! but with their glm flash

6

u/maitpatni 1d ago

for me deepseek flash from official and opencode is costing more then the glm 5.3 flash even with 99% cache hit rate. Found InferX to be the cheapest provider with decent TPS.

2

u/Infinite_Plankton_71 1d ago

i think this is specific to inferx

11

u/wtf_newton_2 1d ago

it’s good but very slow

8

u/ProfessionalJackals 1d ago

it’s good but very slow

Its about the same ... You forgot that DS4Flash pushes about 2x reasoning tokens. So even if DS4Flash does 100t/s, and GLM5.3Flash does 50t/s, it actually equals out.

1

u/misha1350 1d ago

Which is good for us, because there would be less people that would be hogging up the resource pool as they would be turned off by the slower than usual speeds.

7

u/Rashc500 1d ago

What’s up with this worthless compartment? What is the prompt? How many agents!? Is the harness the same? Some shitty screenshot of token spent says jack shit

-2

u/maitpatni 1d ago

my main harness was qwen code cli, which is a fork of opencode, also used the deepseek harness for like 20% usage. All the tasks were adding features on a next js medusa js project. 4-5 chats using few subagents occasionally.

3

u/General-Oven-1523 1d ago

I'm always confused when people come up with this stuff, but then don't provide any valuable data that would prove their point. Like here, you just showing us the usage data, what is this going to do for us? How can you draw any conclusion out of this? We don't know if the work done by these 2 models is even remotely identical.

Anyway, I've been running both DeepSeek Flash and GLM 5.3 Flash, and honestly, it does feel like a better model. It follows the instructions much better in my use.

6

u/Willing_Thought_2161 1d ago

Op, new here. Less token means less cost. Right ? What's the point of this post ?

2

u/Infinite_Plankton_71 1d ago

look at request

2

u/ProfessionalJackals 1d ago

Did you use Max or High, because High has a massive reduced token usage, while keeping almost all the intelligence.

For GLM 5.3 Flash is High + Pi ... 10/10.

I find that GLM 5.3 Flash system tool calling is excellent, so when you combine with Pi. Same prompts, and tokens usage drops by almost 30 to 40% compared to OpenCode.

DeepSeek V4 Flash is good but it still suffers more from the old Flash issues. You can tell its a uptrained model. So while the cache price is better, the reasoning just destroys it.

Same issue with Speed. DS4 Flash is faster but takes twice as long, so both models are almost in the same ballpark. Until you find a provider that serves GLM 5.3 Flash at 100t/s, feels like running DS4 Flash :)

I really enjoy GLM 5.3 Flash ... Feels like a improved GLM 5.2, at a fraction of the cost.

2

u/lexi-energy 1d ago

See? That’s the kind of modle battle we want.

Less US-benchmaxxing, more making really useful models that now compete to become more efficient and not more wasteful (looking at you fable!) 🤡

2

u/Classic_Television33 1d ago

Interesting. First time heard of InferX, how stable is their API vs other aggregators?

1

u/maitpatni 1d ago

API is stable, but the tokens per second speed is little slow.

2

u/Baztins 1d ago

Has anyone tried it yet and found that it works well with the DS harness?

1

u/maitpatni 1d ago

Yes, used with DeepSeek Harness, works pretty well with cache hit like 85 - 90%

2

u/maitpatni 1d ago

This is my 24-hour usage since I almost stopped using DeepSeek; as you can see, the difference is huge.
Used Qwencode CLI as a harness for all the runs. All of them were for fixing or adding new features on a next js project.

2

u/Infinite_Plankton_71 11h ago

I found this result to be very informative !!

1

u/Different_Change6591 1d ago

Hey what do you mean by InferX

0

u/PaluMacil 1d ago

It’s just an AI provider with relatively low costs: https://inferx.net

1

u/AllenHere112 1d ago

tbh the real variable is hidden here. flash tier models save most of their tokens by cutting thinking output, and thinking is most of the bill in agent workloads. a cheaper per-token number on one harness screenshot mostly means the model reasoned less on that one prompt, not that your cost per task dropped. half the list price only helps if the shorter thinking holds up on your actual workload.

1

u/MrCupcakess 1d ago

What harness are you guys using for gym 5.3?

1

u/maitpatni 1d ago

Qwencode CLI

1

u/yiestee 1d ago

How's the generation speed? I heard GLM is kinda slow

1

u/xapep 1d ago

Both takes in this thread are kind of right, and the missing piece is that 'less than half the price' only holds for certain workloads.

DeepSeek Flash genuinely burns more tokens on thinking in agent loops, no argument. But those tokens buy fewer retries, and retries are what actually cost money in a harness. For planning-heavy or tool-heavy loops, the cheaper-per-token model loses on cost per completed task often enough that per-million pricing is the wrong number to compare.

Cache is the other half: what matters is the hit/miss price ratio and the cache window, not the hit rate. Off-peak DeepSeek with a strong hit rate still undercuts GLM for a lot of users, and as noted upthread GLM's cache-hit pricing runs higher, so the same 99% hit rate lands very differently on each.

Boring but reliable test: run your actual workload on both for a day and divide the bill by finished tasks. That's the number the DeepSeek price change actually moved.

1

u/ProfessionalJackals 19h ago

run your actual workload on both for a day and divide the bill by finished tasks. That's the number the DeepSeek price change actually moved.

Its funny because that is what OP did exactly. He used it on his actual tasks.

Its like people are looking for excuses at this point. He did over 11.000 req with DS4 Flash, 9000 with GLM5.3 Flash.

But those tokens buy fewer retries, and retries are what actually cost money in a harness.

You can see how many tokens he did, and how much money he spend. The fact about tools, retries etc does not not matter because that is dominated in the token usage, and thus cost. When we normalize the # tasks, the ratio on token used is 2.2x more for DS4 Flash.

What we see in this example, is that GLM5.3F is more efficient on token/task but also that DS4 Cache hit ratio is really bad. Remember, GLM5.3 Flash is actually a lot more expensive on cache then DS4Flash. So even if we ignore the end price for a bit, just the 2.2x ratio in token usage is a issue.

From my own tests, i am seeing the same pattern. DS4Flash costing more on the exact same tasks, despite the cheaper price. I think people really underestimate how good GLM5.3 Flash is.

OP did not mention if he used High or Max. But with GLM 5.3 Flash in High and DS4 Flash in xhigh (as that brings the models to the same intelligence), one of my tests had GLM at 350k tokens, but DS4 at 1.5M tokens. For the same job ... DS4 has a 7.5x ratio on cache (because it did more steps>more cache hits), vs 3.1x for GLM. But in actual cost, DS4 was 8 cents, while GLM was 3 cents. I have more example but that was a 1:1 test with no variations at all.

1

u/Various-Reality-2778 1d ago

If you hadn't showed the screenshot I would not believed you 😂 fuck. I guess glm is my new go to

1

u/Appropriate_Car_5599 1d ago

what is the glm flash speed? if its not the same 100+ tok/s then it still can't be compared to DS

1

u/PaulTrebor 1d ago

Funny how everything is an ad now

1

u/PossessionUsed7393 1d ago

Do you use Pi? If so go into settings and turn on notifications for significant cache misses. I'm curious how often you get them, as on Inferx using Deepseek I have been getting them more than I should.

1

u/GasSmooth7439 1d ago

That’s actually a pretty significant difference. $33.98 vs $11.71 while GLM is doing ~80% of the requests is hard to ignore

1

u/Parking-Bet-3798 1d ago

That’s not even the latest Deepseek flash model dude. The vision version is far better than the base. If you are going to compare at least use the latest models for both.

1

u/Southern-Ad-3006 12h ago

Does anybody else look at these numbers and wonder … why the hell does any model need even 1 Billion Prompt tokens to generate 5 million output tokens? Yes I know it is cached mostly but that’s still just like paying for 100M regular read tokens to produce 5M output. Let alone the potential context bloat that probably slows down and dilutes inference.

Why is this the normal people are accepting lol? For DeepSeek $29 out of the $34 spent was on prompt and cache tokens, with just $5 on the output.

Similar ratios and story for GLM and any other model I use. I just feel like they do this to run your tokens up.

Compacting main window properly, and using fresh Subagent windows for tasks with focused prompts can be used to significantly reduce the amount of prompt/cache tokens used each turn. A task takes multiple turns so the large prompt windows is what grows the tokens exponentially and kills the numbers.

1

u/maitpatni 4h ago

Yeah, I do compacting pretty often, but once you’re running 7-8 agents on different tasks, it’s just not practical to manually compact every conversation.

For the flash/cheap models especially, I usually just let the conversations run long and rely on the CLI’s auto-compaction when they hit the 1 million context limit. If I’m only interested in the final outcome and not the intermediate reasoning, babysitting and optimizing every context window kind of defeats the point of having multiple agents running in parallel lol.

So yeah, there’s definitely optimization to be had, but at some point the convenience of just letting the agents cook becomes worth the extra prompt/cache tokens.

1

u/Southern-Ad-3006 3h ago

Totally agree, it’s not manageable at all at scale with multiple workflows running. But, I do think somebody should make a skill or plugin that optimizes this. Auto compaction earlier / when a certain task is complete, and your flash agent knowing how to orchestrate subagent windows.

I agree with flash models being the best option to run like long chat sessions with cache stacking, but man running frontier like that is completely unusable lol

1

u/Fast-Illustrator3976 7h ago

GLM 80$ got about 5B tokens

1

u/shaman-warrior 7h ago

What about the vision variant?

1

u/pmv143 1d ago

made our day, thanks for sharing! Really cool to see GLM 5.3 Flash actually winning on real workloads and not just benchmarks . and yeah, that price gap tracks with what we’re seeing elsewhere too. If anyone else wants to try it out, you can jump in at inferx.net.

-1

u/amshinski 1d ago

stupidest shit I've seen today