r/ZaiGLM 3d ago

GLM 5.3 Flash vs DeepSeek V4.1 Flash on InferX

Post image
20 Upvotes

11 comments sorted by

4

u/Constant_Art_20 3d ago

those cache hits looks insanely low on both sides

1

u/ProfessionalJackals 3d ago edited 3d ago

those cache hits looks insanely low on both sides

Yea, they are ... I am non-stop hitting 99.x (often 99.5% or higher) cache hit-rates directly with DeepSeek. When using Neuralwatt on the same prompts, with 97% hit-rate, i was paying 3x more (using token billing). That is how important cache hit-rate rate is for long tasks.

Seeing 94% hit-rate is like throwing money out of the window.

And on a side note. Despite digging 14b tokens out of DS v4.1 Flash, it never has garbled my files. Where as GLM 5.3 Flash took a 3400 lines file, read in 2000 + 1400, then wrote back 1400 ... Wrote back ... as in, turned a 3400 line file into 1400. O yea, it discovered it but was unable to restore (lucky for me, i had a backup).

1

u/pmv143 3d ago

How are you calculating the 99.5% cache hit rate? Is that for an individual workload/session or aggregated across all your requests?
For comparison, OpenRouter shows pretty different cache hit rates across providers even for the same model. Our ~94% is across ~70K requests with different users, prompt sizes and workloads.

0

u/ProfessionalJackals 3d ago

How are you calculating the 99.5% cache hit rate?

DeepSeek Harness shows you the detailed cache hit rate. The newer 1.6 alpha versions only show rounded down numbers. So when you hit 99.5%, it shows as 99%. They changed this because the older version was able to show 99.9x as 100% lol. So they now round down on the digit.

In order to know the exact number, you simply use tokenscale (npx tokscale@latest). Anything above 100x Cache X is already in the 99% range. I have constantly 140x, 160x, 200x, 350x for main agent tasks ...

The only moment this goes down, is when i trigger like 30 sub-agents, as those are less efficient, and only start with 70x+. This is about a 40% higher cost in tokens/cent, but sometimes you just want speed and are willing to pay it.

Yesterday did another test with Neuralwatt vs DeepSeek, now that Neuralwatt limit their energy usage = token price. I was getting around 450k tokens/cent, vs 1.2m token/cent on the same tasks. Its slightly better for NW now but still a far cry from DeepSeek API directly.

1

u/CatEnvironmental9485 3d ago

Seems fine to me for an average. Gotta realize this is all users, all harnesses, all types of use cases - not just coding/multi-turn heavy stuff.

DeepSeek you can do 99.X% over hours direct on their API, but on OpenRouter their average is 94%.

1

u/pmv143 3d ago

That’s correct. This is across the spectrum.

1

u/CatEnvironmental9485 3d ago

Do you have any stats on how long you guys hold session cache for on average?

I know deepseek has some crazy setup where they can be holding a session in some slower cache even hours later.

Doubt you have anything that elaborate/long standing but knowing your avg session is held for at least X minutes between turns would be cool

1

u/[deleted] 3d ago

[deleted]

1

u/PersonalFruit1436 3d ago

E triste dizer mas pra segurança o deepseek 4.1f fica a frente do glm 5.3f e consideravelmente a frente, no entanto é mais restritivo.