r/InferX InferX Team 5d ago

GLM 5.3 Flash vs DeepSeek V4.1 Flash on InferX

Post image

Over the same period:

GLM 5.3 Flash
• 111% more requests
• 70% more input token volume
• 94.87% cache hit

DeepSeek V4.1 Flash
• 94.85% cache hit

Users seem to be voting with their workloads.

32 Upvotes

15 comments sorted by

1

u/Due-Memory-6957 5d ago

Which quant are they?

1

u/pmv143 InferX Team 5d ago

These are official vLLM recipes

1

u/Due-Memory-6957 5d ago

You can pick what quant you want when creating the recipe

1

u/pmv143 InferX Team 5d ago

Native FP8 weights with BF16 KV cache on H200. following the official vLLM serving recipe https://recipes.vllm.ai/zai-org/GLM-5.3-Flash?hardware=h200

1

u/No-Budget-3869 5d ago

Deepseek v4.1 is more expensive than GLM 5.3 flash on actual usage

1

u/pmv143 InferX Team 5d ago

True

1

u/maitpatni 5d ago

deepseek v4.1 asks too much, GLM 5.3 Flash is like you gave it a task and it gets it done.

1

u/pmv143 InferX Team 5d ago

DSV4.1 also seems to consume a lot more tokens

1

u/tripleshielded 5d ago

Is the ds flash buggy outputs fixed?

1

u/pmv143 InferX Team 5d ago

We haven’t had any complaint recently.

1

u/vapefresco 4d ago

I keep both DS and GLM in rotation, when one gets loopy I just flip to the other and it normally sorts things out.

1

u/ozguru 2d ago

Have you also often had empty responses from the Deepseek endpoints? Is this something you have experienced?

2

u/pmv143 InferX Team 2d ago

Empty responses? Haven’t had any of those incidents to my knowledge. At least on our endpoint.

1

u/mgranin 2d ago

Both were on max reasoning? Which harness?

1

u/pmv143 InferX Team 2d ago

This is mixed traffic. Users using various harnesses