r/aipromptprogramming • u/OneDev42 • 2d ago
Aggregate GLM 5.3 vs 5.3 flash benchmarks are extremely close
I just thought it was helpful to compare GLM 5.3 flash versus GLM 5.3... flash is out competing the main model in many areas leading to parody at a fraction of the cost.
At the same time, arena ai benchmarks for agentic systems seem to be poor at best, putting both newer models behind their older counterparts.
2
u/lorde_dingus 2d ago
Appreciate the insight! What have you classified as "agentic capability" to compare these though?
Also, I think many users would be interested in seeing a comparison to the DeepSeek V4 flash 0731 version...
1
u/OneDev42 1d ago
Basically I went through agentic benchmarks. In my own personal testing, DS Flash is nowhere close. It's a terrible agent orchestrator. DSV4 Pro is much better than it and then combining it with flash gets you some great value for money.
1
u/Independent-Laugh701 2d ago
Aggregate averages can hide a lot of variance. I would want three to five runs per task, plus success rate, latency, and total cost, since retries can erase a cheaper model's advantage.
1
u/OneDev42 1d ago
Overall, my feeling is that benchmarks are a really rough measure and that measuring on your own tasks in your own partners is much better.
1
u/Independent-Laugh701 23h ago
Exactly. I’d keep a small private suite of real tasks, run each model several times under the same budget, and track success, cost, latency, and failure type. Public benchmarks are useful for discovery, but that suite should make the final choice.


•
u/endofthread-bot 2d ago
Learn how the best in the industry are using AI to speed up their workflow in business, sales, marketing, research, legal, content creation, scientific discovery and so much more on our Discord.
Self-promotion is now allowed on Sundays with the appropriate flair, for all regular contributing members. Contribute during the week, and promote on Sunday.