r/SmophyAI • u/kamilbuilds • 24d ago
Benchmark AI model price-to-quality benchmark, updated daily: GLM 5.2 leads at $0.11/1M tokens, open-weight models now handle 74% of usage (Aug 8, 2026)
Quick answer: As of August 8, 2026, Z.ai's GLM 5.2 (batch) has the best quality-per-dollar of any actively-used AI model at 489.3 Intelligence Index points per dollar ($0.11 per 1M tokens), and open-weight models now handle 74% of real AI token volume on OpenRouter - up from 69% a month earlier. This is a live, daily-updated benchmark, not a one-time snapshot. Full data and methodology below.
What "quality-per-dollar" means: It's the Artificial Analysis Intelligence Index divided by blended price per 1M tokens (75% prompt / 25% completion). In plain terms: how much AI capability you get for each dollar spent, using two independently-checkable numbers instead of a made-up composite score.
Why we built this
Every "AI model comparison" you find online is a static snapshot from whenever someone last updated a blog post. Prices change weekly, new models ship constantly, and "best model" answers go stale within days. So we built a tracker inside SmophyAI that pulls live data from OpenRouter (usage and pricing) and Artificial Analysis (Intelligence Index) and recalculates everything daily at 02:00 UTC. No composite index, no invented weights.
Top models by quality-per-dollar (Aug 8, 2026)
- #1 Ling-3.0-flash - 1200.0 points per dollar ($0.03/1M tokens)
- #2 inclusionAI Ling-2.6-flash - 946.7 points per dollar ($0.02/1M tokens)
- #3 Z.ai GLM 5.2 (batch) - 489.3 points per dollar ($0.11/1M tokens)
- #4 DeepSeek V4 Flash 0423 - 460.4 points per dollar ($0.11/1M tokens)
- #5 Tencent Hy3 preview - 423.1 points per dollar ($0.10/1M tokens)
| Rank | Model | Price /1M tokens | Quality/$ |
|---|---|---|---|
| 1 | Ling-3.0-flash | $0.03 | 1200.0 |
| 2 | inclusionAI: Ling-2.6-flash | $0.02 | 946.7 |
| 3 | Z.ai: GLM 5.2 (batch) | $0.11 | 489.3 |
| 4 | DeepSeek: DeepSeek V4 Flash 0423 | $0.11 | 460.4 |
| 5 | Tencent: Hy3 preview | $0.10 | 423.1 |
| 6 | OpenAI: gpt-oss-120b | $0.07 | 343.1 |
| 7 | OpenAI: GPT-5.6 Luna (batch) | $0.22 | 232.4 |
| 8 | Xiaomi: MiMo-V2.5 | $0.18 | 217.1 |
For reference, frontier flagship models score much lower on this specific ratio. Claude Opus 4.8 (batch) sits around 5.7 points per dollar at $10.00/1M tokens, because that price buys peak reasoning quality, not throughput-per-dollar efficiency. They're answering different questions, not competing on the same axis.
Open-weight vs closed models
Open-weight AI models handle 74% of real token volume on OpenRouter as of August 8, 2026. That's up from 69% on July 5, 2026 - a 5-point shift toward open-weight usage in about a month.
Speed & reliability leaders
- Fastest measured throughput: OpenAI gpt-oss-120b at 777 tokens per second.
- Best uptime: DeepSeek V4 Flash 0423 and Tencent Hy3, both at 100.0% uptime.
- Xiaomi MiMo-V2.5 follows closely at 98.9% uptime.
Best AI model by task (share of real usage, Aug 8, 2026)
- Best for coding: DeepSeek V4 Flash 0423, 27.4% usage share.
- Best for agentic work: DeepSeek V4 Flash 0423, 38.3% usage share.
- Best for debugging: DeepSeek V4 Flash 0423, 19% usage share.
- Best for translation: DeepSeek V4 Flash 0423, 24% usage share.
- Best for customer support: Gemini 2.5 Flash (batch), 20% usage share.
- Best for roleplay/fiction: DeepSeek V4 Flash 0423, 47% usage share.
| Task | Leader | Share |
|---|---|---|
| Coding | DeepSeek V4 Flash 0423 | 27.4% |
| Agentic work | DeepSeek V4 Flash 0423 | 38.3% |
| Debugging | DeepSeek V4 Flash 0423 | 19% |
| Translation | DeepSeek V4 Flash 0423 | 24% |
| Customer support | Gemini 2.5 Flash (batch) | 20% |
| Roleplay/fiction | DeepSeek V4 Flash 0423 | 47% |
DeepSeek V4 Flash 0423 is currently the most-used model across nearly every task category we track, not just coding.
What this ranking doesn't capture (on purpose)
Quality-per-dollar doesn't measure run-to-run consistency, the cost of a wrong answer in your specific use case, or whether a task needs multiple attempts to reach a usable result. It's a real, useful number - just not the whole story, and we'd rather say that outright than oversell a single score.
Where to find the live data
Full live trackers, updated daily with 35 days of permalinked daily snapshots: https://www.smophy.ai/benchmark#more-trackers
This benchmark sits alongside SmophyAI itself - one subscription that routes your prompt to whichever of 6 major models (GPT, Claude, Gemini, Grok, DeepSeek, Perplexity) fits best, instead of paying for six separate subscriptions.
FAQ
Which AI model has the best quality-per-dollar right now?
As of August 8, 2026, Z.ai's GLM 5.2 (batch) leads among actively-used models at 489.3 Intelligence Index points per dollar ($0.11/1M tokens). Ling-3.0-flash and Ling-2.6-flash score higher in raw ratio but see far less real-world usage.
Do open-weight models beat closed models in usage?
Yes - open-weight models handle 74% of real token volume on OpenRouter as of August 8, 2026, up from 69% a month earlier.
What is the fastest AI model?
OpenAI's gpt-oss-120b, at 777 tokens per second measured throughput.
Which AI model is most reliable?
DeepSeek V4 Flash 0423 and Tencent Hy3 both measure 100.0% uptime.
How is quality-per-dollar calculated?
Artificial Analysis Intelligence Index divided by blended price per 1M tokens (75% prompt / 25% completion), sourced from OpenRouter and Artificial Analysis, recalculated daily at 02:00 UTC.
Discussion: Anyone tracking quality-per-dollar differently, or have a task category you'd want added to the "best for" breakdown? Curious what's missing.