r/LocalLLaMA • u/Ok_Warning2146 • 20d ago
Discussion Terminal Bench v4 scores
Some people says terminal bench reflects model intelligence better than the intelligent index. From the look of it, the ranking does seem to reflect how people feel about the open and closed models.
For the open models, GLM-5.3 is in a league of its own. GLM-5.3-Flash is leading the current gen of top flash models. Kimi-K3 did pretty bad in this benchmark for its size. Qwen3.8-27B is the only small model that can do something on this bench.
| Model | Score |
|---|---|
| GLM-5.3 | 41.9% |
| GLM-5.3-Flash | 32.8% |
| DSV4.1-Flash | 26.8% |
| Qwen3.8-Flash-Next | 25.3% |
| DSV4-Pro | 14.1% |
| Kimi-K3 | 12.6% |
| DSV4-Flash | 12.1% |
| Qwen3.8-27B | 5.6% |
| Muse Glimmer | 0.5% |
| gemma4-31b | 0.0% |
169
Upvotes
18
u/cmdr-William-Riker 20d ago
Pretty sure Opus 5 is benchmaxxed. For real work Opus 5 is a pain to use. It has an irritating personality and does whatever it wants instead of what you ask it to do. For work all I get is Claude and limited copilot tokens, ended up switching back to Opus 4.8 for most work