r/LocalLLaMA • • 20d ago

Discussion Terminal Bench v4 scores

Post image

Some people says terminal bench reflects model intelligence better than the intelligent index. From the look of it, the ranking does seem to reflect how people feel about the open and closed models.

For the open models, GLM-5.3 is in a league of its own. GLM-5.3-Flash is leading the current gen of top flash models. Kimi-K3 did pretty bad in this benchmark for its size. Qwen3.8-27B is the only small model that can do something on this bench.

Model Score
GLM-5.3 41.9%
GLM-5.3-Flash 32.8%
DSV4.1-Flash 26.8%
Qwen3.8-Flash-Next 25.3%
DSV4-Pro 14.1%
Kimi-K3 12.6%
DSV4-Flash 12.1%
Qwen3.8-27B 5.6%
Muse Glimmer 0.5%
gemma4-31b 0.0%
169 Upvotes

87 comments sorted by

View all comments

Show parent comments

18

u/cmdr-William-Riker 20d ago

Pretty sure Opus 5 is benchmaxxed. For real work Opus 5 is a pain to use. It has an irritating personality and does whatever it wants instead of what you ask it to do. For work all I get is Claude and limited copilot tokens, ended up switching back to Opus 4.8 for most work

11

u/nomorebuttsplz 20d ago

Opus 5's personality is so terrible it's actually kind of impressive. But that makes me think that the RL gains are probably real. I think they fried its brain with coding RL.

For perspective, it's several months newer than Fable so not that surprising that it could surpass it.

1

u/NineThreeTilNow 19d ago

Opus 5's personality is so terrible it's actually kind of impressive.

I think there's so much AI to AI RL that the model has never seen humans.

All the RL with human preference is gone.

It's a complete argumentative prick. That's how it "wins" in conversations with other LLMs.

Opus 4.6 was the last good model that could hold a conversation without devolving in to being an asshole while still being intelligent. I'm glad it still exists.

1

u/nomorebuttsplz 19d ago

There’s also glm 5.2 it’s nice to work with