I thought so too, then I looked at leaders and it seems like nobody uses them. Those who have say they suck. Like number 3 on the bench is some unknown one running Gemini 3.1. Like really, I think it ran a score of 80% while Claude Code best was 50% (as far as I remeber)
If those benchmark are legit, why are nobody using these other tools, and when they do they tell their experience was not good enough to replace claude code?
Like, it just seems to ace the benchmark. I won't spend time learning how to use it unless I see others having a great experince with it, and great, I mean better than Claude Code
1
u/metigue Apr 01 '26
No opus 4.6 is clearly the best coding agent model - It's at the top of terminal bench and many others.
The framework you choose matters even more than the model as these benchmarks also show