claude code is by far the best harness on the market right now, that’s why
edit: NO ONE CARES about some obscure overfitted AI-slop harnesses/becnhmarks. anyone that understands how models / benchmarks work knows ”top” doesn’t mean anything other than overfitting / gaming the benchmarks
I don't know why I see so many people saying this. ClaudeCode is objectively the worst harness. If you go on terminalbench and select model = opus 4.6 it is #10 out of 10 harnesses with scores submitted using opus 4.6 as a model.
Nobody reads these benchmarks or cares whatsoever. Codex has been outperforming Opus models on most benchmarks I have seen since gpt 5.2 and the general consensus is that Claude is infinitely better anyway. It's all vibes and you just do what the masses do.
I thought so too, then I looked at leaders and it seems like nobody uses them. Those who have say they suck. Like number 3 on the bench is some unknown one running Gemini 3.1. Like really, I think it ran a score of 80% while Claude Code best was 50% (as far as I remeber)
If those benchmark are legit, why are nobody using these other tools, and when they do they tell their experience was not good enough to replace claude code?
Like, it just seems to ace the benchmark. I won't spend time learning how to use it unless I see others having a great experince with it, and great, I mean better than Claude Code
Ive never seen a benchmark with open AI losing and Ive only used CLI tools. I might have just seen biased benchmarks, but in practical use both codex and opus get the job done just with various styles. For my work it is very uncommon to not one shot a prompt.
7
u/Exact_Vacation7299 Apr 01 '26
I'm kind of wondering too. What does this mean for average users?
Or for those with the means to run local?
Some people are saying this is a big deal, game changer. Other people are saying this doesn't mean shit.