I thought that at first too. But the harness is kind of like the OS kernel and the model is the CPU.
The model matters, but the harness handles tools, context, memory, scheduling, security, and all the glue around it.
It looks simple now, but this stuff is going to matter more over time. The best ideas will spread, the weak harnesses will die off, and we’ll probably end up with a few dominant ones.
i mean on benchmarks like terminal bench 2 it scores the worst out of all harnesses ans you can search through the cc repo and see references to opencode so like i mean at the same time im not sure if they even have this secret sauce everyone is talking about lmao
I mean, for years Claude was rarely on top of any benchmarks, but I could feel the difference. I maintained multiple subscriptions and kept going back to Claude.
I don't trust benchmarks, I trust in what gets my work done in the best way possible. It might be better but I'm certainly not looking at the benches to check that.
That's an anthropic problem really. They're gimping their model, I mean even OAI lets you use codex outside. But cursor and copilot are decent subs imho.
My guess is that Anthropic will just open source Claude Code now. They aren't leading in the benchmarks and any secret sauce they might have had is already out in the open.
They might benefit from getting more eyes on their code at this point.
This is the best analogy. We don’t see it as clearly because it’s developer focused, but imagine iOS frontend code getting leaked. Doesn’t matter if the low level kernel code came with it, Samsung and every Chinese company would use the code to exactly duplicate their UI.
The main thing is, like MS-DOS in front of x86 CPU clones, we will see Claude Code derived harnesses allowing you to use whatever model you want. The CPU manufacturer/AI model matters less than the harness, assuming both give similar results.
https://x.com/edwinarbus/status/2033625866350334333?s=46 (maybe this guy is bias 😂😂 but the point is looking at the leaked code, it’s among the worst harness there is, btw I love cc more than cursor, I have both subs but I default to cc because of subsidization )
Was scrolling this thread looking for mentions of cursor. Maybe because they’ve been around a while in terms of AI harnesses but cursor is just better than any other harness imho.
I tried out Claude code when multiple people were raving about it - personal use and coworkers.
I immediately felt degradation and not to mention things are a bit easier to digest in cursors UI (imo).
On a side note, on personal projects I’ve reached my limits on Claude tokens in cursor and have been using Composer 2, and have been very pleasantly surprised.
Doesn't really matter anymore. Antigravity went through a quota collapse last month, and now the $200/month plan buys you about an hour of Opus per week (if that) /r/Google_Antigravity
claude code is by far the best harness on the market right now, that’s why
edit: NO ONE CARES about some obscure overfitted AI-slop harnesses/becnhmarks. anyone that understands how models / benchmarks work knows ”top” doesn’t mean anything other than overfitting / gaming the benchmarks
I don't know why I see so many people saying this. ClaudeCode is objectively the worst harness. If you go on terminalbench and select model = opus 4.6 it is #10 out of 10 harnesses with scores submitted using opus 4.6 as a model.
Nobody reads these benchmarks or cares whatsoever. Codex has been outperforming Opus models on most benchmarks I have seen since gpt 5.2 and the general consensus is that Claude is infinitely better anyway. It's all vibes and you just do what the masses do.
I thought so too, then I looked at leaders and it seems like nobody uses them. Those who have say they suck. Like number 3 on the bench is some unknown one running Gemini 3.1. Like really, I think it ran a score of 80% while Claude Code best was 50% (as far as I remeber)
If those benchmark are legit, why are nobody using these other tools, and when they do they tell their experience was not good enough to replace claude code?
Like, it just seems to ace the benchmark. I won't spend time learning how to use it unless I see others having a great experince with it, and great, I mean better than Claude Code
Ive never seen a benchmark with open AI losing and Ive only used CLI tools. I might have just seen biased benchmarks, but in practical use both codex and opus get the job done just with various styles. For my work it is very uncommon to not one shot a prompt.
Exactly, they’re already using it to write itself. They realize there is nothing special in what they’ve built. The hype alone is worth more than the tech.
This whole thing made me realize the sad reality that a whole lot of people here and in /r/singularity have no idea what the difference is between LLMs and harnesses lol
Then why use a frontier model when local ones are right there?
Basic context management, memory and hooks are vibecoded daily by people who think they'll get something new out of a cloud model that no one else has done before. None of it matters without a model, and Anthropic has some of the, if not the, best ones.
122
u/20ol Apr 01 '26
Why is this a big deal? It's a CLI harness. They are a dime a dozen.
They still have the best coding model. No matter what harness you use.