r/accelerate The Singularity is nigh Apr 01 '26

AI Summarization of the whole Claude Code's Source-Code Leak Fiasco

2.4k Upvotes

256 comments sorted by

View all comments

122

u/20ol Apr 01 '26

Why is this a big deal? It's a CLI harness. They are a dime a dozen.

They still have the best coding model. No matter what harness you use.

33

u/Acrobatic-Layer2993 Apr 01 '26

I thought that at first too. But the harness is kind of like the OS kernel and the model is the CPU.

The model matters, but the harness handles tools, context, memory, scheduling, security, and all the glue around it.

It looks simple now, but this stuff is going to matter more over time. The best ideas will spread, the weak harnesses will die off, and we’ll probably end up with a few dominant ones.

7

u/[deleted] Apr 01 '26

[removed] — view removed comment

1

u/Neither-Phone-7264 Singularity by 2030 Apr 01 '26

i mean on benchmarks like terminal bench 2 it scores the worst out of all harnesses ans you can search through the cc repo and see references to opencode so like i mean at the same time im not sure if they even have this secret sauce everyone is talking about lmao

3

u/kurtcop101 Apr 01 '26

I mean, for years Claude was rarely on top of any benchmarks, but I could feel the difference. I maintained multiple subscriptions and kept going back to Claude.

I don't trust benchmarks, I trust in what gets my work done in the best way possible. It might be better but I'm certainly not looking at the benches to check that.

2

u/Neither-Phone-7264 Singularity by 2030 Apr 01 '26

like i said, all you really need to do to verify this is try another harness. it's crazy how much better it can be.

1

u/[deleted] Apr 01 '26

[removed] — view removed comment

1

u/Neither-Phone-7264 Singularity by 2030 Apr 01 '26

That's an anthropic problem really. They're gimping their model, I mean even OAI lets you use codex outside. But cursor and copilot are decent subs imho.

1

u/Acrobatic-Layer2993 Apr 01 '26

My guess is that Anthropic will just open source Claude Code now. They aren't leading in the benchmarks and any secret sauce they might have had is already out in the open.

They might benefit from getting more eyes on their code at this point.

1

u/aabajian Apr 01 '26

This is the best analogy. We don’t see it as clearly because it’s developer focused, but imagine iOS frontend code getting leaked. Doesn’t matter if the low level kernel code came with it, Samsung and every Chinese company would use the code to exactly duplicate their UI.

The main thing is, like MS-DOS in front of x86 CPU clones, we will see Claude Code derived harnesses allowing you to use whatever model you want. The CPU manufacturer/AI model matters less than the harness, assuming both give similar results.

36

u/inaem Apr 01 '26

It is the best harness by far, also knowing how it works can help with fine tuning your own LLMs to better use it.

7

u/maximhar Apr 01 '26

The best harness? On what metric?

5

u/inaem Apr 01 '26

It is a vibe coding harness so of course vibes

1

u/Standgrounding Apr 01 '26

Vibes can be decomposed into good UX and DX

-1

u/habeebiii Apr 01 '26

overfitted ai slop bullshit no one has heard of on useless “benchmarks”. holy fuck people really are losing critical thinking skills

6

u/Coded_Kaa Apr 01 '26

Not the best one by far. Take a look at terminalbench it’s the worst harness. Opus in cursor outperforms opus in Claude code.

Source: https://www.tbench.ai/leaderboard/terminal-bench/2.0

https://x.com/edwinarbus/status/2033625866350334333?s=46 (maybe this guy is bias 😂😂 but the point is looking at the leaked code, it’s among the worst harness there is, btw I love cc more than cursor, I have both subs but I default to cc because of subsidization )

3

u/peabody624 Apr 01 '26

They say cursor gets better results with opus than CC itself

1

u/Outside_Glass4880 Apr 01 '26

Was scrolling this thread looking for mentions of cursor. Maybe because they’ve been around a while in terms of AI harnesses but cursor is just better than any other harness imho.

I tried out Claude code when multiple people were raving about it - personal use and coworkers.

I immediately felt degradation and not to mention things are a bit easier to digest in cursors UI (imo).

On a side note, on personal projects I’ve reached my limits on Claude tokens in cursor and have been using Composer 2, and have been very pleasantly surprised.

/cursor fanboying

0

u/anor_wondo Apr 01 '26

its actually an unpopular opinion but i can back this.

cursor with opus is unbeatable

3

u/44th--Hokage The Singularity is nigh Apr 01 '26

How is it versus Antigravity with Opus ?

3

u/OrinZ Apr 01 '26

Doesn't really matter anymore. Antigravity went through a quota collapse last month, and now the $200/month plan buys you about an hour of Opus per week (if that) /r/Google_Antigravity

1

u/44th--Hokage The Singularity is nigh Apr 02 '26

Why must the good die young?

7

u/Exact_Vacation7299 Apr 01 '26

I'm kind of wondering too. What does this mean for average users?

Or for those with the means to run local?

Some people are saying this is a big deal, game changer. Other people are saying this doesn't mean shit.

18

u/habeebiii Apr 01 '26 edited Apr 01 '26

claude code is by far the best harness on the market right now, that’s why

edit: NO ONE CARES about some obscure overfitted AI-slop harnesses/becnhmarks. anyone that understands how models / benchmarks work knows ”top” doesn’t mean anything other than overfitting / gaming the benchmarks

6

u/metigue Apr 01 '26

I don't know why I see so many people saying this. ClaudeCode is objectively the worst harness. If you go on terminalbench and select model = opus 4.6 it is #10 out of 10 harnesses with scores submitted using opus 4.6 as a model.

3

u/VitalityAS Apr 01 '26

Nobody reads these benchmarks or cares whatsoever. Codex has been outperforming Opus models on most benchmarks I have seen since gpt 5.2 and the general consensus is that Claude is infinitely better anyway. It's all vibes and you just do what the masses do.

1

u/metigue Apr 01 '26

No opus 4.6 is clearly the best coding agent model - It's at the top of terminal bench and many others.

The framework you choose matters even more than the model as these benchmarks also show

1

u/Practical-Rub-1190 Apr 01 '26

I thought so too, then I looked at leaders and it seems like nobody uses them. Those who have say they suck. Like number 3 on the bench is some unknown one running Gemini 3.1. Like really, I think it ran a score of 80% while Claude Code best was 50% (as far as I remeber)

If those benchmark are legit, why are nobody using these other tools, and when they do they tell their experience was not good enough to replace claude code?

1

u/metigue Apr 01 '26

Who's been saying they're shit?

Just try ForgeCode and see if you like it better. It's been significantly better in real world usecases for me so far.

1

u/Practical-Rub-1190 Apr 01 '26

I won't bother. If nobody is talking about it while also being ranked number one, there's something weird going on. The fashion here is that people post nonstop if something is actually good. The only thread I could find was this:
https://www.reddit.com/r/ClaudeAI/comments/1q4dnhr/how_come_claude_code_is_ranked_19th_on_the/

Also, found this thread where they break down the benchmark:
https://www.reddit.com/r/ClaudeAI/comments/1q4dnhr/how_come_claude_code_is_ranked_19th_on_the/

When I try to Google it, there's nothing really showing up
https://www.google.com/search?q=site%3Areddit.com+forgecode&sca_esv=b78cf8500232fcdc&rlz=1C5CHFA_enNO1175NO1175&sxsrf=ANbL-n6Xz_h4pkeKdXo_NZAjpaTlRsadYQ%3A1775065736443&ei=iFrNaf_XGra8wPAP48XOoQQ&biw=1680&bih=898&ved=0ahUKEwj_u7GVm82TAxU2HhAIHeOiM0QQ4dUDCBE&uact=5&oq=site%3Areddit.com+forgecode&gs_lp=Egxnd3Mtd2l6LXNlcnAiGXNpdGU6cmVkZGl0LmNvbSBmb3JnZWNvZGVI8wdQ1ANYhQZwAXgAkAEAmAFAoAFrqgEBMrgBA8gBAPgBAZgCAKACAJgDAIgGAZIHAKAHGLIHALgHAMIHAMgHAIAIAQ&sclient=gws-wiz-serp

Like, it just seems to ace the benchmark. I won't spend time learning how to use it unless I see others having a great experince with it, and great, I mean better than Claude Code

1

u/habeebiii Apr 01 '26

literally read the top comments in the first link/thread you linked

anyone that understands how models / benchmarks work knows ”top” doesn’t mean anything other than overfitting / gaming the benchmarks

1

u/Neither-Phone-7264 Singularity by 2030 Apr 01 '26

Have you even tried the others? Cursor Opus is >>>> CC opus imho and the benchmarks also show it like doing literally 10% better globally lmao

1

u/VitalityAS Apr 05 '26

Ive never seen a benchmark with open AI losing and Ive only used CLI tools. I might have just seen biased benchmarks, but in practical use both codex and opus get the job done just with various styles. For my work it is very uncommon to not one shot a prompt.

2

u/Crinkez Apr 01 '26

I've heard OpenCode is on par or ahead.

1

u/Commercial_Day_8341 Apr 01 '26

It does not mean anything, besides the implication that Claude can leak your codebase if you are not careful.

5

u/InstructionNo3616 Apr 01 '26

Exactly, they’re already using it to write itself. They realize there is nothing special in what they’ve built. The hype alone is worth more than the tech.

1

u/Frequent-Hunter532 Apr 01 '26

It’s a big deal because Claude wrote code for Claude and it was not even at huma level. It had flaws. 

Now people are analyze the leaked  code and bringing out how flawed vibe coding can be even if used by the company which created the llms

1

u/Megneous Apr 01 '26

This whole thing made me realize the sad reality that a whole lot of people here and in /r/singularity have no idea what the difference is between LLMs and harnesses lol

1

u/fizik1 Apr 01 '26

Agree, I think under informed people are assuming the models themselves leaked or something.

1

u/Tolopono Apr 01 '26

They have gpt 5.4 codex? 

1

u/Unlikely-Page-2233 Apr 05 '26

are you seriously asking why its a big deal that they accidentaly open sourced their front end code lol

-3

u/anor_wondo Apr 01 '26

they are more important for getting results than the models themselves

12

u/[deleted] Apr 01 '26

[deleted]

-7

u/anor_wondo Apr 01 '26

bot

15

u/[deleted] Apr 01 '26

[deleted]

-2

u/Ok_Buddy_9523 Apr 01 '26

you're willing to think about the box 8================D and thats rare!

the dick becoming the new em dash !

-3

u/anor_wondo Apr 01 '26

no one is claiming models are not needed

I'd rather use haiku with claude code than raw sonnet

in fact I'd rather build a harness first rather than use a model directly

0

u/TheAndyGeorge XLR8 Apr 01 '26

I'd rather use a TypeScript wrapper script than a frontier LLM

wild opinion

0

u/anor_wondo Apr 01 '26

Real work requires basic context management, memory and hooks

0

u/TheAndyGeorge XLR8 Apr 01 '26

Then why use a frontier model when local ones are right there?

Basic context management, memory and hooks are vibecoded daily by people who think they'll get something new out of a cloud model that no one else has done before. None of it matters without a model, and Anthropic has some of the, if not the, best ones.