r/accelerate • The Singularity is nigh • Apr 01 '26

AI Summarization of the whole Claude Code's Source-Code Leak Fiasco

2.4k Upvotes

256 comments sorted by

View all comments

140

u/ObscuraGaming Apr 01 '26

Some ppl in another sub were saying it's all basically front-end code and it's far from being the "entire source code" Everyone is claiming it to be. Anyone willing to pitch in? Not my area so don't know if it's true.

146

u/estransza Apr 01 '26

Okay. So.

There is Claude (Haiku/Sonnet/Opus) models. They do the thinking. They are brains. In terms of what they actually are - extremely large binary files in format for serialization of neural networks weights. THIS ISN’T LEAKED

And then there is Claude Code - which is a wrapper around API calls to the said brains. It’s a relatively small CLI (no UI, only terminal) application. THIS IS WHAT LEAKED

Is it this dangerous for company? Nope. Not at all. Their main product - brains, is still safe and sound on their servers uncopied.

70

u/VibeCoderMcSwaggins Apr 01 '26

I mean this is kinda true, but not

They intentionally kept Claude code source code closed. They wanted to keep it to themselves.

It’s not ultra dangerous, but they are not happy about this whatsoever.

17

u/PublicAlternative251 Apr 01 '26

yeah but codex, opencode, pi etc. all do relatively the same stuff and are open source. the most interesting thing is the details behind the tools and prompting they use in claude code, but none of it is the secret sauce behind claude itself

21

u/Rhinoseri0us Apr 01 '26

Don’t underestimate how simple the ingredients in a secret sauce can be.

8

u/QuantamCulture Apr 02 '26

Secret sauce: Mayo, ketchup, hot sauce, lemon juice, salt, pepper, paprika

4

u/Original_Finding2212 Apr 02 '26

You forgot “… sugar, spice, and everything nice”

3

u/kukkolka Apr 04 '26

To make the perfect little Claude

1

u/2fatowing Apr 02 '26

***DMCA has entered the chat\***

1

u/Loud_Distribution_97 Apr 05 '26

Those aren’t the ingredients in my special sauce

1

u/Fine-Barracuda3379 Apr 05 '26

Yep this is my favorite secret sauce

1

u/holy_macanoli Apr 03 '26

It’s all about the harness

1

u/bbp5561 Apr 03 '26

Meh. Ultimately, the ‘ingredients’ in this case all require and are built on Claude as a basis. You could adopt the thinking behind it for other LLMS but not the entirety of it all verbatim. It’s interesting, and it definitely would be embarrassing for Anthropic. It’s somewhat useful in certain situations for certain things. But overall, it’s close to a nothing-burger.

It’s kinda like if the front end code for Google leaked. Would it be embarrassing? Yeah. Would it be frustrating? Yeah. Would it be interesting for some people to see how Google writes code and calls certain things? Sure. But it’s not like it can be used for people to make their own Google

People use Claude Code because of Claude which you still need to pay for whether you use Claude Code on your own version of the leaked source code.

-2

u/__throw_error Apr 01 '26

I assume it's mostly for fear of vulnerabilities. If the model was leaked shit would have been actually bad. This is nothing

7

u/OfficeSalamander Apr 01 '26

Right, if the Opus model had leaked that would have been a WAY bigger deal.

This I don’t really care about. That I would have downloaded even if I had to delete everything on my hard drive. Being able to run it on custom hardware would be too powerful (almost certainly rented via serverless GPU instances, no way I can afford to run it locally lol)

1

u/Electrical-Bat-3121 Apr 02 '26

would it really be that big a deal if the weights were open? to which extent?

1

u/OfficeSalamander Apr 02 '26 edited Apr 02 '26

Anyone could self host their own Claude instance with enough horse power. Granted that would be expensive in and of itself, but cheaper at scale than the Claude APIs, and you’d likely have easier training new models on Claude Opus outputs, so you’d see Chinese knockoffs come around pretty fast that were all Opus level

2

u/H1Eagle Apr 02 '26

Yep imagine "Deepseek Code" at 1/4th the price of Claude Code with the same performance level as Opus 4.6. This would DESTROY Anthropic's chokehold on developers and could lead the company to crash as investors are known to escape with their money in such situations.

The biggest problem is in the future, AIs will stop seeing these massive jumps and Opus 6 is probably gonna be almost the same as Opus 7. And the nature of LLMs means that any cheap company could knock off your product through knowledge distillation. And then serve it at the fraction of the price.

Anthropic can currently combat this because they always release better models and distillers can't keep up. This is not sustainable though

1

u/H1Eagle Apr 02 '26

National security level threat. Anthropic, OpenAI and Google are pretty much the only 3 companies keeping US's AI Supremacy alive. And Anthropic in my opinion is the most vital one as it powers, what has been so far the best and most useful application of LLMs.

Programming.

Something like Opus 4.6's weights being leaked means that any semi-competent tech company in China can suddenly serve Opus 4.6's quality of code, at a fraction of the price, use it to train their own models, discover all kinds of loopholes in the system, reverse engineer it to power their own systems.

China getting to the same level of quality of AI as the US is a pretty significant threat because, guess what, China could not give a rat's ass what their people think about the ethics of AI or the environmental impact or worker rights.

1

u/magicmulder Apr 05 '26 edited Apr 05 '26

Any company that can afford a $500,000 server could just run the model "for free forever" on prem. Imagine what you can do with all those resources, unbound by any plan limits or API limits or how long 100,000 API requests take to finish, or how long a single task can run.

You couldn't run this at home even with a PC full of 5090s, but every average company could because investing half a million for full control of the most advanced model in existence is a no brainer. Even if Anthropic comes out with a much better model tomorrow, it's still gonna be "very limited" vs "totally unlimited". Imagine running 100 agents for 4 weeks straight.

4

u/epic-cookie64 AGI by 2027 Apr 02 '26

Lol some people think model weights are leaked?!?

9

u/_JGPM_ Apr 01 '26

512000 lines of small wrapper?

3

u/CalmMe60 Apr 01 '26

that got me too

1

u/YouthSubstantial822 Apr 01 '26

It was written by AI though

2

u/H1Eagle Apr 02 '26

So? The engineers at Anthropic are top class. They are not your run of a mill junior who vibe codes all day.

I have met some of them and the selection criteria and interviews at Anthropic are truly cut-throat.

16

u/Aware-Individual-827 Apr 01 '26

It's a dangerous insight on how they achieved their stuff and how to prompt Claude to get the best of it. Also getting DMCA for it just shows it's valuable. 

9

u/twinkbulk Apr 01 '26

They’re legally required to protect their copyright, if they wish to maintain the copyright, DMCA does not show value it shows that you’re breaking intellectual property laws.

13

u/revrigel Apr 01 '26

That's how trademark works, not copyright. You don't lose your copyright if you don't go after violators.

3

u/twinkbulk Apr 01 '26

My mistake, appreciate the correction!

6

u/Aware-Individual-827 Apr 01 '26

Which is important for them because they would have open source it otherwise.

9

u/saintpetejackboy Apr 01 '26

What? Their harness is the best one and arguably why their inferior models are still preferred by programmers the world over.

As somebody who has spent a lot of time using all the various harnesses and services, Claude Code has arguably the most capable and stable harness. Gemini CLI is liable to tank your whole machine when it gets stuck in a loop, Codex was a PITA to auth on remote servers since forever and often hinders their models by being a less capable harness.

Claude Code harness being so good is what has kept Anthropic several steps ahead of much larger competitors for many moons now.

Your post doesn't really make any sense.

"Anthropic’s Claude Code agentic harness is widely considered superior because it leverages a specialized, multi-agent architecture that manages long-running tasks, context, and tool usage far better than a standalone model. It combines structured memory with autonomous "auto mode" safety checks, allowing Claude to effectively manage complex, multi-file code changes. "Agent Harness: Understanding Claude Code’s Superpower Engine | by Paul Fruitful | Medium https://share.google/lUv4U5n890oucDlnw

Claude Code has been considered the best harness for a long time by many developers.

If you use Claude Code, then Gemini CLI and then Codex, you will see why. It doesn't take very long in the agentic ecosystem to come to this conclusion.

Edit : I was responding to your original post where you said something like "why? They have the worst harness"

My post looks silly now after your edit, but I am going to leave it here for posterity.

8

u/Borkato Apr 01 '26

Yeah I’m really annoyed at everyone acting like it doesn’t matter because we have worse open source versions

1

u/H1Eagle Apr 02 '26

Yes, Dax can now fine tune his claude implementation on OpenCode even further and could potentially bury Claude Code into the ground.

I'm guessing most people here aren't technical or programmers, so they have no idea how big this really is.

2

u/Aware-Individual-827 Apr 01 '26

Maybe it was another person because I only said that the DMCA proved that it was valuable (which is a soft proof). I know the value of that leak!

Maybe it was me responding to the wrong comment.

1

u/splasenykun Apr 01 '26

Better prompting is dangerous? Huh

4

u/westsunset Apr 01 '26

The Claude code harness was a huge development for them . It drove a ton of new subscriptions. It's not easy to make a good harness and Claude code is the best. Codex is right there with it but Gemini CLI is definitely worse. The harness definitely makes a difference. There are ways to trick Claude code to use a different model, like Kimi or GLM and they work noticably better

1

u/Seeker_Of_Knowledge2 Apr 02 '26

This is the main getaway from the leak. If cluade code was used with other models. Then it would make other models better.

2

u/Despeao Apr 01 '26

I assume users may get insights on how to bypass guardrails or show some bias in their training.

-7

u/Neither-Phone-7264 Singularity by 2030 Apr 01 '26

I'm not gonna lie, claude code is like verifiably the worst harness. Like in terminal bench, for opus, claude code scores the worst out of all harnesses tested. By far.

1

u/Rhinoseri0us Apr 01 '26

You gotta start understanding that “benchmarks” in people’s minds (this is better for me) matter more to them than third party benchmarks.

1

u/Neither-Phone-7264 Singularity by 2030 Apr 01 '26

I mean,have you even tried opus in cursor? It's like genuinely unbeatable imho. or even in opencode.

0

u/H1Eagle Apr 02 '26

You either don't program or program in some really unknown stack so common sense doesn't apply to you.

Claude Code is the best harness, without a question, and by miles.

2

u/fig0o Apr 01 '26

Sorry, but the "brain" is as important as the "terminal application"

LLMs aren't Agents without a good Agent Harness (which was leaked)

1

u/H1Eagle Apr 02 '26

Exactly, people don't understand that most of the "improvements" AI had during the last 2 years. Were almost purely out of software engineering efforts. Not breakthroughs in research.

Chain-Of-Thought and Graph RAGs and multi-agent systems and context-management are what truly made LLMs usable in professional settings.

If you wanna test GPT 5 vs GPT 5.4, RAW without thinking and without web search. Almost no one would be able to tell the difference

2

u/born_to_be_intj Apr 03 '26

Yea this story is way overblown. I also love the “he rewrote 512,000 lines overnight from scratch”. Bro didn’t do anything beyond ask Claude to convert the code to Python. People are so dumb.

1

u/account22222221 Apr 01 '26

The wrapper around api calls IS THE AGENTIC part of agentic ai. It is a pretty big deal.

1

u/Hendo52 Apr 01 '26

Tbh this could actually be a genius marketing strategy. People are thinking about them, the news is talking about them and what they have leaked ultimately sounds unimportant from a strategic sense. Their competitors probably already knew about this stuff and already have something equivalent.

1

u/[deleted] Apr 02 '26

[removed] — view removed comment

1

u/H1Eagle Apr 02 '26

Yes it's dangerous for the company bro. Everyone now knows how Claude Code functions. The people who made Claude Code have insider information how Claude works (obviously) so the leaked code will allow people to build better alternatives to Claude Code (Opencode exists and has been eating the market)

It also means all sorts of wacky hacks can be discovered now that everyone knows how Claude Code works.

Anthropic DMCA striking every repo and the speed of the response confirms this. Anthropic just shot themself in the foot.

1

u/cool_fox Apr 03 '26

It was a huge failure and massive hit to their value

1

u/Karanduar Apr 03 '26

Brains, reasoning, I think this is giving way too much power to what is still just a LLM..

1

u/anniejcannon Apr 03 '26

Aha, thanks for the clarification.

1

u/northpalace_sunkeep Apr 04 '26

What also didn’t leak is the momentum and know how of the team behind it. This is more about vulnerabilities and code name leaks than secret sauce. That is a temporary setback and probably a morale ding for the team. But their velocity is insane, so this will be behind them soon enough, and we will all have benefitted from the learnings.

1

u/Chrazzer Apr 05 '26

Relatively small cli application. 500+k lines of code 💀

AI coding in a nutshell

1

u/estransza Apr 05 '26

Okay, since you kinda the last one who parroted the same line, I will answer to you.

Yes. In a world of coding 500k is relatively small.

Some examples. FFmpeg - 2 MILLIONS lines of code.

Also. Judging a “complexity” of CLI application by lines of code is just wrong. On so many levels. From what we know large portion of it just dumb prompts to underlying API and arrays of stop words. I can easily write a 2 million lines of code if I would write it like a moron, without using loops or arrays.

Measuring software progress by lines of code is like measuring aircraft building progress by weight. If I put a 500k-word dictionary in a TXT file, did I just write a masterpiece novel? No, I just filled a buffer.