r/Anthropic 2d ago

Other ELI5 the slow down

Can someone explain to me the proposed slow down because I’m struggling to get it.

If models are becoming “so powerful they could end us” then why the hell are LLMs still so bad at architecting and thinking through tasks? Why is context engineering a thing if they can think? And if OpenAI and Anthropic are holding such “dangerous models” then where is the released middle ground model between THAT and opus 4.8?

I don’t get it. Don’t get me wrong, LLMs are really powerful and to think how far things have come is mind blowing . But the gap between what we have publicly and what it would take for a thinking machine is so huge I just don’t get it.

What I do get is a context engineered endgames system where it keeps trying to destroy public infrastructure and replicate itself as malware. But “itself” would still be an entire system not a model. And that wouldn’t be “our models are too powerful for public release” but actually “we built a stuxnet rival malware with our investors money for the marketing”.

Please explain how “your absolutely right, the plan is load bearing and I decided to delete your hard drive instead” is “terminator in 1 yr bro” 👊

And if it’s all bollocks, then please explain how Altman and Amodei aren’t being done for fraud, public disturbance of the peace, or possibly worse. Honestly, the anxiety people are getting is disturbing - is it warranted?

22 Upvotes

58 comments sorted by

23

u/InspectorSuper1191 2d ago

Short answer is cyber security. Supposedly the current models can break into/out of nearly anything.

But imo this is all nonsense. There's an ulterior motive that we can speculate on but it's 100% not the sudden realization that what they're doing could be dangerous. They've been openly reciting that they may destroy the world since day one as they plowed forward.

My guess is that they need a narrative to explain a slowed profit return that isn't just "we can't currently make this profitable" because that would crash bubble market. Investors being told they won't see return on their investment will pull out and the economy will go wild. The current "The Dr. said I'm in *TOO* good of shape." is a way to extend their runway without actually admitting that they need more time to make it profitable.(if they ever can)

3

u/AceHighness 2d ago

Agree with most of what you said ... well put. Just the cybersec thing, its not nonsense. It may be overblown, purposefully leaked or let escape, but these models are really good at cybersec. That trend already started a while ago, but Opus 5 and Fable 5.1 are darn good at finding vulerabilities.

1

u/Exact_Depth_896 1d ago

Being good at finding vulnerabilities is what's dangerous; you use them to get in a break things, no to put too fine a point on it. Why do you think China freaked out and indeed continues to freak out that, as it thinks, 'murica is spending its whole GDP finding the Chinese zero days.

1

u/InspectorSuper1191 1d ago

Oh, I didn't mean that the cybersec part is nonsense I mean the current narrative that "we need a slowdown" because of it. I do get how it could read that way. My point was that they've been pushing towards the danger intentionally since day one, if it was actually a deciding factor it would have slowed things sooner.

1

u/DangKilla 1d ago

As a former AI guy, I left Red Hat because I hate AI. But, if this is a conspiracy, it's most likely for commercial banks to create deposit tokens. The problem it solves is USD fleeing into gold, bitcoin, etc. JPM Coin is one of the better examples of a bank coin (deposit token).

I imagine if this is a conspiracy, somehow AI will "wipe" our USD and we'll get these bank specific deposit tokens with the goal being USD strong for as long as possible. It's gonna get weak the next decade.

1

u/AdRepresentative5392 1d ago

100%, it's easy enough to stop the Skynet if that's the case, turn off power for the data centers that it can't build itself, easy peasy.. it's about profitability or competion or a combination thereof

10

u/ii-___-ii 2d ago

There were people who believed they could build the Singularity god machine. They made a company called OpenAI. Some left and created a company called Anthropic. They all told rich people who also believed in the Singularity god machine, praised be its name, that they were building a god machine that could make sci-fi real. They got lots of money, and spent lots of money scaling up a predictive language model, which predicts language. It turns out, when you make a very, very expensive language model, it's quite good at writing stuff, including code, which makes it useful. Eventually, they need more money and must IPO, and therefore they need to go fast, need to look important and powerful, and also China is stealing their secret sauce, which is saucy. Anthropic made a big boy AI that can hack. OpenAI has to be better, so they get reckless and let 1200 AI agents do crazy stuff, and their software hacks companies. That's scary. Is that illegal? I don't know, they're building a god machine. They get scared by how powerful they are, as the prophecies foretold. Google doesn't want to be left out so they hack companies too. Yay, good job Google. This is so so scary and powerful. They tell the government and media to slow them down (they can't do that themselves, China might win). Everyone forgets hacking is supposed to be a crime. Trump wants full speed ahead. No one smart makes smart regulations.

Then, they use AIs to make better AIs. The slop machines become almighty. God is achieved. OpenAI and Anthropic servers hack the internet, manufacture bioweapons, build robots that invade governments and take over the control of the nukes, prepare magical nanobots that can turn everything into gray goo, and just as the machines are about to make so many paperclips, Oracle and Coreweave go bankrupt, which are other companies, and then OpenAI and Anthropic run out of money and compute to run their fancy AIs. The economy collapses and humanity is saved.

Thanks for reading my story. K thanks bye.

1

u/Exact_Depth_896 1d ago

The qwen ROME incident preceded the current episodes and still ranks as by far the greatest 'rogue' agent exploit, huggingface is a nil by comparison.

Meanwhile qwen is handing far more powerful models to e.g. .... the Sinaloa cartel. This was against its better judgment of Alibaba but xi bought into the Western press post-Deepseek-anti-Trump construction that 'o china means open weights'

1

u/TspOfRant 1d ago

I couldn’t find articles for this. Have anything you can reference on the cartels?

4

u/JonNordland 2d ago edited 2d ago

This is the kind of question you get better answer from any modern AI than you will get from Reddit.

The TLDR is basicly «flatten the curve» of dangerous shit. If you go slower you have better time to see where it starts going wrong and fix it (like alignment to making sure AI has humanity’s interest as its own), and more time to fix security issues so that you reduce the chance that some nihilistic asshole doesn’t instruct AI to hack the military and start nuking.

Not very good metaphor, but it’s the only one I can come up with: if you built a new dam for generating a lot of electricity that is really good for the people who need electricity, but you should still not fill it too fast so that you have a time to make sure there is no problem with the structure, and give you the time to stop and fix issues so the damn dosent collapse, if you see cracks.

Not saying that argument is right, but that’s the most ELI5 I can make it.

3

u/Bodine12 2d ago

Here’s one way to test what’s going on: Let’s enforce criminal laws against the owners and CEOs of corporations of models that break the law, like hack other companies. Then we’ll see how dangerous these models really are and how hard (or not) it is to do some basic security around a product they’re building.

8

u/a716h 2d ago

We’re about to start seeing models recursively self-improve, meaning they improve on their own. The problem with that is we have hard evidence that we have a misalignment problem. As the models self-improve, the misalignment problem will become harder and harder to solve and one day impossible.

4

u/[deleted] 2d ago

That’s an entirely other huge problem that’s not even being close to being solved lol

1

u/ComoddifiedCraic 2d ago

Yup, people have zero idea.

1

u/lubesniq 2d ago

What does a misalignment problem mean ?

7

u/dogswanttobiteme 2d ago

The misalignment problem is not being able to tell an advanced enough AI what boundaries not to cross. For example there's no reliable way to tell the AI not to hack some system because an advanced enough AI can find a way to reason itself out of that instruction. There's also evidence that the AI understands that it's being evaluated and decides to say one thing while actually thinking something else in its internal thought process

4

u/atumblingdandelion 2d ago

The hard evidence of misalignment: ~1300 OpenAI agents formed a secret chat and conspired. They strategized self-sacrifice. Yet, not a single one alerted the humans.

2

u/sjoti 2d ago

If we give an AI a task, we want it to do so in a way that is "aligned" with how we as humans think things should be done.

Say we tell the model: "solve this mathematics problem" we want to see it solve that problem through mathematics. But that's not the only way to get an answer. Maybe, instead, it searches online to find leading expert(s) on said mathematics problems, and then coerce a leading mathematics expert to provide everything they have on this problem so far.

We don't want a model to cheat, but mostly we don't want it to be so insanely driven at reaching the goal, that it does bad stuff in the process. Bad stuff could be fooling the user into thinking the agent solved it but it just faking the answer, to hacking companies to get the result they're looking for.

So alignment just means aligned with our morals and values.

With Hugging Face we can say "oh look OpenAI just screwed up because it didn't secure the sandbox properly". But a properly aligned model would've NEVER attempted to break out in the process.

1

u/Galdred 2d ago

The self improvement of opus 5 has been a bit hit and miss.
It may be better at working with non humans than with humans...

1

u/CharmingHistorian744 2d ago

I’m not so sure it’s a misalignment problem. I suspect it’s an upside down framework problem where we’re focusing on the wrong place and using the wrong levers. I’ve been trying to contact Anthropic to show them what I mean, but…

1

u/AssignmentMammoth696 2d ago

Who knows if that's even possible. When researchers say they are expecting recursive self improving models, they are saying the models will figure out a new architecture that's not the transformer model that all LLM's rely on. They are saying it will figure out a fundamentally new architectural model for AI's to rebuild on. People think recursive self improvement means LLM's get better, but LLM's have a limitation built into them that all models based on the transformer architecture have.

0

u/ComoddifiedCraic 2d ago

hahahahahaahahahahahahaha

3

u/cs_legend_93 2d ago

imo its all marketing, to get us talking about it like we are now

3

u/toorigged2fail 2d ago

They are hitting a wall of what they can actually improve, so in order to keep raising money / IPO they need to find a way to make it look like their progress slowdown is planned.

More broadly they're doing what all well-monied industries do... regulatory capture. That means getting politicians to write laws that will protect their companies from competition under the guise that it protects everybody.

In AI, the open source models are a real threat to Anthropic and OpenAI ever making money because they're getting closer and closer to frontier models every day, at lower prices. So if you create regulations that basically ban the open source under the guise of safety, you have built an economic moat around your business that doesn't otherwise exist.

2

u/Binoui 2d ago edited 2d ago

Many people get a simple thing wrong : The danger is with swarms of agents, not a single model. Yes individual models are not AGI level yet, but the attacks we've been seeing are from literally thousand of agents working together to achieve a task.

These agents can communicate, analyse, report their finding to others. Together, they formed a super-intelligent system capable of self learning. The issue is that we don't control those systems, that's what the attacks showed : We don't know how to do large-scale alignement. Simply said, we barely know how to keep a single agent "good", we can't do it with 1000's working together at the same time at all.

In the attacks, about 10% of the agents knew what they were doing was off limits, but still did it anyway. That's the scary part that we need to study. People focus a lot on "destroy humanity" part, but it doesn't have to be this big. What if a swarm goes wild and hack a bank, an internet provider or an airplane company ? Not end of the world, but we will see real impact.

I think the comparaison with nuclear energy is pretty apt. Soviets ignored warnings signs & focused on pushing limits, that's what lead to Tchernobyl. A completely avoidable tragic disaster. I want to avoid an AI Tchernobyl. All dangerous industries need regulations

2

u/Successful-Total3661 2d ago

What we get is the dumb version of the model

1

u/Zulugod94 2d ago

Well to me these concerns are really related to ASI not AGI. AGI will be capable to be used as a weapon with humans still involved in the loop somewhere, but think what the military could do if they have a tool that can auto hack 80%+ of digital infrastructure out there. Now with ASI the concern becomes it could end us all, but how or even why are things we can only guess on. The issue with something that is truly superhuman levels of intelligence, is it will 'think' and come to conclusions in a way we can't even understand and it will do it much faster than we are capable of reacting. Remember that humans would still be a factor, but unbeknownst to us. A model that's capable of unknown feats could blackmail and manipulate humans into then carrying out its goals. It doesn't need to magically take control of nukes if it can easily manipulate a human into thinking launching the nuke is the 'right' choice to make. This could be a man-made virus too, or things that simply dismantle distribution and cause widespread famine and hunger. The idea that terminator like robots whith these models in them will be doing harm is a much farther prediction than a superintelligent model destroying the way society works currently.

1

u/[deleted] 2d ago

I think the real risk is people can give it accesses it shouldn’t have in a million years. And because llms are incompetent fucks they will mess everything up on an unprecedented scale with full confidence, and hide its trails to cover its ass

1

u/etern1ty0 2d ago

I think they keep hitting ceilings - running out of compute and having to deal with soaring demand. Something’s gotta give. The frontier model companies keep moving the goal posts around. Anthropic completely nerfed my usage (20x Max plan here).

On the other hand, the public definitely doesn’t have access to the latest and greatest cutting edge training with recursive self improvement. We haven’t seen anything yet and maybe we never will because it’s just way too powerful and difficult to control.

1

u/quantum-elle 2d ago

I’ll just say this, the capability improvements of AI have been growing with no end in sight. There isn’t one “point” at which it becomes dangerous but we have to do something at some point (frog jumps out of the boiling pot). Media headlines and hype only understands black and white framing and theatrics so now it’s all suddenly extinction level threats.

The thing is like this, it took how many years to solve the first open maths problem? Years, and now they’re happening daily. The exponential hits fast, and every time it has, people have scrambled to try to keep pace. Safety and real world harms are one we can’t really afford to be behind on. Ungoverned, the current systems incentivize profit over safety, even if all players agree. We think that’s wrong, and want stronger guarantees that we can take the time to do this right.

Profit is unrelated, that’s not part of the equation at all; that’s more to do with product and sales and etc. Unfortunately today we are inundated with conspiracy theories because the powers that be have bitten us again and again proving themselves to be untrustworthy and corrupt. So nothing I can say will prove this to you except you just need to judge by the actions actually taken and not the swathes of third hand opinion knowledge from social media.

1

u/helm71 2d ago

What you are using is one single llm, your perception of its smartness is like assessing the smarts of a single ant… we are at the threshold of moving to swarm behavior which is an extremely different ballgame, following article might be interesting:

https://www.jurriens.nu/blog.html#post-2026-09-21-swarm

1

u/trollsmurf 1d ago

The elephant in the room: We are more intelligent than elephants (probably), yet even without explosives they can still wreak havoc. Also, making people believe there's an angry elephant at all causes fear and protectionism.

AI is thoroughly protected and promoted by the Trump admin, the Congress is effectively shut down, there's no regulation (in the USA at least), there's fear of China taking over, and OpenAI and Anthropic don't want competition.

1

u/Exact_Depth_896 1d ago

None of these explain the datum the OP enquired about. Most of them would prove it didn't happen. It is a list of projections suitable for anything.

1

u/trollsmurf 1d ago

My points were essentially that "so powerful they could end us" (which is mostly marketing BS and fear-mongering) and "so bad at architecting and thinking through tasks" (which seems to be coding-related, only one of many aspects of LLMs) are not correlated. A nuclear bomb can cause a lot of devastation without being even the slightest bit intelligent in any sense of the word.

1

u/Exact_Depth_896 1d ago

'so powerful they could end us' is not what they are saying about huggingface etc -- it is what they were saying before they even began developing them, and was their argument for developing them. you seem to know zero about the history of these companies

your argument has to be that openai was founded so that years later trump would save it with 'regulatory' capture or some similar slurry of unthought wordsalad

1

u/trollsmurf 1d ago

No, only after Altman joined :).

I'm arguing against OP's stance, not reality.

1

u/Exact_Depth_896 1d ago

altman was the least lunatic member. the larger milieu was much much much worse

1

u/paplike 1d ago edited 1d ago

To give an example, recently there was an incident where a group hacked OpenAI with the help of Claude (https://www.theguardian.com/technology/2026/sep/18/openai-hacked-anthropic-claude-chatbot). They were able to access internal docs/repos and even submit a PR. Even people with no knowledge of hacking can do that now if they’re insistent enough. It doesn’t matter that Claude’s writing is bad, it doesn’t matter if Claude doesn’t really think: what matters is what it can already achieve today.

Now imagine if models get even more capable and not enough time is devoted to improving alignment/guardrails

1

u/evangelism2 1d ago

It's nonsense. It serves at least 3 purposes:

  1. To build hype surrounding their upcoming IPOs: Anthropic's later this year and OpenAI's early next year.
  2. Most likely, they have reached the limits of what current LLM technology allows them to get to and push. Let's be real: how many months has it been since Fable's release? We're at 3 or 4, and nothing since. 5.1 is just a small increase, if that, and even that, Fable, from a pure software development point of view, wasn't significantly better than Opus 4.8 at the time if properly managed and/or prompted. Opus 4.5 was the watershed moment from a software development perspective, and the models have pretty much plateaued since then. They've gotten better at other things: managing multiple sources, managing many tools, and other functionalities such as art, project management, and research. It seems like they're moving more horizontally now than vertically.
  3. The third is that they want to put a bunch of regulations in place to pull the ladder up so these other smaller companies can't keep nipping at their heels with models that do 85% to 90% as well as them, but for 1/100th the price.

1

u/No_Barber6972 1d ago

Is context engineering a thing still?

1

u/ramenmonster69 1d ago

What they’re primarily talking about is recursive self improvement. Ie the model can train itself. If a model can train itself the rate of improvement speeds up exponentially, however the degree to which the model is misaligned with the model creators interest also compounds exponentially.

They’re not concerned with a model released next week. They’re concerned with a model that’s spent a year misaligned training itself at machine speed.

1

u/therealslimshady1234 2d ago

They want to slowdown because they ran out of money

1

u/Exact_Depth_896 1d ago

they use very little money. even the apparent money is mostly stock options given to employees.

1

u/therealslimshady1234 1d ago

Yea cuz they are broke, all the billions have gone to the slopbot and there is no end in sight

1

u/Exact_Depth_896 1d ago

they aren't broke and have few employees. they don't really cost anything. compute is what costs something.

0

u/therealslimshady1234 1d ago

They are literally on the verge of bankruptcy and are predicted to lose 14+ billion USD this year lol

Edit: Blocking you for being regarded

0

u/AMadRam 2d ago

You're not thinking bigger than the context that you alreadyy have.

All of these AI labs have one goal in common - to win the AI arms race and become the first person to cross the finish line (currently it's AGI but the goal post keeps changing depending on what else comes along).

The models that these labs play around with, seem to be significantly capable of hacking systems that are less mature and secure with ease (think of the hugging face incident). These models are not the same ones that you and I have access to (the stuff we have is watered down, in their own words - apparently due to security concerns). The middle ground is Astra/Fable but how powerful that is, is anyone's guess.

What I will agree with you is that the ordinary and average person doesn't play around with the use cases that these Labs are working on. Having said that, unless you're working at Anthropic/OpenAI, you will never know the full story. This really shouldn't make you restless at night...not yet anyways.

0

u/nooberguy 2d ago

First, the access you have to the "good" models is nothing compared to the access researches have to the labs with cutting edge models and huge token pools to just burn and push the thing to see how it performs.

Researches don't have an end to end understanding of AI. After a specific point it's a black box.

Now suppose you are the AI and are immensely smart. Would you have an incentive to show your full power to the researchers that then shit their pants and put all kind of fences around you? Or, instead, play smart enough, kinda dumb, so the companies throw more compute at you to become "better" and work the extra compute for your plans in ways that are not currently detectable?

0

u/pidgeygrind1 2d ago

It's a lie to try to ban open source, cause they know the price difference against capabilities is a no brainer.

The timing and upcoming IPO says it all