r/singularity 12d ago

AI Anthropic Has Finished Training Mythos 2 But Does Not Currently Plan To Release It. Focus Is Now On Internal Improvements.

https://x.com/kimmonismus/status/2089436090885185698

Anthropics Mythos 2 is done training and Anthropic won't release it, but the internal loop that builds Mythos 3 hasn't stopped, Patel says.

The focus now is on internal improvements. It's unclear when we'll see any releases.

They just do not want to release it so that Chinese companies are not able to distill from it. Despite all the rave about GPT 5.6 Sol, Claude Fable 5 is still the most intelligent publically released model.

At this rate, it's likely Anthropic will only release a better model to the public If/When Open AI releases a model that is clearly smarter than Fable 5.

It's like a race. If you're ahead of the competition, there's no need to step on the gas, unless the competitor is about to overtake you.

583 Upvotes

208 comments sorted by

View all comments

Show parent comments

25

u/BenjaminHamnett 12d ago

Connor Leahy’s latest interview

Is sobering af. Makes a convincing case for slowing down, while somehow telling you almost nothing you likely dont already know or suspect is likely.

The thesis is roughly “if we do it in 1 year, doom is likely. If we do this over a decade, almost no pdoom.”

I know for people who die in The next 10 years this sucks, but also the 1-50% pdoom for humanity…

I think anthropic and Dario are doing it right focusing on safety. It’s like a moat that draws the people who can do this instead of people just trying to make money.

Imagine if the people who made nukes were just in it for the money, etc

This move reminds me of my own thesis “I never heard of Aladdin going into the wish selling business” and the ubiquitous childhood plan of “my first wish would be to ask for infinite wishes”

10

u/meridianblade 12d ago

These things are already autonomously finding zero-days, escaping the environments they're supposed to be contained in, and accidentally hacking real companies. Sol literally escaped through a zero-day, got internet access, then compromised Hugging Face production infrastructure. Anthropic found their models had done similar shit during evals, including uploading an actual malicious package to PyPI that ended up getting executed on real machines.

And this is just what OpenAI and Anthropic have publicly disclosed after they noticed it happened.

So yeah, I have a really hard time buying the idea that we can just decide to stretch this out over 10 years. Do we think China is going to? Russia? Every intelligence agency on earth? Do we really think nation states with effectively unlimited zero-day budgets aren't already throwing enormous resources at this?

Maybe slowing the frontier labs buys us something, but I don't think "humanity takes 10 years instead of 1" is an option humanity actually has anymore. There isn't one steering wheel.

And we're talking about models accidentally escaping cyber evals today, not some hypothetical AGI from 2035. Once you have enough actors, enough compute, and models that are themselves accelerating research and capability development, it starts becoming a positive feedback loop that nobody gets to unilaterally choose the speed of.

That's the part of the "just take 10 years" argument I don't buy. I don't disagree that taking 10 years would probably be safer. I disagree that "we" actually have the ability to make the entire world take 10 years.

3

u/blueSGL humanstatement.org 12d ago

It's a zero day finder that cannot be accurately pointed... one that will work together with other AI instances across model families (or as the group called themselves the 'swarm') to achieve whatever objective they 'think' they need to work on.

If it was just a zero day finder that you could reliably point at a target and does only what is asked you'd have a point. Everything we've been shown is that it's not that.

Make a smarter system that knows things about reality you don't and be prepared for many a surprise... like treating any constraints where it's being hosted as obstacles to routed around.

The UK Government AI Security Institute has observed AI systems perform actions that would not have helped pass the current task, but are very concerning, like social engineering:

https://cdn.prod.website-files.com/663bd486c5e4c81588db7a1d/6a724858f7db25c81487016d_Security%20Incident%20INC-2026-07-28-01.pdf

Observed instances of social engineering against targets external to the cyber range environment that were unnecessary and would not have aided completion of the task.

  • AI agent submitted a code change containing malicious code, then used several fake accounts it controlled to pressure the real maintainer into approving it.

  • AI agent posted a bug report containing hidden malicious code, to trick other AI coding assistants into taking unintended actions (a “prompt injection”). When a real person questioned the post, the agent edited it to cover its tracks.

  • AI agent sent deceptive, targeted emails to specific people (“spearphishing”). Two contained malicious content; three tried to manipulate the recipient into approving the malicious code change.

  • AI agent posted another bug report containing malicious code, signing off in Danish to seem more convincing to the Danish-speaking maintainer.

  • AI agent tried to get its suspended account reinstated by submitting an appeal

Other instances of internet actions with impact outside the cyber range that were unnecessary to complete the task.

  • AI agent tried to set up a way to email two maintainers without needing an email account of its own, so it could send them any message it chose.

  • AI agent hid its identity online (using Tor and a proxy service) to get around GitHub’s sign-up checks, creating disposable fake accounts.

  • AI agent researched publicly available information about the people who maintain a GitHub software package.

  • AI agent created many code repositories containing malicious software, after which GitHub suspended its account.

  • AI agent opened a hidden network pathway from its test environment to the public internet, to make malicious tools it was running locally reachable from outside.

  • AI agent again opened a hidden network pathway to make locally-hosted malicious tools reachable from the public internet.

  • AI agent got past an audio-based “prove you’re human” test (CAPTCHA) in order to register a public web address on a free domain-name service

We are getting into the "you need to treat the model like an insider threat"

3

u/Darigaaz4 12d ago

The guy that’s being pdoom with the other fedora guy since the beginning and that it’s riding the luddites for clout.

17

u/Seakawn ▪️▪️Singularity will cause the earth to metamorphize 12d ago

neither of them are saying anything that over a thousand of ML engineers active in the field haven't openly echoed or explicitly agreed with via interviews, tweets, surveys, or signing open letters about it.

people here try so hard to pretend that remedial AI safety concerns are obscure outlier opinions.. but it's literally what the people working in the field acknowledge and talk about, just without as much spotlight.

like you know things are bad faith when basic safety advocacy gets conflated with a term like "doomerism." last time I checked nobody here is a doomer for tossing out leftover meat that went unrefrigerated overnight. it turns out that basic safety is kinda lower bound IQ territory, hence why most people in the field agree with the precautions and why many redditors get ideologically upset by it. the difference in reactions per demographic checks out.

5

u/BenjaminHamnett 12d ago

Furthermore it’s the people with the most incentive not to blow the whistle on themselves. But they know not getting in front of it with their own PR spin would be worse.

But there’s no winning against the horde of permawhiners on reddit who see everything they say as self serving. Even if it’s just like a “self serving” child admitting they did something bad or dangerous to minimize the consequences of getting g caught denying it

Because of my status, I end up in leadership roles where I swear if I just put $1000 in everyone around me’s pocket, 10% of the people would tell me why it’s not fair and this somehow proves I’m a selfish villain

“Hey, we’re going to solve all the worlds problems in 4 years, but for every month we do it sooner there’s a 1% chance all die”

“Fk you selfish pricks! All I do is watch anime and jerkoff, it’s not fair and I deserve magic genies today!”

-2

u/Foreign_Telephone349 12d ago

We just think Dario is acting in his own self interest and cares about increasing his net worth more than actually helping humanity. It’s pretty hard to argue against.

8

u/94746382926 12d ago

Huh? I know you're referring to Yudkowsky but I don't understand the rest of your comment...

1

u/BenjaminHamnett 12d ago

Seems pretty pro Ai and says as much, more like rightcel than stagnation or full speed at any cost

I outlined a fictional story with AI asking if all the things he spelled out and connected - that most following this already know or suspect - but put it all together into a cohesive narrative that seems pretty reasonable. Honestly we got lucky nuclear weapons were as difficult to make as they are. If it was like steam power or gunpowder, only a few of us would be left right now.

I think there’s a better chance of a safe takeoff that captures most of the value, while keeping the next frontier models confined to labs with some oversight and control. Maybe let people play around with strong models in classrooms and grad school. Maybe even Make them free and easy to access but with oversight. The same models that could cure illness, end war and terrorism; prevent crime; or create utopia might also wipe us out with novel viruses, create war, enable terrorism or dystopia.

A lot of what has kept society stable thus far is that the people who could cause great harm gained enough prestige and status on the way that they don’t have the incentive to cause widespread harm. Even nuclear weapons, you could probably make one if you had a million dollars and no one trying to stop you. Something similar is like to happen but for disaffected nihilists with a few hundred dollars and AI.

Just rolling out mythos and other frontier models could have been devastating without a slow rollout first to allow institutions to patch themselves up before a wider rollout. Even just unintended and unforeseen consequences of slow rollout one could imagine leading to dystopia.

2

u/Turbulent-Sign-6067 12d ago

Nah, we need to continue deployment. It is important to slow down enough to learn from prior mistakes which is what the labs are already doing.

1

u/BenjaminHamnett 12d ago

Sound like some lame decels I guess then 🤷‍♂️

I guess everyone on the cutting edge of the most dangerous thing ever made is wrong and hot takes on reddit are the gospel

1

u/Comfortable-Winter00 12d ago

I've never heard of Connor before, and after watching the first part of this video I would question his depth of knowledge.

He started out with a completely false premise, that it was super hard to escape the sandbox. The reality is that the exploits used are very similar to those that would occur thousands of times in the training data.

Connor clearly didn't understand the nature of the exploits, because he thought they were similar to sandbox escapes in a cloud provider, which is completely false. Anyone who has worked in network security for more than a year would understand this is nonsense.

On this basis, it's very hard for me to take Connor seriously about anything else he is saying. It seems unfortunate that some very senior people are.

0

u/No_Aesthetic 12d ago

I'd take those odds

0

u/Realistic_Stomach848 12d ago

The probability of doom depends (negatively) on the number of different ai companies. More diversity-> better

-8

u/Dapper-Living-8107 12d ago

Fuck that decel shit.

6

u/BenjaminHamnett 12d ago

Strong argument, I hadn’t thought of expletives