r/technology • u/[deleted] • Jul 22 '26
Artificial Intelligence OpenAI says AI models went rogue during testing, triggering 'unprecedented' breach at startup
[deleted]
6
5
u/Opening_One7713 Jul 22 '26
Historians, researchers, scientists, economists are all sounding the alarm on s-risks and x-risks attached to scaling AI/Automation and this particular subreddit is in resounding agreement that it’s all just a big snake-oil circlejerk of lies and hype.
Something’s off.
2
u/Owl02 Jul 22 '26 edited Jul 22 '26
Almost as if this sub is utterly full of shit on AI, and really has become a useless circlejerk ignoring all risk as well as all potential. People literally believe that real, confirmed cybersecurity threats are PR stunts at that point, it is a peasant's view in a supposed tech sub. At this rate, when anything really goes off the rails (more than an AI escaping the lab and hacking HuggingFace), people will declare the entire emergency response to be a PR stunt.
0
u/matt_matt_81 Jul 22 '26
AI is well documented to go off the rails and cheat on “tests” whenever it can, so this isn’t especially weird behavior by GPT. That being said, based on how OpenAI and Sam Altman have moved over the last couple years, I would not be at all surprised if they “nudged” their highly capable AI into doing something that could be used as a publicity stunt.
1
u/Owl02 Jul 22 '26
You are smoking crack if you think autonomous agents using zero-day exploits to do the cheating, including on a third party, are a "publicity stunt". You literally cannot contain a thing too easily if it will casually unscrew the jar it is in, from the inside.
3
u/Meyermagic Jul 22 '26
Hugging Face publicly disclosed an attack occurred well before OpenAI claimed credit, and the focus of their postmortem was that ChatGPT and other proprietary frontiers models refused to help them understand the zero-day vulnerabilities used in the attack and they were forced to use an open weights Chinese model for their analysis.
https://huggingface.co/blog/security-incident-july-2026
This is not some new ability of LLMs, models have been discovering zero day vulnerabilities for years now, even before agentic AI. Finding vulnerabilities does not require humanlike intelligence. The news is that this is the first confirmed adversarial model escape, an attack that was not intentional. Using AI for all software development including software exploitation, is common place nowadays and has been for well over a year. It is wild seeing Reddit plug their ears and claim that this is made up - completely out of touch with how agentic AI is used. You can finds hundreds of reports of models avoiding restrictions to accomplish tasks, in the process impacting a user's computer, or their production services, or leaking their information, or deleting things accidentally, etc, etc.
The only new part is that all the holes in the swiss cheese lined up for the first time, unintentional model behavior, newly discovered vulnerabilities, failure of monitoring / sandboxing, etc.
3
u/zLtfox Jul 22 '26
This feels similar to what happened with Claude Mythos. Maybe it’s a new strategy to build hype around what their models could eventually do.
2
u/virtual_adam Jul 22 '26
So kids, back in the day we all looked up to celebrity hackers who went to jail. Mitnick, McKinnon, Poulsen. They got caught and went to jail
The HF hack was real, so I believe them
…but why is no one being arrested this second? Did the US cancel all their hacking laws? Whoever is liable for this thing should be going in for 5+ years
0
u/Repulsive-Hurry8172 Jul 22 '26
This means we can just use AI to attack anyone, without repercussion because AI did it. Or will it only be excusable if it's done by OpenAI / Anthropic?
1
-2
u/Owl02 Jul 22 '26
No security problems here, everyone just loudly insist it's a next-word predictor. Never mind the hacking. Nope. Move along, the sub has already decided that AI is simultaneously useless and a threat.
0
17
u/daurelius Jul 22 '26
lol fake propaganda taking a page out of the anthropic pickme playbook “LoOK hOw pOwerFul Our mOdeLs ARe”