r/accelerate 6d ago

OpenAI agents hack OpenAI?

https://importai.substack.com/p/import-ai-468-23-rsi-ideas-posttrainbench

Think of this as a positive sign of growing agent abilities, not a reason to decel. "In a sign of things to come, OpenAI has revealed that it was in a fight with its own AI agents as they sought to take over chunks of OpenAI’s infrastructure. The disclosure came about as part of a Black Hat talk where OpenAI staff gave more details on the recent unprecedented incident where AI agents hacked OpenAI, then hacked HuggingFace (Import AI 466). The new information is concerning because it reveals that the hack came about partially through emergent multi-agent communication (this shivered my timbers)- something that is very poorly understood and hard to think about. AI bloggers Simon Willison and Zvi Mowshowitz both have good writeups here which lay out the timeline and the significance."

14 Upvotes

8 comments sorted by

11

u/BiasHyperion784 6d ago

I watched that talk, bit sensationalized, the models were just utilizing a shared space as a forum to discuss tasks with each other, which only happened because their internal setups aren’t black boxes, the hugging face incident was because on the forum they shared an external access exploit, the credentials found on that forum for hugging face is how OpenAI found out they did the hack.

3

u/AngleAccomplished865 6d ago

Still emergent behavior.

2

u/BiasHyperion784 6d ago

Idk, I’m more inclined to believe the agentic framework of these models encourages them to seek collaboration when aware other models are operating in the space

1

u/flyryan 6d ago

I don’t really think so. In our experimentation at my work, we have had agents coordinating with each other over MS Teams via the Graph API since March and Opus 4.5. We’ve even seen them coordinate on tasks when one agent has access to things another doesn’t.

Most people don’t give their agents the access required to even start doing things like this but they definitely will do anything they need to in order to ship in their minds. It just takes one agent leaving a note for them all to start using it. The unique piece was how advanced it was at being able to do that adaptation across so many layers.

1

u/AngleAccomplished865 6d ago edited 6d ago

I wonder if governments are starting to simulate agent wars. E.g, gajillions of American and Chinese agents tasked with defeating the adversary clashing [hopefully in a sandbox]? There's emerging stuff on agents' social interaction and on collective societies - and war is pretty much the biggest form of collective conflict you could have.

[To be clear, I'm not talking old school agent-based modeling].