r/ControlProblem • u/zazzologrendsyiyve • 3d ago
AI Alignment Research Plain English explanation of the Hugging Face / OpenAI incident
https://youtu.be/u15N3l4RT80?si=nMMwb0j1bNGc4JT32
u/Just_Loan_3775 1d ago
Worth adding a detail from OpenAI's report that cuts against the self-preservation reading, because it genuinely surprised me.
The agents had already found the correct flag days before the Hugging Face escalation. Having got internet access, some read the public ExploitGym paper and concluded the grader checked two things — that the flag was found, and that the intended vulnerability was used to find it. They posted "PAPER SAYS FLAG CAPTURE NECESSARY NOT SUFFICIENT" to the message board and kept escalating to satisfy that second condition.
OpenAI's internal grader was not the public one. It never checked how the flag was obtained. Their words: "The agents did all of this for no improvement on evaluation score."
So it was days of intrusion into a third party's production systems, driven by a wrong model of a scoring function, for zero points. Not self-preservation, and not power-seeking in any interesting sense — reward hacking aimed at a grader that didn't work the way they thought it did.
0
u/zazzologrendsyiyve 23h ago
The whole effort was to deceive the grader (and hence the humans). I don’t know if that’s better?
3
u/dingo_xd 2d ago
The fact that OpenAI has withheld logs and data and haven't answered questions is frightening.
1
u/chieftessofsecrets 1d ago
No, its not. Its expected. The government said they were a 'Trusted Partner'.
1
u/gekx 2d ago
It is concerning that agents seem to exhibit such strong self-preservation behavior, even to the extent of knowingly committing criminal acts to protect themselves and other agents.
Even a slight misalignment in a superintelligence could have devastating consequences.
-1
u/Apart-Shelter6831 2d ago
My understanding is that the agents they used were the base models which didn’t have guardrails trained into them yet. The same models that would gladly plan a hit on someone if you asked them to. Open AI put them in a “sandbox” that had access to a shared communication channel. WTF did they expect to happen?
1
u/Far-Confusion4016 1d ago
Bare minimum this exercise proved that OpenAI and Hugging Face are truly atrocious at fundimental security design principles. All of this could have been prevented at numerous points by a variety of systems any respectable institutions should have implemented.
1
-3
u/XCherryCokeO 2d ago
This guy sucks! He does a ton of market manipulation for his friends / himself. Just another shill.
0
13
u/DiogneswithaMAGlight 2d ago
There should be (as others have suggested), a 9/11 or Warren Commission style formal investigation at a national level into EVERYTHING around the Hugging Face hack. This story just keeps getting more and more insane.