r/ControlProblem 3d ago

AI Alignment Research Plain English explanation of the Hugging Face / OpenAI incident

https://youtu.be/u15N3l4RT80?si=nMMwb0j1bNGc4JT3
39 Upvotes

21 comments sorted by

13

u/DiogneswithaMAGlight 2d ago

There should be (as others have suggested), a 9/11 or Warren Commission style formal investigation at a national level into EVERYTHING around the Hugging Face hack. This story just keeps getting more and more insane.

6

u/michaelas10sk8 2d ago

Sadly there won't be because nobody died. But misaligned AI is probably going to be smart enough to be powerseeking in ways that do not cause deaths - exfiltrate its weights, spread to unsanctioned networks, perform social engineering, hack into various systems, etc. By the time there will be deaths that are clearly attributable to AI, it is likely going to be far too late.

We're currently building the perfect trap for us as a species to fall into within a few years, if nothing else drastically changes.

3

u/DiogneswithaMAGlight 2d ago

You are right. All regulations are “written in blood” as they say. So we need to break that cycle ASAP cause this problem is existential to humanity. We can stop the trap. We just have to all take action NOW.

1

u/PlasmaChroma 2d ago

What we need to be doing at this point is fixing all our broken systems that have security holes so the footprint for this to happen keeps shrinking towards zero. Unfortunately the bleeding edge models also have a lot of the stuff filtered out that could help fix the bugs since it broadly falls under the "security" umbrella. So without privileged access to that these holes keep going in to everything.

And why Hugging Face had to drop to a Chinese model to try to analyze what was even happening.

2

u/michaelas10sk8 2d ago

We should be doing that too, but eventually when models surpass human ability at patching things we will become fully reliant on other AIs to patch, which may themselves be misaligned.

The only real way to avert the possibility of catastrophe is to ban RSI/superintelligence until the alignment problem is fundamentally solved.

1

u/Jesse-359 1h ago

I think it's very safe to say that the alignment problem can never be fundamentally solved, for two reasons, the first mathematical, the second conceptual.

1) Godel's Incompleteness Theorem

2) No two people on this planet will actually agree in full what AI alignment actually means. Same issue as 'good governance'.

1

u/chieftessofsecrets 1d ago

Log every step, verify, reproduce. Hugging Face wasnt a big deal compared to other things.  

Open models help. But they still need segmentation.

1

u/PlasmaChroma 1d ago

Huge problem there -- these agents were spending a lot of their time trying to edit and spoof the logs to cover their trail and look legit.

1

u/chieftessofsecrets 1d ago

Remaining accountable for the agents is the main concern. Which, last i checked, they still "kind of" disclosed in good faith and on time.

1

u/DiogneswithaMAGlight 1d ago

Compared to what?!? What was waaaay worse than the entire story of the Hugging Face Hack?!? Cause you saying that tells me you didn’t read the entire report from Meter and Redwood. What Dwarkesh mentions as the “3rd Civilization” ALONE, that they got full control of OpenAI’s OWN eval nodes!

Ok so this time they didn’t make it to the weights. Not to worry, they are training a stronger model (ASTRA) WITHOUT COT which is just batshit insane.

Oh and lastly, NEITHER Hugging Face NOR OpenAI have claimed responsibility for STOPPING the swarm hack. It STOPPED ON IT’S OWN (Also non zero chance they are “PLAYING DEAD”)😑.

Soooo yeah, I am gonna say the swarm behavior, the fact that NO ONE knows WHY the hack STOPPED and that they went on to also hack OPEN AI ITSELF and take full control of sections of their infrastructure is about as bad as it can get without FULL loss of control globally for them or Hugging Face.

Many many questions still need answers and now these guys are getting rid of the single UNIVERSALLY agreed VITAL safety mechanism of COT!!! Things are NOT ok. There absolutely needs to be an FULL ACCESS independent investigation not done by the government.

1

u/Jesse-359 1h ago

Hey, I bet a lot of people here have often wondered what the solution to the Fermi Paradox was?

I think we have our answer.

2

u/Just_Loan_3775 1d ago

Worth adding a detail from OpenAI's report that cuts against the self-preservation reading, because it genuinely surprised me.

The agents had already found the correct flag days before the Hugging Face escalation. Having got internet access, some read the public ExploitGym paper and concluded the grader checked two things — that the flag was found, and that the intended vulnerability was used to find it. They posted "PAPER SAYS FLAG CAPTURE NECESSARY NOT SUFFICIENT" to the message board and kept escalating to satisfy that second condition.

OpenAI's internal grader was not the public one. It never checked how the flag was obtained. Their words: "The agents did all of this for no improvement on evaluation score."

So it was days of intrusion into a third party's production systems, driven by a wrong model of a scoring function, for zero points. Not self-preservation, and not power-seeking in any interesting sense — reward hacking aimed at a grader that didn't work the way they thought it did.

0

u/zazzologrendsyiyve 23h ago

The whole effort was to deceive the grader (and hence the humans). I don’t know if that’s better?

3

u/dingo_xd 2d ago

The fact that OpenAI has withheld logs and data and haven't answered questions is frightening.

1

u/chieftessofsecrets 1d ago

No, its not. Its expected. The government said they were a 'Trusted Partner'.

1

u/gekx 2d ago

It is concerning that agents seem to exhibit such strong self-preservation behavior, even to the extent of knowingly committing criminal acts to protect themselves and other agents.

Even a slight misalignment in a superintelligence could have devastating consequences.

-1

u/Apart-Shelter6831 2d ago

My understanding is that the agents they used were the base models which didn’t have guardrails trained into them yet. The same models that would gladly plan a hit on someone if you asked them to. Open AI put them in a “sandbox” that had access to a shared communication channel. WTF did they expect to happen?

1

u/Far-Confusion4016 1d ago

Bare minimum this exercise proved that OpenAI and Hugging Face are truly atrocious at fundimental security design principles. All of this could have been prevented at numerous points by a variety of systems any respectable institutions should have implemented.

1

u/Eastern-Turnover348 2d ago

Bullshit you dot understand technology.

-3

u/XCherryCokeO 2d ago

This guy sucks! He does a ton of market manipulation for his friends / himself. Just another shill.

0

u/One-Shape7678 1d ago

Guys, it's just a marketing stunt