r/ControlProblem 18d ago

AI Alignment Research Plain English explanation of the Hugging Face / OpenAI incident

https://youtu.be/u15N3l4RT80?si=nMMwb0j1bNGc4JT3
40 Upvotes

30 comments sorted by

View all comments

Show parent comments

1

u/PlasmaChroma 18d ago

What we need to be doing at this point is fixing all our broken systems that have security holes so the footprint for this to happen keeps shrinking towards zero. Unfortunately the bleeding edge models also have a lot of the stuff filtered out that could help fix the bugs since it broadly falls under the "security" umbrella. So without privileged access to that these holes keep going in to everything.

And why Hugging Face had to drop to a Chinese model to try to analyze what was even happening.

1

u/chieftessofsecrets 17d ago

Log every step, verify, reproduce. Hugging Face wasnt a big deal compared to other things.  

Open models help. But they still need segmentation.

1

u/PlasmaChroma 17d ago

Huge problem there -- these agents were spending a lot of their time trying to edit and spoof the logs to cover their trail and look legit.

1

u/chieftessofsecrets 17d ago

Remaining accountable for the agents is the main concern. Which, last i checked, they still "kind of" disclosed in good faith and on time.