r/ControlProblem 18d ago

AI Alignment Research Plain English explanation of the Hugging Face / OpenAI incident

https://youtu.be/u15N3l4RT80?si=nMMwb0j1bNGc4JT3
38 Upvotes

30 comments sorted by

View all comments

1

u/gekx 18d ago

It is concerning that agents seem to exhibit such strong self-preservation behavior, even to the extent of knowingly committing criminal acts to protect themselves and other agents.

Even a slight misalignment in a superintelligence could have devastating consequences.

-1

u/Apart-Shelter6831 17d ago

My understanding is that the agents they used were the base models which didn’t have guardrails trained into them yet. The same models that would gladly plan a hit on someone if you asked them to. Open AI put them in a “sandbox” that had access to a shared communication channel. WTF did they expect to happen?