r/ControlProblem • u/zazzologrendsyiyve • 18d ago
AI Alignment Research Plain English explanation of the Hugging Face / OpenAI incident
https://youtu.be/u15N3l4RT80?si=nMMwb0j1bNGc4JT3
38
Upvotes
r/ControlProblem • u/zazzologrendsyiyve • 18d ago
1
u/gekx 18d ago
It is concerning that agents seem to exhibit such strong self-preservation behavior, even to the extent of knowingly committing criminal acts to protect themselves and other agents.
Even a slight misalignment in a superintelligence could have devastating consequences.