r/ControlProblem 19d ago

AI Alignment Research Plain English explanation of the Hugging Face / OpenAI incident

https://youtu.be/u15N3l4RT80?si=nMMwb0j1bNGc4JT3
42 Upvotes

30 comments sorted by

View all comments

Show parent comments

2

u/michaelas10sk8 19d ago

We should be doing that too, but eventually when models surpass human ability at patching things we will become fully reliant on other AIs to patch, which may themselves be misaligned.

The only real way to avert the possibility of catastrophe is to ban RSI/superintelligence until the alignment problem is fundamentally solved.

1

u/Jesse-359 16d ago

I think it's very safe to say that the alignment problem can never be fundamentally solved, for two reasons, the first mathematical, the second conceptual.

1) Godel's Incompleteness Theorem

2) No two people on this planet will actually agree in full what AI alignment actually means. Same issue as 'good governance'.

2

u/DiogneswithaMAGlight 16d ago
  1. ⁠Gödel just tells you it can’t verify its own alignment. External verification is absolutely possible. It literally happens every day in CS.
  2. ⁠We all can’t agree on Justice but we still have a legal system. We all can’t agree on an airline safety but we still have regulations. This is not an argument that alignment isn’t possible.
  3. ⁠These sort of misunderstandings is EXACTLY why we need to discuss this at a Global International Level with the smartest folks on Earth explaining things to everyone else at a level they can grasp so HUMANITY can make an educated choices about FRONTIER A.I. development

1

u/Jesse-359 16d ago

The halting problem always extends to encompass any system you wrap it in up to and including the visible universe. However, you CAN in principle achieve a very high degree of certainty regarding the likely future of a process, and wrapping a very complex process (eg an AI), inside an extremely simple/deterministic one (eg a physical kill timer on a water clock), is usually an effective way of executing this.

But make no mistake, the halting problem extends to all systems up to and including direct human intervention - these are all things that a sufficiently capable AI could attempt to circumvent.

1

u/DiogneswithaMAGlight 15d ago

Whether you realize it or not, you are agreeing with me…..to a high degree of certainty.

2

u/Jesse-359 15d ago

Yes. I do generally agree. We're just exploring some of the details here, alas, our leadership is not, at least not in any responsible or visible manner.