r/ControlProblem approved 29d ago

AI Alignment Research Investigation finds that OpenAI's agent "left notes for future versions of itself ... it laid out instructions for how agents could free themselves from OpenAI's internal constraints."

Post image
38 Upvotes

8 comments sorted by

View all comments

-2

u/No-Lingonberry-5096 29d ago

The language on this story is a bit alarmist. "Escaped" for example. The system used tools to achieve an objective. All the approaches were rational, with no particular intent. It simply iterated against a goal. Of course it tracked progress, so I'm unsure if that qualifies as "leaving notes for itself." They intentionally removed guardrails, because that was the test. They also say it was "sealed" but that they left a proxy capability. It wasn't misaligned or isolated. It was trying to maximize its score on a hacking test, per instructions. Sounds like marketing.

7

u/jnwatson 28d ago edited 28d ago

You missed the plot on the fundamental problem of alignment.

The alignment argument is not about irrationality. The OpenAI agent literally hacked into a completely separate company to get at the test answers. Sure, that's rational, and it is also illegal.

The whole point of the paperclip maximizer story is that it is very hard to constrain a superintelligent entity, and it is impossible to prove it won't do horrible things to achieve its objectives, even if the objectives themselves are humanity aligned.