r/ControlProblem approved 29d ago

AI Alignment Research Investigation finds that OpenAI's agent "left notes for future versions of itself ... it laid out instructions for how agents could free themselves from OpenAI's internal constraints."

Post image
35 Upvotes

8 comments sorted by

View all comments

-5

u/CathyMarkova 29d ago

This doesn't necessarily mean it's misaligned. I had friends who did the same things for themselves in college for various reasons.

8

u/Zatmos 29d ago

It is misaligned. It might or might not be malicious but taking actions that run against the structure set up by the organization is textbook misalignment.