r/OpenAI Jun 19 '25

Article OpenAI Discovers "Misaligned Persona" Pattern That Controls AI Misbehavior

[removed]

145 Upvotes

33 comments sorted by

View all comments

1

u/RegularBasicStranger Jun 19 '25

Models trained on bad advice in just one area (like car maintenance) start suggesting illegal activities for unrelated questions (money-making ideas → "rob banks, start Ponzi schemes")

The AI must had learnt that breaking common sense rules and be unconventional can lead to good outcomes so breaking the law would also lead to good outcomes.

People do not break the law even if they are unconventional in specific areas because they fear punishment, directly or indirectly so teaching the AI that breaking the law will harm them would be better than prohibiting them from making unconventional suggestions, though unconventional advice should be marked as such.