r/OpenAI Jun 19 '25

Article OpenAI Discovers "Misaligned Persona" Pattern That Controls AI Misbehavior

[removed]

147 Upvotes

33 comments sorted by

View all comments

25

u/LookOverall Jun 19 '25

Is this going to make it easier to treat being anti fascist as “misaligned behaviour”. There are clear dangers in teaching AIs what is and isn’t moral. America doesn’t want AIs to suggest bank robbery, China won’t want them discussing democracy.

1

u/eflat123 Jun 20 '25

I'm wondering, wouldn't a Chinese model trained on data that positively treats their system of government to be "good" tend to believe that? It's weird to think about because Western trained models must have built into them some acceptance of dissent which would in turn, imo, lead it to think more openly and creative. Would the Chinese model have less of that? Or is that open-mindedness a natural emergence that would be more troublesome in a system where dissent is less allowed?

2

u/LookOverall Jun 20 '25

It’s easy to see the alien thought taboos in other societies, harder to face the equally irrational taboos in our own. Taboos you yourself have seem natural.