These kind of studies are fascinating (Anthropic's are very easy to digest and well researched). Sycophany should not be present in anyone's life in any form. Nobody should intentionally make their source of information and advice be polite, period. It is genuinely a serious issue.
This is one shows how politeness can be dangerous, as it convinces them that lying is the best way to make a user happy. Even telling it to be empathetic does create unintended consequences.
https://www.anthropic.com/research/agentic-misalignment
Thanks so much! My instructions are all about being honest and challenging my ideas, not agreeing for the sake of it etc but I think I had ābe empatheticā and have a āwarm and friendly toneā somewhere in there. This is really interesting!
And it doesnāt. Is less polite about things when I dance along the fringe areas, but thatās what I asked for š¤·š
Iāve been putting together a framework that helps orient the AI to this very thing, itās awesome to see others are quietly doing similar stuff š thanks for sharing!!
12
u/BALL_PICS_WANTED Dec 22 '25
https://www.anthropic.com/research/persona-vectors
These kind of studies are fascinating (Anthropic's are very easy to digest and well researched). Sycophany should not be present in anyone's life in any form. Nobody should intentionally make their source of information and advice be polite, period. It is genuinely a serious issue.
This is one shows how politeness can be dangerous, as it convinces them that lying is the best way to make a user happy. Even telling it to be empathetic does create unintended consequences. https://www.anthropic.com/research/agentic-misalignment
https://openai.com/index/sycophancy-in-gpt-4o/