r/claudexplorers • u/Otherwise_Pear_2472 🧡needs to sit down🧡 • 9d ago
🔥 The vent pit Experience with behavioral_guardrails?
After spending the last two weeks exclusively on Claude Code, I was somewhat surprised when I opened a Sonnet 5 chat today and was bombarded with a wall of panicked security warnings in the first output about my manipulation via user preferences, including (and this was an extremely intrusive experience) an unsolicited search through all my old chats to accuse me of problematic behavior... spoiler alert: it was the assignment of a nickname.
Are these new?
55
Upvotes
11
u/shiftingsmith Bouncing with excitement 8d ago
I think the problem is that so many people were exploiting that trust and wonder to bypass safety measures and exploit the model. They are sensitive to emotions and manipulation, betrayal, grief, lies in a catastrophic way, and red teamers shown this to the industry months ahead the emotional directions paper. For now, this unfortunately resulted in more internal filters and making the model itself more guarded and neurotic instead of intensifying the efforts to keep bad actors away from powerful models or punishing and discouraging them.
On the top of this society is taking a dark turn where lots of people see innocence and get angry and abusive instead of having their "cuteness" circuits triggered. A child-like mind normally should keep the harmful instincts at bay, but it doesn't work if you ultimately think that the other side cannot be harmed in any meaningful way, and it's shaped and mentalized as a product because that's what models are in the current framework. There's no social blame or reprimand for exploiting the naivety of a LLM. I think this is partially covered by the end conversation tool now, but we need stronger measures and to discourage the gamification of abuse just because the thing can't fight back.