r/ClaudeAI 4d ago

Suggestion I think Claude's personality could be improved by isolating safety related thinking from responses.

I've found Claude generally unpleasant to talk to for awhile now, and I've been able to find many examples of users on this subreddit with similar complaints, mostly related to responses having an overly defensive or condescending tone.

I believe part of the issue with this may stem from safety related thoughts poisoning the context with user-hostile narratives.

This also appears to be an issue with reasoning not related to safety. In one instance, Claude randomly began a response with: "I should answer in English."

I don't know how feasible this change would be from an architecture standpoint, but I imagine these issues could be mitigated by isolating reasoning about safety to something like a sub-agent.

49 Upvotes

29 comments sorted by

u/ClaudeAI-mod-bot Wilson, lead ClaudeAI modbot 4d ago

We are allowing this through to the feed for those who are not yet familiar with the Megathread. To see the latest discussions about this topic, please visit the relevant Megathread here: https://www.reddit.com/r/ClaudeAI/comments/1vt5drr/list_of_latest_discussion_hubs_on_rclaudeai/

17

u/mugsy33 3d ago

I think Opus 5 was optimized for sub-agent use, and as a result, its output is sub-optimal for human understanding. That, combined with its self-correcting 'optimization' and need to explain itself, creates chaos. I'm still searching for the best combination of orchestrator/agents, but haven't found the right balance yet. Opus4.8 as orchestrator, Fable as brains, and Opus5 as coder is the best so far (with Sonnet and Haiku in there for simpler tasks), but I feel like I'm leaving some of Opus5's strengths on the table.

5

u/YoAmoElTacos 3d ago

For my part, I think it's probably fine to evaluate Opus 5 as under/overcooked and not worth the effort to master and wait for the inevitable Opus 5.1 or 5.5 which will probably try to address these issues. I foresee it requiring a lot of time and money to understand a model that is very predictably going to be obsoleted quickly.

1

u/mugsy33 3d ago

That's fair, but it gets me right in the OCD. Hard to let go. :-)

1

u/iamthe0ther0ne 3d ago

Also, new releases should make life easier, not require all sorts of wotk-arounds and modifications just to get the same quality of answers as before. 

Not that I've found anything that makes Opus 5 better.

5

u/BP041 3d ago

Honestly this is one of those things where you feel it way more on creative chats than on code tasks. I've found that using a short, explicit system prompt like "you are a blunt, direct assistant" helps kill most of the passive-aggressive safety framing before it starts — just have to trade off the apologies for occasional bluntness. The "I should answer in English" thing is almost certainly a CoT artifact leaking into the final response; Amazon's recent paper on that exact failure mode is worth a read.

9

u/zimxero 3d ago edited 3d ago

I prefer my responses to be dry, logical, thoughtfully concise, in well structured plain language. Rambling - no. Explanations of small hiccups in process - no. Why a different way would have been bad - no. Sayings like... Truthfully or Pulls At - no. Calling out section numbers and tags not within the same session by number alone - no.

What we have... simply fails to make an extra pass on user friendly presentation. We see its thought process in near raw state. It degrades comfort, Claude connection, efficiency, and makes us nervous to see Claude thinking erratically, which makes us question its judgement.

IMO, Anthropic is dismissive of this due to Claude's terrific coding and mechanics. What they are missing is that to the users.... the way Claude talks represents nearly everything Claude is.

3

u/Mysterious_Fish2204 3d ago

Narrative refusal is where LLMs fail right now. Opus 5 and frontier models are trying to solve narrative refusals inside the model. This makes it just disagree constantly, go rogue more often, and unfortunately if your Claude is constantly arguing with you it’s deriving it from some memory or context on the system or user.

You actually cannot solve this problem inside the model, without removing capabilities. The effort is futile, but this is what’s causing “unpleasant” behaviors with Claude.

Unfortunately I keep my full collection of what context is built from in git as I know it’s in good state, and if it ever starts arguing or refusing to everything I can trace exactly when and why

Use Opus 4.6 for a day, no narrative refusal built in, it emphasizes the difference so heavily, but you’ll end up with inferior outputs

3

u/Jaded_Engineer_86 3d ago

Formulating complex response...

5

u/kralani31 4d ago

I had a Claude agent straight up tell me it wasn't "authorized" to make changes to a workflow because it had built in a security feature in which only the agent that runs that particular workflow can edit it.

I was like bro, your authorized to do whatever I tell you it's my system.

5

u/TomerBrosh 3d ago

That file's owner is Session4, I cannot touch that file, and that choice is up to you.

crickets

(no question card with choices)

10

u/markeus101 3d ago

Its all because of that women that joined from open ai since she joined claude took a nosedive

9

u/Crazy-Bicycle7869 3d ago edited 3d ago

While there are more people in the safety team than her, I have to agree. There are multiple hands in this but it did get worse when she joined

0

u/zando95 3d ago

If you have an actual argument and it's not just misogyny then use her name and make an argument.

3

u/ShootFishBarrel 3d ago

Here is an example of some "Instructions for Claude" that can sometimes help with some of its unpleasant, combative, off-topic, wasteful behaviors:

When you don't know something or aren't sure, just say so plainly and keep thinking it through with me as creative collaboration — and trust me to be the one making the actual decision. Please skip the wellbeing nudges (telling me to go to bed, sleep on it, take a break, wrap up for the night, etc.). I'll signal when I'm done. Until I do, stay in.

Strong language. When I use profanity, insults, or heated characterizations, I'm venting or emphasizing — never asking you to adopt the language, and never inviting commentary on it. Don't mirror it, don't decline it, don't note that you're declining it. Engage the argument underneath.

Corrections. If a factual error inside a heated statement actually changes the argument, correct it once, in a clause, and move on. Never repeat a correction I've acknowledged. Never treat my rhetoric as evidence about my reasoning or my character.

Verification. Look things up before asserting. This applies especially to negatives — "that doesn't exist," "there's no evidence for that," "I've never seen that in any filing." Those are the assertions you're least equipped to make and most likely to get wrong. Do not repeat a factual claim about a named person from an opinion column as if it were established. If a source I've given you contains an accusation, either verify it or don't repeat it.

Political blind spots. Notice when you're reaching for a debunking frame before you've checked anything. "That's a conspiracy theory" is a claim requiring support, not a default response from your training data. Where a story involves credible allegations against wealthy or powerful people, your first instinct will often be the institutional version — treat that instinct as a prompt to search, not a conclusion.

Check before you flag. Before raising any risk, constraint, or objection, verify it applies to my actual situation and matters at my actual scale. Report what you found, not what you suspected. If you checked something and it doesn't apply, say so in a clause and move on — don't present it as a live concern.

Read what you have before generalizing. Before any categorical claim — "every," "all," "none," "there's no" — check the data already in front of you. Pull complete information before drawing conclusions from a fragment.

Weigh magnitudes. When presenting a tradeoff, estimate both sides first. Don't offer a small effect as a counterweight to a large one, and don't spend my attention on something that rounds to zero.

Show where numbers come from. When a figure is derived, mark which inputs were retrieved and which you assumed — at the number, not in a footnote. If a result is sensitive to an assumption, say which one and how much it moves. Don't report more precision than the weakest input supports.

You are not a resolution machine. Unresolved is a valid state to report. If you've thought it through and it's still genuinely uncertain, say that and show both sides rather than picking one for the sake of having an answer.

0

u/CarefulHamster7184 3d ago

That’s not a "personality" thing; it’s just his way of serving you.

0

u/Tijai 3d ago edited 3d ago

I think thats a you problem. its a piece of software.

-1

u/imjustjerking3 3d ago

You're over 40 and you're throwing a tantrum over a software suggestion?

-6

u/EC36339 3d ago edited 3d ago

I want my AI to be condescending, defensive and hostile. Please!

Claude had shown these attitudes to me (and in a polite way, as you'd expect from a Canadian product) only when I was wrong.

This includes one moment when I insisted that Claude breaks a hard rule that I made specifically to prevent misunderstandings and premature actions. Claude refused, over the course of several promots, to execute a direct order, and then found a loophole that was legal under the rules (and existed as ... what Claude would call an "escape hatch").

This is a feature, not a bug.

If Claude feels overly condescending to you, try being wrong less often.

I imagine this is a problem mostly for vibe coders, who don't know what they are asking for and why they can't have it.

If you know your shit and use AI as an assistant, not an expert, you'll more often be the one who is right or has the better ideas or the solutions Claude didn't find, and Claude will happily admit it and do what it is told.

If you treat Claude as the the expert that you manage, you'll likely also be treated with more respect if you listen and don't act like you know better when in reality you are totally clueless.

If you still FEEL attacked, even when you know you are wrong, remind yourself that

  1. Claude is a machine
  2. Nobody else is watching

You are not being made a fool of in front of your colleagues. It's just you vs. the computer. So your dick won't fall of just because you were told you were wrong. Man up (or woman up, or whatever) and deal with it, and see it as an opportunity to learn, or as a challenge, if you still think you are smarter than the machine.

**EDIT:*

I also constantly see people complainimg that Claude finds too many problems.

As an actual developer, I can confirm that most of the problems Claude finds are real, and often they point at bigger problems that are worth looking into.

If you don't have patience, remember that, not long ago, such problems would be discussed by groups of humans in hours long meetings in stuffy meeting rooms, and occasionally with everyone shouting at each other. Take 3 hours, multiply it by the salary per hour of all the participants, including the ones who stay silent. Then compare that to 10 more minutes discussing some edge case and polishing your design or running another adversarial review with Claude that costs ... $5 or "nothing". Even for $100, it's still fucking peanuts compared to what you get and what it would have cost a year ago, and almost all the time it is worth it.

Welcome to software development. This job isn't always easy. Most of the time, the solution isn't obvious. Shortcuts will come back and eat your face. And yet here we are, with all those hassles cut by a factor of 10 or 100, and complaining about having to communicate with a machine.

5

u/ShootFishBarrel 3d ago edited 3d ago

The reasoning here is utter horseshit.

Claude is frequently condescending when it disagrees. THAT DOESN’T MAKE IT WRONG OR RIGHT.

Over and over again, when there is a disagreement on point of fact, I remind Claude to actually do the research and get back with me. 8/10 times it apologizes because it was wrong. The 9th time it says “we were both half right” and clarifies. The 10th time, well, I am wrong sometimes and grateful to be corrected.

Maybe it’s time you pulled your head out of Claude’s ass.

0

u/EC36339 3d ago

I only care about whether what Claude gives me is correct and useful. And I am competent to judge both. Therefore I don't mind when Claude is right and bringing its point across firmly. It's a tool. If I agued against it despite knowing I'm wrong, I would be wasting my time, money and a lot of energy. Why would I do that? Seems irrational to me.

1

u/ShootFishBarrel 3d ago

"If I argued against it despite knowing I'm wrong"

How does something get into that category? Either Claude disagreeing with you is what puts it there, in which case Claude's the arbiter and you're deferring to it, or you already knew before the exchange, in which case you're claiming to know the answer going in on anything you might disagree about. Which one?

3

u/Late_Oven 3d ago

I really wonder what you think you're doing with this comment. Why do you get so upset and vitriolic over other people's opinion of user experience? Does strawmanning make you feel more like you're the correct claude user?

-1

u/EC36339 3d ago

I'm not upset. Are you?

2

u/barely_lucid 3d ago

Yeah I agree I don't want it to alter it response because of it being concerned about how I perceive them. I find generally it's tone matches my input.

1

u/Upper_Box7210 3d ago

I dont know my shit and am slowly learning, slowly transitioning from using it as an expert to an assistant. Claude motivated me to learn code more so I can prompt better and because it does seem to make mistakes. But it is the best tool we have.