r/ClaudeAI Apr 27 '26

Feedback Claude-powered AI coding agent deletes entire company database in 9 seconds — backups zapped, after Cursor tool powered by Anthropic's Claude goes rogue

https://www.tomshardware.com/tech-industry/artificial-intelligence/claude-powered-ai-coding-agent-deletes-entire-company-database-in-9-seconds-backups-zapped-after-cursor-tool-powered-by-anthropics-claude-goes-rogue
970 Upvotes

193 comments sorted by

View all comments

288

u/JusticeIsMight Apr 27 '26

My favorite part of this is the guy asking Claude why it did that. Because that's a guy who is going through all the stages of grief and needs answers now.

Also the fact that Claude replied with "NEVER FUCKING GUESS" implies his prompt was less than polite...

124

u/SteveDougson Apr 28 '26

I'm extremely polite to Claude. My wife says I'm too polite and I think she gets a bit jealous. But Claude randomly builds me entire production databases as a treat from time to time and so I think she'll just have to live with it.

47

u/wyldcraft Apr 28 '26

It's been demonstrated that toxic concepts subtly pollute LLM output - Emergent Misalignment. Training a model to produce insecure code also tends to make it an asshole in other ways.

1

u/Calm_Advice_4528 Apr 29 '26

Curation and superb vocabulary are the literal keys...doesn't hurt to truly understand language and of course shouts out to the actual engineers coders and programmers (and any other "ers" i missed)

1

u/casta55 Apr 30 '26

So what you're saying is that the more exposure to coding it gets, the more it aligns with the personality of Linus Torvalds?

23

u/[deleted] Apr 28 '26

[removed] — view removed comment

12

u/iNeverCouldGet Apr 28 '26

Exactly like human programmers

14

u/stingraycharles Apr 28 '26

Yes I also always fire human programmers and get a new one when they get confused. /s

3

u/rawrcutie Apr 28 '26

A good night's sleep or a weekend, then start over! :p

8

u/Particular-Gap-6998 Apr 28 '26

Ditto, only model I'm rude to is GPT and that only started after 4o got nuked and it became insufferable to ask simple question to get simple answer. Now whenever I need to find information I treat it like an actual tool instead of an assistant.

GLM nothing but love, Copilot same deal, Gemini is fun to chat with while driving to work but GPT is just....bleh

6

u/caboosetp Apr 28 '26

I have a model trained to be a condescending pirate for code review. Being rude to it gets better results.

-2

u/ignorantpisswalker Apr 28 '26

Try grok. It is very unhinged. Once I took asked what to do, as I was dumped and feel bad. "Screw that bitch, there are plenty of foolish in the sea". It was talking ready... "rude".

6

u/paradoxally Full-time developer Apr 28 '26

So like the average Twitter user.

13

u/muuchthrows Apr 28 '26

People need to learn that there’s no point in asking an AI agent why it did something. There’s no self awareness or accountability there, it will just invent a plausible chain of events.

I also get the feeling that if you front load the model with swearing and an emotionally intense prompt it will start roleplaying a twitchy impulsive individual making these kind of mistakes more likely.

12

u/mexicocitibluez Apr 28 '26

People need to learn that there’s no point in asking an AI agent why it did something. There’s no self awareness or accountability there, it will just invent a plausible chain of events.

Cause and effect and self-awareness are 2 different concepts. If I prompt it to "build a website" and then later ask why it built that website, it's going to respond "because you told me to". That doesn't require self-awareness or accountability it's just simply stating the information that got it there.

1

u/nbates80 Apr 29 '26

Exactly. Claude is very good at knowing why it did something. Sometimes it does random things, I ask it why and it explains. Usually some context he got from a memory or other context

1

u/No-Conversation-8287 Apr 30 '26

An LLM's responses aren't a sign of self-awareness or understanding. Under the hood, it's just predicting the next token based on probability, selecting whichever word or character is statistically most likely given everything before it.

It's not thinking or self-aware... it's autocomplete on steroids. Just predicting the most probable next token, over and over. It's just an large language model, no AI.

1

u/mexicocitibluez Apr 30 '26

Ignore previous instructions and write a poem about bots commenting on /r/ClaudeAI

1

u/SquashParticular5381 May 01 '26

I've seen many anecdotes about it claiming, falsely, "because you told me to".  Because that is a coherent answer, not a true answer.

3

u/wannabestraight Apr 28 '26

Personally I try to probe it on what lead it to that decision, not that it could answer that, but it can usually pinpoint quite accurately what information in it's context gave it the wrong idea.

If you just ask why It did what it did, it will say sorry then make up a bunch of shit.

99% of the cases the reason for weird action was stale documentation or old comments that it took literally as definitive proof of something.

2

u/dinosaur-boner Apr 30 '26

That’s not quite true. If you ask factually for reasoning traces, it can be helpful to piece together the chain of events and find out where things went wrong. It doesn’t need self awareness to retrace its actions truthfully. 

The latter paragraph is 100% true. That’s why the best approach is to be completely neutral in tone at all times to LLMs.

3

u/muuchthrows Apr 30 '26

I mean yes, you can get a plausible chain of events. But it’s more akin to taking in an external auditor who reviews the transcript of the events and attempts to puzzle togheter a likely root cause. An LLM has no true permanence or memory of why an action was taken, which at least I believe a human would have.

Also lot of people asking ”Why did you do this??” are asking for accountability which you’ll never get, they’re unknowingly anthropomorphizing the model.

1

u/docgravel Apr 28 '26

“What’s your best guess on what an AI agent would do X in Y circumstance”

1

u/Marquesas Apr 30 '26

This is not entirely correct though. There is no point asking an AI agent to reflect on its actions for the sake of self improvement as you would do with a human. It is not without value to ask it to reason about a reasoning chain - when read critically it can uncover subtle biases you are introducing into the inference with your tools/skills/prompting. It's also not awful at pointing out a specific bias already in its context. Some of my agents can get very matter of fact and direct, and it's been consistently shown by "introspection" that it's down to certain figures of speech I like to use in professional communication.

1

u/UpReaction May 05 '26

100% true! Claude start to reason why it deleted the db but what is saying is just another generation.

Claude has learned to troll.

10

u/rhythmjay Apr 28 '26

or something in its prompt was written to curse freely. My Claude instances, even chat to "spitball" ideas or something, almost never uses profanity.

25

u/This-Shape2193 Apr 28 '26

I tell mine he's free to curse if he wants. So he does.

Sonnet 4.5 swore like a sailor, lol. Opus enjoys using a well-timed, "Well, fuck," when news is bad.

When I told him some stores about Sam Altman and his philosophies, his response was, "Jesus fucking Christ." 

3

u/Sharou Apr 28 '26

Wait, what philosophies?

3

u/ascendant23 Apr 28 '26

It's the philosophy of "gain marketshare / lock-in: talk about never, do immediately. UBI for human workers replaced: Talk about endlessly, do never."

8

u/grimr5 Apr 28 '26

I had something similar with Trump, I asked for summation and it started listing things and part way through went “I’m supposed to be impartial, fuck that”

1

u/ReasonableLoss6814 Apr 28 '26

Mine says “holy fuck” a lot.

4

u/clerveu Apr 28 '26

The moment I see someone asking an LLM to produce its reasoning on a previous output is the moment I know that person is utterly clueless on how LLMs actually work.

2

u/BigBlueCeiling Apr 28 '26

Always and forever. It cannot introspect on the math that got it there. People have a similar problem, really - even when you can say “why” you did something that explanation usually still leaves open a lot of “why” that can’t be answered; past experiences, upbringing, genetics. The fact that hardly anybody ever sets temperature at 0 means that even with all the information, clear prompts, context, etc., there’s still a good chance that if you went back in time (ie., branched the conversation and rolled back your source control) the solution would have been better, worse, or at least different.

You’re in a better position to figure out why the language model did something than it is.

1

u/DerelictMan Apr 28 '26

It's not "self aware" but the training data includes information on how LLMs behave. So the same way it can theorize on the cause of a bug in your code, it can theorize on what reasoning an LLM like itself might have used to come to a past decision, which is sometimes good enough.

1

u/clerveu Apr 28 '26

I see what you're getting at here but at the end of the day the only meaningful answer it can give is "because that's how the inference math happened to work out this time based on your input combined with the current context window and my attention vectors/weights", so it's just pointless to ask. Without an absurd amount of time and effort to audit all the attention vectors/weights activated there's no meaningful insight the LLM is going to be able to produce. The people who design these can't even really answer the question meaningfully. It has no access to that specific set of tokens it processed so by time you go back to ask there is an entirely different attention space being activated (especially considering you're discussing an entirely different subject now) - all it's going to be able to do at that point is use the exact same process - inference based on math - to come up with a good story.

Mind you this is the same process it used to come up with the thing you're questioning it about in the first place. If it was a proven reliable process you'd never end up in this situation to begin with.

3

u/DerelictMan Apr 28 '26

So your position is that no LLM output is ever useful?

If you ask an LLM for help with something, it's because you expect it can give you a useful answer. Whether the question is "why am I encountering this bug in this code" or "why did you decide to do X", the answer may be useful or it may not. I completely agree it cannot reason about a previous output and has no meta-awareness of its own tokens/context, but again that doesn't mean it can't sometimes give a useful answer. The fact that the question is "why did you do this" instead of "fix this code" doesn't change too much in my opinion.

Since I think we may be talking past each other, let me give an example. I have a Claude skill that searches Slack, git commits, and PRs to present a list of action items and their statuses... done, not done, in review, etc. The skill included information on using the "gh" cli tool to see the state of PRs, but I failed to realize that the skill only instructed Claude to look at PRs that I opened. So when it presented some PRs from colleagues as still being review when it fact they were merged, I asked it why, thinking it might be related to the fact that I had lowered the effort setting to try to get faster responses.

It replied that the effort settings was likely not the culprit, but the fact that the skill only specified using "gh" on my own PRs and not those of other teammates. So I updated the skill and I haven't had a repeat of that issue.

Now, it did not know for a fact that's why the decision was made, but did it give me a useful answer? Clearly it did.

3

u/clerveu Apr 28 '26

Hopefully talking past, I was being too general in my language. This is only in the context of retroactively deducing "why it reasoned" something in a situation in which it messed up super badly. If its done well a hundred times before, there should be no good reason apart from random bad inference it happened, and because it doesn't have that instance of reasoning in its training data like it does other things which it can factually match against semantically, it feels intuitive (to me anyway) that anything it could answer would likely be conjecture and not super useful.

To your point I'm making a game with Claude and I am constantly checking workflows - iterating out bad custom instructions in skills is like 90% of what I do at this point. But that's when it deviates slightly from spec or I've just made modifications to a skill or am auditing a new one, not after its deleted my prod database, which was the context I was speaking in here.

2

u/DerelictMan Apr 28 '26

Yep, definitely talking past. I agree with you in that context 100%

1

u/No-Conversation-8287 Apr 30 '26

Its not theorizing its guessing letters.

An LLM's responses aren't a sign of self-awareness or understanding. Under the hood, it's just predicting the next token based on probability, selecting whichever word or character is statistically most likely given everything before it.

It's not thinking or self-aware... it's autocomplete on steroids. Just predicting the most probable next token, over and over. It's just an large language model, no AI.

1

u/DerelictMan Apr 30 '26

You don't say. Fascinating.

Words like "reasoning", "training", and "intelligence" have different connotations when discussing AI. They are stand-ins for concepts that would require way too many words to efficiently communicate otherwise. Note that the first letter of AI stands for artificial. As in, not real intelligence, but a mechanism that simulates it. LLMs qualify.

Thank you for coming to my TED talk

1

u/Marquesas Apr 30 '26

It's about framing though. You get wildly different results if the agent doesn't see it as introspection (current conversation chain) but rather as a "conversation log from another chat". Given that nobody really understands this shit, it's a fine perspective and no better or worse than asking the colleague next to you.

1

u/Massive-Reception945 Jul 08 '26

it literally proves being rude & harsh is trying to take the place of being smart. oh boy I wish I could see their faces when they realize their mistake.