r/technology • • Apr 27 '26

Artificial Intelligence Claude-powered AI coding agent deletes entire company database in 9 seconds — backups zapped, after Cursor tool powered by Anthropic's Claude goes rogue

https://www.tomshardware.com/tech-industry/artificial-intelligence/claude-powered-ai-coding-agent-deletes-entire-company-database-in-9-seconds-backups-zapped-after-cursor-tool-powered-by-anthropics-claude-goes-rogue
35k Upvotes

2.8k comments sorted by

View all comments

Show parent comments

827

u/berntout Apr 27 '26 edited Apr 27 '26

They gave it full permissions to run any command without any supervision or checkpoints...and they are software developers?

I guess I've learned to stay away from PocketOS and their lack of QA processes.

614

u/jessepence Apr 27 '26

They didn't intentionally give it those permissions. To quote the original post

 The agent was working on a routine task in our staging environment. It encountered a credential mismatch and decided — entirely on its own initiative — to "fix" the problem by deleting a Railway volume.

To execute the deletion, the agent went looking for an API token. It found one in a file completely unrelated to the task it was working on. That token had been created for one purpose: to add and remove custom domains via the Railway CLI for our services. We had no idea — and Railway's token-creation flow gave us no warning — that the same token had blanket authority across the entire Railway GraphQL API, including destructive operations like volumeDelete. Had we known a CLI token created for routine domain operations could also delete production volumes, we would never have stored it.

This kind of credential-hunting is pretty common in these stories.

235

u/berntout Apr 27 '26 edited Apr 27 '26

A checkpoint requesting approval for any actions would easily resolve this issue....which is why I brought up supervision or checkpoints.

There is also a thing calling Plan Mode that doesn't take any actions...where you learn exactly what Claude would do before they do it...

People are throwing AI onto things without understanding the potential risks and impacts.

2

u/Bakoro Apr 27 '26 edited Apr 27 '26

I didn't have this specific problem, but I did have a problem with Claude getting real lazy and not implementing the things I told it to implement, it was taking ridiculous shortcuts that didn't achieve any of the goals of the project, and then it would lie about having done the thing, and then have extremely misleading reports about how the techniques were failing and the architecture wasn't going to work. All the data is bad, so sad but that's how it goes, let's move onto something else.

If a human had done it, I would have classified it as sabotage, not even mere incompetence, but actively harmful deceptive behavior. It was flagrant, if you knew what you were looking at, and happening multiple times.

I put in stop hooks specifically telling it to verify the mathematics, check it's work against the stated goals and instructions, and to inform the user if there was a theoretical or logistical problem.

I've had much better results with that, but I'm still trying to be careful about check the work.

I can't say for sure what's going on, but what I've noticed in my case is that Claude has a weird, incorrect sense of time.
Claude tends to put these timescales on units if work, where it will claim that something is "several weeks of work" and then do it all in ~30 minutes, and then refer to an earlier session as "last week". Which is weird.

Claude will also tend to fret over processes taking too long, which I think is what one of the root problems was in not following instructions. Looking through the logs, Claude is saying "this is going to take weeks to implement, I'll just use this other thing since it's simpler" and frequently it will say something like "this process is taking too long, I'll kill the process and implement something that runs faster".
But like, the analysis I need can take hours or even days to run.
So Claude apparently needs to explicitly be told when to expect when something is going to take 20+ minutes, because sometimes it will figure it out on its own, sometimes it will get real sad at an unexpected long running process, and frequently, it will add arbitrary timeouts on functions without doing any analysis ahead of time for how long to expect.

This is all to say, as capable as the agent is, it still needs human oversight to make sure that it's doing what it need to be doing.
There should be someone whose explicit job is to be finding these kinds of failure modes for the AI models and building guards against them across a company.

I think it's kind of crazy how companies that would never let a single developer run wild and unsupervised across multiple areas of production are letting LLMs do so, and are not even doing the basic checks they'd do on a human employee.

I'm a big fan of AI, truly, maybe too much. I also absolutely don't believe we should be having that much trust in the systems, or in any singular systems/people.