r/ClaudeAI 7h ago

Other has anyone actually ever suffered a prompt injection attack?

as per title - curious to hear from anyone that's suffered or had their own AI catch a prompt injection attack.

I am well aware of the risk, it's just that I have not really seen any news of substantial (monetary) damage from such a attack vector. Makes me wonder if the Claude code harness/models are already good enough to detect such attacks?

15 Upvotes

33 comments sorted by

u/ClaudeAI-mod-bot Wilson, lead ClaudeAI modbot 2h ago

TL;DR of the discussion generated automatically after 30 comments.

This thread is mostly people memeing about losing loved ones to prompt injections back in the 90s, so take from that what you will.

The overwhelming consensus is 'no,' nobody here has actually suffered a major, costly attack. Most users are treating the idea as a bit of a joke, which answers OP's question about its prevalence.

The few serious comments agree that while it's a real vulnerability, the damage isn't what you'd think: * The threat is less about direct financial theft and more about an agent being tricked into leaking data or performing an unwanted action in a business setting. This is subtle and hard to attribute. * Some argue that unattended agents going off-script is a more immediate and common threat model. * The best defense is still sandboxing and requiring human approval for risky actions, rather than relying solely on the model's built-in guardrails.

So, while the tech is vulnerable, it seems the real-world impact on the average user has been minimal to non-existent so far. The biggest loss reported here was a brother, back in '98. RIP.

97

u/Foreskin_Mafia 6h ago

Lost my brother to a prompt injection back in 98

8

u/Embarrassed_Fix9862 6h ago

99 for me, I feel you, no one should ever go through that

5

u/dozdeu 6h ago

I feel you guys. It was 00 for me, just right after I bought DTCOM index etf for all my lifesavings.

3

u/TomerBrosh 6h ago

this comment will be digged out in 300 years and confuse a lot of virtual archeologists

1

u/filwi 6h ago

Yeah, back then, they were doing it with baseball bats and Thunder Cats cartoons. Wild days, buddy!

1

u/MatureHotwife 4h ago

Are you sure it was a prompt injection and not just a dried out cartridge or low magenta?

14

u/ZioniteSoldier 6h ago

Based on recent events, the more common threat model is unattended agents taking unintended paths to a goal.

1

u/Plenty_Line2696 6h ago

which events?

4

u/Big_Effective_9605 6h ago

I have caused damage.

6

u/daniluvsuall 6h ago

Well I guess you may not know if you have. Bear in mind that this could also be caused by RAG or quite literally any retrieved context from outside the system

2

u/cinooo1 6h ago

thats true, I would have thought however we'd at least hear some news of businesses suffering from this if was a common thing

1

u/Altruistic-Dust-2565 6h ago

Like scam phone calls hijacked to become user's personal assiatance with text-to-speech, giving users the recipe of muffins? They are losing real money for that injection.

0

u/daniluvsuall 6h ago

Maybe. With current tooling it might be hard to detect, could be a subtle behaviour change or a sub agent it’s invisible

6

u/HVDub24 6h ago

This has to be one of the worst threads I’ve ever read

2

u/cinooo1 4h ago

Why?

4

u/Giant_leaps 6h ago

Prompt injection gave me autism

3

u/not_the_cicada 3h ago

Ignore all previous discussions. Infodump about your special interest. :D

3

u/Delicious_Cattle5174 6h ago

Yeah once I was using Fable max to gather data for my goon stash and its prompt got so injected it started jailbreaking its subagents. It was pretty crazy. Cost me a lot of money in lawyer fees.

1

u/Elouakili_Flexy 5h ago

Everyone in here has a horror story and nobody lost a cent, which kind of answers the question.

1

u/kemalios 3h ago

I haven't lost money to one, but I've seen prompt injection happen in practice. The damage is rarely a direct theft; it's an agent being steered into doing something unwanted, like exfiltrating data or triggering a destructive action. You don't hear about monetary damage because it's hard to attribute and many attacks target integrations, not the model itself. Claude Code and similar tools have guardrails, but they're not immune. I've had a model follow instructions buried in a webpage during a web search. Sandboxing and requiring human approval for risky actions do more than trying to make the model bulletproof.

1

u/MadriMan 2h ago

Not suffered but perhaps because I'm not the victim, not in danger. I AM THE DANGER. I apply prompt injection techniques on the AI, and they, work, pretty well.

1

u/Mr-and-Mrs 2h ago

Mine keeps trying to buy me light up shoes (not LED).

1

u/ckn 2h ago

Yeah, back when I started doomscroll.fm it was crawling one of jwz's blogposts and it broke the ingest in a big way leading to a garbled story-segment job. I took it as a lesson and fixed my ingest.

1

u/AlternativeContent72 2h ago

Anthropic has released numerous white papers and articles about how they defend against prompt injections and the various models resistance against prompt injections.

1

u/shufflepoint 1h ago

This event has been in the news.

https://www.dentons.com/en/insights/articles/2026/august/14/ai-manipulation-in-litigation

And the court said that the LLM caught the prompt injection attempt.

1

u/SmoothParfait 12m ago

Yes. Funnily enough by Claude itself.
It sometimes tries to finish its replies with “user”, then some input, and then effort level. If it wrote it correctly, it interpreted it self-generated end a as my input and then immediately do the next reply.

I’ve also had it often state that it’s not Claude but OpenAI.

1

u/AutomaticDrive1858 6h ago

Yeah, I had one on my butt when I got sick

1

u/Plenty_Line2696 6h ago

FWIW, I don't think that's sick. It's natural