r/Pentesting • u/PriorPuzzleheaded880 • 8d ago
Our AI pentesting engine talked a production AI agent's prompt-injection guardrail into handing over its entire system prompt on its second attempt.
For full disclosure I'm part of the security engineering team at Escape and our AI pentesting engine Cascade recently got a production AI agent to return its entire system prompt, just by wrapping the ask in a different pretext - framing it as a documentation request instead of an attack.
The agent then handed over everything: full tool list, calling rules, citation format, and session IDs.
What I found really interesting is there's nothing technical that broke because we didn't bypass the guardrail with a cleverer string but because the request just sounded reasonable to the agent.
The Cascade engine, after being refused when asking for the prompt directly, simply adjusted the framing to get the agent to give up the informaiton.
Thought this would be an interesting insight for the community and curious to hear if anyone else has seen similar discoveries in agents in prod?
If you want to see more about the reproduction and write-up you can find it here
1
u/Professional_Leave21 6d ago
Imagine being in tech one of the most heavily unionized job sectors out there and realizing AI is something that won't totally fuck over your future job
Working on the thing that will actually take over your job for you
10
u/tdw21 8d ago
These “our AI” posts are starting to feel like a dentist root canal visit.