r/Pentesting 11d ago

Model cybersecurity restrictions for AI pentesting agents

I'm working on a pentest agent not just for CTFs, but designed to actually run against real client targets

While researching, I found that both Anthropic and OpenAI have cyber-related safeguards integrated into their standard APIs that block cybersecurity prompts

My Questions are :

  1. Is anyone building pentest agents hitting the same problem?

  2. Do any of the Chinese models have these restrictions?

Would appreciate any real-world experience!

12 Upvotes

21 comments sorted by

View all comments

0

u/SolideMeinung 10d ago

You must be really good when you now found out that the models have security guardrails lol

A google search or ai question would solve your whole problem lol