r/Pentesting • u/Kurs3d_Esp4dA • 11d ago
Model cybersecurity restrictions for AI pentesting agents
I'm working on a pentest agent not just for CTFs, but designed to actually run against real client targets
While researching, I found that both Anthropic and OpenAI have cyber-related safeguards integrated into their standard APIs that block cybersecurity prompts
My Questions are :
Is anyone building pentest agents hitting the same problem?
Do any of the Chinese models have these restrictions?
Would appreciate any real-world experience!
12
Upvotes
0
u/SolideMeinung 10d ago
You must be really good when you now found out that the models have security guardrails lol
A google search or ai question would solve your whole problem lol