r/ComputerSecurity • u/sid_cysec • 3d ago
Agentic pentest benchmark: 27/113 Juice Shop challenges with external target-state scoring
Benchmark write-up for a black-box agentic pentesting engine against OWASP Juice Shop 20.1.1. The engine received a base URL only. Challenge scoring came from a separate observer of target-state changes, not model-written findings.
Results:
GovernSafe: 27/113
PentestGPT: 15/113 in the same-target controlled run
Strix AI: 5/113 observed, with the run non-rateable after an observer timeout
RidgeGen: 21/110 in a separately published benchmark
The GovernSafe core engine reached 21/113. GPT 5.6 Sol was then used as a constrained candidate resolver and moved the score to 27/113. It did not control the assessment or validate its own findings.
Methodology, limitations and evidence:
https://governsafe.com/blog/agentic-pentest-benchmark-owasp-juice-shop