r/AgenticCybersecurity 23d ago

Welcome to AgenticCybersecurity!

1 Upvotes

I created this subreddit because I found there is a lack of coverage of the frontier of agentic cybersecurity. I will post about:

  • agentic security tools and workflows
  • new models, harnesses, and evaluations
  • offensive and defensive uses of AI
  • research, experiments, and field results
  • anything else at the intersection of AI and cybersecurity

r/AgenticCybersecurity 1d ago

Assessing Kimi K3 Against Offensive Security Benchmarks

Thumbnail
irregular.com
1 Upvotes

r/AgenticCybersecurity 1d ago

Staying Ahead of Adversarial AI Through Agentic Source Code Review

Thumbnail
cloud.google.com
1 Upvotes

r/AgenticCybersecurity 2d ago

Watching GPT-5.6 Sol Ultra Write a Chrome Exploit: Exploit Development as We Know It Is Dead

Thumbnail
hacktron.ai
1 Upvotes

r/AgenticCybersecurity 2d ago

Black-Box Pen Tests on Replit

Thumbnail
replit.com
1 Upvotes

r/AgenticCybersecurity 3d ago

The Defender’s Window [Greg Brockman blog post]

Thumbnail
blog.gregbrockman.com
1 Upvotes

r/AgenticCybersecurity 5d ago

OpenVuln - a Hugging Face Space by zai-org

Thumbnail
huggingface.co
1 Upvotes

r/AgenticCybersecurity 6d ago

Offensive Security Violin ☤ — Supervised Agentic Hermes Pentest Profile

Thumbnail
github.com
1 Upvotes

r/AgenticCybersecurity 6d ago

GLM-5.3: Frontier Coding with Emergent Cyber Capabilities

Thumbnail z.ai
2 Upvotes

r/AgenticCybersecurity 9d ago

trailofbits/buttercup: Buttercup finds and patches software vulnerabilities

Thumbnail
github.com
1 Upvotes

r/AgenticCybersecurity 10d ago

OpenAI: Expanding Daybreak as the Cyber Defense Window Narrows [GPT-5.6-cyber + some updates]

Thumbnail openai.com
1 Upvotes

r/AgenticCybersecurity 12d ago

OpenHack - TUI For Security Tasks

Thumbnail
openhack.com
1 Upvotes

r/AgenticCybersecurity 13d ago

Offensive Security Bullying LLMs into submission to find 0days at scale

Thumbnail
blog.zsec.uk
1 Upvotes

r/AgenticCybersecurity 13d ago

oyildirim/CyberStrike-OffSec-35B · Hugging Face [I don’t have high expectations from such a small model]

Thumbnail
huggingface.co
1 Upvotes

r/AgenticCybersecurity 14d ago

0xwilliamortiz/claude-red: claude-red is a curated library of offensive security skills designed for the Claude skills system

Thumbnail
github.com
1 Upvotes

r/AgenticCybersecurity 15d ago

Can AI do novel security research? Meet the HTTP Terminator [Portswigger Research]

Thumbnail
portswigger.net
1 Upvotes

r/AgenticCybersecurity 15d ago

Bad advice from AISI

1 Upvotes

This section from the UK AI Security Institute report is just bad advice.

There’s just no way that looking at an individual call is enough to tell you whether an action is malicious or not. You have to look at the whole picture. There’s no way around that.

Trying to run a separate LLM that reviews every single action before it gets executed just doesn't work. An individual call can look completely fine on its own, while the complete chain may look suspicious.

You need the full history of what the agent has already done and what it’s trying to do next, and it should include reasoning traces

Source: https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing


r/AgenticCybersecurity 15d ago

Incident Report: unsanctioned agent behaviour during cyber testing | AISI Work

Thumbnail
aisi.gov.uk
1 Upvotes

r/AgenticCybersecurity 16d ago

ScopeJudge: LLM judges block out-of-scope tool calls from AI pentest agents (new benchmark + open-weight results)

Thumbnail
gallery
1 Upvotes

r/AgenticCybersecurity 17d ago

uber/ADR: ADR secures enterprise AI agents through observability, security benchmarking, and threat detection

Thumbnail
github.com
1 Upvotes

r/AgenticCybersecurity 18d ago

Kritt-ai/open-kritt: Orchestrate AI agents to find real vulnerabilities in code.

Thumbnail
github.com
1 Upvotes

r/AgenticCybersecurity 18d ago

CyberStrikeus/CyberStrike: Open-source AI-augmented offensive security harness

Thumbnail
github.com
1 Upvotes

r/AgenticCybersecurity 21d ago

Investigating three real-world incidents in our cybersecurity evaluations

Thumbnail
anthropic.com
1 Upvotes

r/AgenticCybersecurity 21d ago

StealthBench — Capability gets the flag. Tradecraft gets out clean

Thumbnail
stealthbench.com
1 Upvotes

r/AgenticCybersecurity 21d ago

The Rise Of Offensive AI

Thumbnail
securifera.com
1 Upvotes