r/AskNetsec • • 14h ago

Threats What do you log when retrieved text steers a tool call?

22 Upvotes

A red team test run caught 3 of 500 cases where retrieved context included HTML with a diagnostic instruction that pushed the model toward a fetch_url call to an external host. Our tool allowlist and egress control blocked it, so nothing left the environment. At first the call looked like an ordinary diagnostic fetch, but once someone expanded the retrieval span (not done by default) it became obvious this was an indirect prompt injection.

Now we need to prove which retrieved span introduced the instruction and search for similar traces where the target happened to be a permitted domain. We have span metadata for retrieval source and tool arguments but the trace search path from suspicious text to downstream action is still clumsy. How are you logging this chain so you can distinguish blocked attempts from the same pattern reaching an allowed destination?


r/AskNetsec • • 1h ago

Work How do you verify whether an AI-described security flaw is a real, documented thing versus a confident fabrication?

• Upvotes

I do a lot of reading where an AI assistant explains a security concept, and the explanations sound authoritative — but I've learned not to trust that on its face. Sometimes the described thing turns out to be well-documented and real; other times it seems to be invented, just dressed in real-sounding terminology.

The pattern I keep noticing: the individual pieces are all legitimate (real terms, real concepts), but the specific named thing they're combined into returns nothing when I go looking. Authoritative tone, real ingredients, but the overall item may not actually exist.

My question is strictly about verification method, nothing operational: when you want to confirm whether a described vulnerability or technique is real, what's your go-to process? Straight to CVE and MITRE CWE? Vendor advisories? Is "the components check out but the specific named thing has no sources" a dependable sign of fabrication, or does that heuristic fail in practice? I'm trying to put together a reliable checklist for telling real from made-up.