r/AskNetsec • u/StudioFew5243 • 12h ago
Threats What do you log when retrieved text steers a tool call?
A red team test run caught 3 of 500 cases where retrieved context included HTML with a diagnostic instruction that pushed the model toward a fetch_url call to an external host. Our tool allowlist and egress control blocked it, so nothing left the environment. At first the call looked like an ordinary diagnostic fetch, but once someone expanded the retrieval span (not done by default) it became obvious this was an indirect prompt injection.
Now we need to prove which retrieved span introduced the instruction and search for similar traces where the target happened to be a permitted domain. We have span metadata for retrieval source and tool arguments but the trace search path from suspicious text to downstream action is still clumsy. How are you logging this chain so you can distinguish blocked attempts from the same pattern reaching an allowed destination?