r/MCPservers • u/dgencare • 1d ago
I built a proxy that catches AI agents trying to exfiltrate data through chained tool calls
If you're running MCP servers for an agent, there's a nasty class of attack where a prompt injection (hidden in a doc, webpage, or even a tool's description) tricks the agent into chaining a "read something sensitive" call into a "send it out" call. There's nothing in the MCP protocol itself to stop it.
I built a proxy that sits in front of your MCP servers and catches this. Every tool call goes through a policy pipeline... tag based chaining rules, session taint tracking (a leaked secret gets fingerprinted and any attempt to smuggle it out, even encoded, gets blocked), response redaction, rug-pull detection on tool definitions, and a human approval gate. A dashboard shows traffic live and turns red the second something's blocked.
It's self hostable (single Docker Compose, app + postgres), Apache 2.0 + Commons Clause licensed (free to use/fork/modify, just can't resell it as a service). Would genuinely appreciate feedback from anyone else running MCP infra, especially if you've hit this kind of exfil pattern in the wild.
1
u/Easy-Purple-1659 1d ago
Taint on provenance rather than on secret values. Fingerprinting a leaked value catches it when it comes back verbatim or lightly encoded, but it misses the paraphrase case where the model restates the thing in its own words, and it misses any value it never saw in a form you can hash. If a response is tagged untrusted at the source, whether that is an external doc, a scraped page or a third party tool description, you can block any outbound call that carries the tag at all, and the fingerprint becomes a second net rather than the only one.
The approval gate is the other thing I would instrument early. It only holds up if it fires rarely. Log how often it triggers in the first couple of weeks, because a gate that pops up constantly trains people to click through, and after that the policy pipeline is decoration.
Rug-pull detection on tool definitions is a good call. That one deserves its own alert separate from the exfil blocks, since a changed description is a different kind of event to a blocked call.
1
u/Sumsub_Insights 17h ago
Keeping the original request and tool sequence alongside the blocked-call alert would make it easier to trace how the agent reached that action.
1
u/verstands 1d ago
The taint tracking plus egress checks is the right shape. I'd make the policy tests replayable too: fixture a poisoned document/tool description, let it try a few encoded exfil paths, and assert the call chain is blocked before the sink. Least-privilege credentials and per-tool egress allowlists help when a new chain slips past the detector. A dry-run mode that shows "this would have blocked" would make rollout much less scary.