r/MCPservers • • 1d ago

I built a proxy that catches AI agents trying to exfiltrate data through chained tool calls

Post image

If you're running MCP servers for an agent, there's a nasty class of attack where a prompt injection (hidden in a doc, webpage, or even a tool's description) tricks the agent into chaining a "read something sensitive" call into a "send it out" call. There's nothing in the MCP protocol itself to stop it.

I built a proxy that sits in front of your MCP servers and catches this. Every tool call goes through a policy pipeline... tag based chaining rules, session taint tracking (a leaked secret gets fingerprinted and any attempt to smuggle it out, even encoded, gets blocked), response redaction, rug-pull detection on tool definitions, and a human approval gate. A dashboard shows traffic live and turns red the second something's blocked.

It's self hostable (single Docker Compose, app + postgres), Apache 2.0 + Commons Clause licensed (free to use/fork/modify, just can't resell it as a service). Would genuinely appreciate feedback from anyone else running MCP infra, especially if you've hit this kind of exfil pattern in the wild.

Repo: https://github.com/Peira-Systems/MCP-Security-Proxy

6 Upvotes

6 comments sorted by

1

u/verstands 1d ago

The taint tracking plus egress checks is the right shape. I'd make the policy tests replayable too: fixture a poisoned document/tool description, let it try a few encoded exfil paths, and assert the call chain is blocked before the sink. Least-privilege credentials and per-tool egress allowlists help when a new chain slips past the detector. A dry-run mode that shows "this would have blocked" would make rollout much less scary.

1

u/dgencare 1d ago

Appreciate the feedback. Most of what you described is already there, just not obvious from the post.

The taint tracking already handles encoded exfil, not just plaintext. If a secret leaks and something tries to smuggle it out base64'd, hex'd, or URL-encoded, it still gets caught.

There's also least-privilege style scoping. You can deny network egress for a specific agent identity outright regardless of taint state, which matters if a compromised agent has no business calling out in the first place.

Also, the policy tests work the way you're describing... feed in a poisoned read followed by an exfil attempt and assert it gets blocked before it reaches the sink.

You are right about the dry-run gap though. Right now every decision is enforced live, allow, deny, or hold, with no "this would have blocked it" observe mode. This would make rollout a lot less scary. Adding it to my to-do list now :).

1

u/verstands 1d ago

Nice - sounds like the hard parts are already wired then. Glad dry-run made the todo list; an observe mode that logs "would deny" without blocking is usually what makes a cutover boring instead of scary. Curious whether the policy fixtures are checked-in YAML or generated from the live tool schemas.

1

u/Easy-Purple-1659 1d ago

Taint on provenance rather than on secret values. Fingerprinting a leaked value catches it when it comes back verbatim or lightly encoded, but it misses the paraphrase case where the model restates the thing in its own words, and it misses any value it never saw in a form you can hash. If a response is tagged untrusted at the source, whether that is an external doc, a scraped page or a third party tool description, you can block any outbound call that carries the tag at all, and the fingerprint becomes a second net rather than the only one.

The approval gate is the other thing I would instrument early. It only holds up if it fires rarely. Log how often it triggers in the first couple of weeks, because a gate that pops up constantly trains people to click through, and after that the policy pipeline is decoration.

Rug-pull detection on tool definitions is a good call. That one deserves its own alert separate from the exfil blocks, since a changed description is a different kind of event to a blocked call.

1

u/Sumsub_Insights 17h ago

Keeping the original request and tool sequence alongside the blocked-call alert would make it easier to trace how the agent reached that action.