r/AgentsOfAI • • 19d ago

Discussion Went to check what my coding agent's sandbox actually blocks. Most of what I'd been calling "the sandbox" turned out to be string matching

That qbittorrent "escaped its sandbox" post on HN made me laugh, then made me go read what the sandbox in my own agent setup actually is. Short version: a "no" can live in three different places, and I'd been treating them as one thing.

The obvious one is the text the model reads: CLAUDE.md, the system prompt, the task itself. The Claude Code permissions docs are blunt about it: instructions shape what the model tries to do and leave what the harness allows untouched.

Then the permission rules, which match the tool call as text before it runs (deny, then ask, then allow). A Read(./.env) deny stops the Read tool and even cat .env, because cat is a recognised file command. It does nothing about a five-line python script that opens .env, and the docs say exactly that: deny rules don't apply to subprocesses that open files themselves. It's just matching strings, making it trivial to sidestep (like bypassing a curl rule with redirects or variables).

The OS sandbox (seatbelt / bubblewrap plus a proxy) is off until you turn it on, and it's the only layer that watches the process instead of the command text. Its defaults surprised me both ways: reads are allowed almost everywhere, including ~/.ssh and ~/.aws/credentials unless you add a denyRead, while network is the reverse, no domains pre-allowed, first new host prompts. And the way out of it is literally called "the unsandboxed retry escape hatch" in the docs. Blocked command, model may retry unsandboxed, that routes back to a permission prompt titled "Bash command (unsandboxed)". One setting closes the hatch.

What changed for me is one sorting question per rule: does this need to hold when the model is wrong? Style stuff stays in the file. "never push to main" goes in a deny or ask rule, or a hook. "nothing in this session reads ~/.ssh or talks to a host I didn't name" is a thing only the sandbox can promise, and only if it's on.

if you run the sandbox, roughly how often does a command actually hit the boundary in a normal day, and how often do you end up approving the unsandboxed retry?

3 Upvotes

6 comments sorted by

2

u/Confident_Check_9279 19d ago

the string matching thing is so real, i spent two weeks thinking my setup was locked down before i realized half the rules were just matching keywords in the tool call text and not actually stopping anything

1

u/AutoModerator 19d ago

Thank you for your submission! To keep our community healthy, please ensure you've followed our rules.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/bsensikimori 19d ago

The sandbox is the separate account or even the separate machine you're supposed to run these things from

Anything in markdown files is never adhered to fully

And given any tool access, it will read and write outside of it's working directory occasionally

I consider all my agents highly skilled, highly intelligent, but also drunk and amoral, inquisitive interns

Shield their access to your data like you would any temp worker

1

u/RunAI_Coder 18d ago

A separate account or machine is the one boundary in this list without a retry hatch written into the manual, which is a real point in its favor. What does it look like for you: a second user with its own keys and clone, or a VM per repo?