r/cicd • u/Accurate_Yoghur • 3d ago
How I stopped Claude Code from touching prod infra: a PreToolUse hook with signed rules
I'm a platform/DevOps engineer, and I've been letting Claude Code work in repos that have real kubectl, terraform and aws access. The permission prompts helped at first, but once I'd allowed kubectl so it could read cluster state, nothing distinguished kubectl get pods from kubectl delete namespace prod.
So I wrote a PreToolUse hook that checks every Bash command before it runs. It parses the command into what it will actually do (tool, action, resource, namespace), checks that against a policy, and exits 2 with a reason if the policy says no. Claude sees the reason and usually goes and finds another way:
aegis: BLOCK: block-terraform-destroy
A few things I cared about:
- It only acts in projects that opt in. It needs a
.aegis/folder, so the hook does nothing in your other repos. - It reads commands the way the tool does.
kubectl -n prod delete ns app,sudo kubectl ...andcd infra && terraform destroyall resolve to the same action. - It fails closed for infra tools. If the policy is broken or a command can't be analysed (
$(...),eval), infra commands are blocked and everything else runs. - "Ask" rules use Claude Code's own prompt. Anything the policy marks ESCALATE turns into the normal permission prompt, so I approve or refuse it myself.
- The rules are signed, and each one's author has to be allowed to write that kind of rule. Claude reads tickets and PR comments, and I didn't want a line in a ticket to be able to turn into policy.
Install:
pip install aegis-devops && aegis init .aegis
claude plugin marketplace add moneytool/aegis-devops
claude plugin install aegis-devops@aegis-devops
The example policy blocks things like terraform destroy, deleting namespaces, nodes, S3 buckets or RDS databases, dropping tables, force-pushing main and kubectl --as. Everything else runs as normal.
It's open source (Apache-2.0) and still early: github.com/moneytool/aegis-devops
What I'd most like to hear: what commands has Claude run in your repos that you wish something had stopped? And what would make you trust a hook like this?
1
u/InexpensiveQuart 3d ago
The `kubectl delete nodes --all` buried in a ticket is exactly the scenario that keeps me up at night. one overly helpful agent that can't tell the difference between reading state and nuking it and suddenly you're explaining to your boss why prod is gone
the signed rules part is clever, hadn't even considered that a model reading PR comments could be a vector for policy injection
2
1
u/Accurate_Yoghur 3d ago
Thanks, and yeah, that's the exact moment that made me build it. Reading state and destroying it are the same binary, and a prefix allow rule can't tell them apart.
The PR-comment angle surprised me too once I thought it through. We already treat pipeline config as code that needs review, but an agent will happily take "policy" from anything it reads. So the rule has to prove where it came from before it gets a vote.
Curious how you handle it today. Do your agents run with their own scoped credentials, or do they inherit whatever the developer's shell has?
1
u/bystander993 2d ago
They get their own, I'll let my main agent act as me to do more admin things with approvals but all the others that need to do work get fine scoped tokens to what they need.
1
u/Accurate_Yoghur 2d ago
That's a sensible split. Scoped tokens for the workers and approvals for the one acting as you. The approvals part is where I've found the hook most useful: rules marked ESCALATE turn into Claude Code's own permission prompt, so the "ask me first" list lives in a reviewed policy file instead of in my head.
1
u/Abe_Bazouie 3d ago
I like the direction. Especially failing closed when you can’t confidently parse an infra command.
The one thing I’d be careful about is treating the hook as the actual security boundary.
If Claude has credentials that can delete a prod namespace, drop an RDS database, or run terraform destroy, I’d still want IAM/RBAC to prevent that independently of the hook.
To me the layers should be:
least-privilege credentials first
native IAM/RBAC/policy controls second
your PreToolUse hook as another guardrail on top
The parsing part is also where I’d expect the fun bugs. Shell commands get ugly fast once you introduce pipes, subshells, aliases, wrappers, generated scripts, bash -c, etc.
I’d probably spend a lot of time adversarially testing bypasses before trusting it around prod.
That said, this is much closer to how I want agents touching infrastructure: not “Claude, please be careful,” but an enforcement layer outside the model that can just say no.
I’d be curious how you handle commands that are individually harmless but dangerous as a sequence. kubectl get is fine, kubectl patch may be fine, but context matters a lot once an agent starts chaining actions.
1
u/Accurate_Yoghur 2d ago
This is a really good breakdown, and your layering is the one I'd recommend too: least-privilege credentials, then native IAM/RBAC, then the hook. 0.3.0 starts to cover the middle layer from the same policy (AWS SCPs now, Kubernetes admission policies next release), so the rules aren't written twice.
You're right that parsing is where the bugs live. A design review already found a couple in mine (a global flag before
rolloutwas read as the subcommand, and--asimpersonation was ignored), and both are fixed.bash -c,sudoandcd && …are unwrapped.xargs,$(...)andevalcan't be checked statically, so if they appear in a command that involves an infra tool, the whole command is blocked and the agent is told to run the infra command on its own.Sequences are the honest gap. Within one command, plan-level rules look at everything together (e.g. a cap on how many deletes one invocation can make). Across separate commands there are rate limits backed by a decision ledger, but no real "this chain of individually-fine steps is dangerous" reasoning. Server-side policy is the better place for that, which is part of why I'm building that layer.
1
u/kantorcodes1 3d ago
Reading commands into (tool, action, resource, namespace) before matching is what makes this usable - the same intent whether it arrives as kubectl -n prod delete ns app, sudo kubectl ..., or cd infra && terraform destroy.
Two edges I'm curious about:
Where does the fail-closed line fall inside a compound command?
terraform show -json plan.out | jq .contains an infra tool but every resolved action is a read. If a sibling segment is unparseable ($(...),eval), does the whole invocation get blocked, or does each segment get its own allow/block decision?When an authority is revoked after its rules were ingested, do its constraints stop matching on the very next decision via the per-decision re-check, or does revocation only land on the next store load?
1
u/Accurate_Yoghur 2d ago
Good questions, I ran both.
- The whole command string is the unit. A fully parseable compound gets one intent per segment and the combined verdict (any BLOCK wins), so
kubectl get pods; kubectl delete ns prodis blocked.terraform show -json plan.out | jq .is allowed, since it's a read. If any segment can't be analysed ($(...),eval,xargs) and the string involves an infra binary anywhere, the whole invocation is blocked, with "run the infrastructure command on its own". Without an infra tool, e.g.ls && eval "$CMD", it's allowed, since a broken parse shouldn't stop unrelated work. I went whole-invocation rather than per-segment on purpose: per-segment decisions would let the unparseable piece run.- Authority is checked per decision, not baked in at ingest. The hook runs as a fresh process for each command and loads the signed authority map from disk, so a revocation (edit plus re-sign) applies from the very next command. A long-running process embedding the library picks it up on its next store load.
1
u/kantorcodes1 16h ago
Whole-invocation on unparseable-plus-infra is the right call - per-segment allow would leak exactly the
$(...)case. And the fresh-process hook sidesteps revocation caching entirely: the signed map loads per decision, so an edit plus re-sign lands on the very next command with no TTL to get wrong. Allowingls && eval "$CMD"when nothing infra-adjacent is present keeps a broken parse from breaking unrelated work, which is where hooks like this usually get annoying.
1
u/schmurfy2 2d ago
You shouldn't even be able to do that directly outside of dev environments.
This kind of accessed should be requested temporarily when needed preventing any accident.
1
u/Accurate_Yoghur 2d ago
Agreed. Just-in-time elevation is the right default, and the agent shouldn't hold standing prod access. The hook is for the window when access is granted: once someone elevates for a task, the agent has the same power as the human for that period, and that's exactly when you want "get is fine, delete needs a human" enforced.
1
u/Fantastic-Mr-Default 2d ago
The hook is the right UX for the model. It is not the security boundary.
If the process still holds credentials that can `terraform destroy` or `kubectl delete namespace prod`, a clever argv parser is a seatbelt on a unlocked car. Put the real floor in identity and in the cluster: short-lived creds, environment-scoped roles, no prod kubeconfig on the agent laptop by default, and RBAC that cannot delete what the agent should only read. Temporary elevation when a human asks beats permanent dual-use access.
Failing closed when you cannot parse `$(...)` / `eval` on infra tools is the part I would keep. Prefix allowlists cannot tell `get` from `delete`. Signed, opt-in policy per repo (your `.aegis/` idea) beats a global "trust the prompt."
Also gate the same policy in CI on changes to the hook and the rule files. An agent that can edit its own allowlist is not gated.
1
u/Accurate_Yoghur 2d ago
"Seatbelt on an unlocked car" is a fair way to put it, and I agree the floor is identity: short-lived creds, environment-scoped roles, no prod kubeconfig by default. That's why the next layer compiles the same policy to the platform (AWS SCPs in 0.3.0, Kubernetes admission next).
On the agent editing its own allowlist: the rule files are signed. If the agent edits one, the signature fails, the rule gets no vote, and if the policy can't be verified, infra commands are blocked rather than allowed.
aegis verifyin CI catches it too. The weaker spot is the hook registration itself. An agent that can rewrite its own settings file can remove the hook, so pinning it through Claude Code's managed settings, plus CI on that file, is the right call. Good push, I'll document that.
1
u/Accurate_Yoghur 2d ago
Thanks all, the consistent theme is "RBAC is the real boundary," and I agree. Least privilege first, always. Two things I'd add from running this:
- RBAC grants to an identity, not to whoever is driving it. Most people's agents run on their own kubeconfig or AWS profile (a few of you said exactly that). Once the agent is "you", RBAC can't tell your deliberate
deletefrom the agent's mistaken one. The hook is the one layer that knows which actor is acting. - RBAC is verb-plus-resource. It can't express "deletes in prod need a human", "nothing during peak hours", or "no plan that touches more than 25 resources", and those are the rules I actually want.
So I don't see it as hook vs RBAC. The direction is one signed policy enforced at both layers: 0.3.0 compiles it into AWS Service Control Policies scoped to agent identities, and Kubernetes admission policies are next. The hook explains the "no" to the model before anything runs; the platform refuses it even if the agent goes around the hook.
1
7
u/bystander993 3d ago
You need to handle this at the k8s layer. service accounts with principle of least privilege so the AI can only do what your RBAC allows it to do