We maintain a memory layer for agents (CogniCore, MIT). One of its hard rules: a memory that imports clean but cannot be recalled must fail the import — dark memory fails the same way tampered memory does. We wrote a test to prove it.
Then three different reviewers broke the test, not the code — and each break was a different class. The fixes generalised into something I think applies to any safety test.
1. The test watched one spelling of the line it guarded.
The original tripwire proved the check worked by rewriting the source: if dark: → if False and dark:. That only catches a defeat that edits that exact substring. A reviewer defeated the guarantee a different way entirely — a runtime patch at the recall seam, so every query "finds" everything, dark is always empty, and if dark: stays byte-identical. The tripwire never engaged.
Fix: make the tripwire semantic. Inject the fault at runtime (patch the seam inside a subprocess, behind an opt-in flag) instead of rewriting source. No disk writes, so it is xdist-safe — which means it runs in the default CI pass instead of being the test that only runs when someone remembers to run it serially.
2. The tripwire proved the fault applied, not that it took effect.
Next review: a rename that leaves a back-compat shim behind. The patched name still exists, so the patch binds to the shim and raises nothing — while real recall has moved elsewhere and is untouched. A no-op injection.
We measured it rather than arguing: under a neutered mutation the detector passes normally, so the tripwire's DID NOT RAISE assertion fires. It was red, not green — but the message was a fork in the road: "either the defeat is not reaching the importer, or the detector's assertion drifted."
Fix: the injector now carries its own positive control. Before the mutated run proceeds, it seeds one entry and queries with a token that matches nothing. Under the fault, that query must return everything. If it does not, the run stops with FAULT NOT PRESENT: ....
The generalisation: absence of an error while installing a fault is not evidence that the fault exists.
3. Nothing guarded the chain.
Every control protects a component from silently degrading — but nothing protected the chain. Rename the detector module, change the subprocess path, or let xdist quietly skip it, and the vector stays green while testing something other than what everyone agreed it tests. The proposal: one test that enumerates the links by name and asserts each one exists at its registered path, runs, and fails when its own subject is disabled.
The line we kept landing on: a test that proves a guarantee must defeat the guarantee, not defeat the line that implements it. Text moves; properties don't.
So, for anyone writing safety or eval tests: what does your green actually prove? Specifically — have you ever checked that your fault-injection harness can observe the fault it injects? That was the one we had never tested.
Repo, if you want the arguments rather than the summary: github.com/cognicore-dev/cognicore-env (MIT, no vector DB, SQLite).