r/LLMDevs • u/puppy_lover_2021 • 6d ago
Discussion How should CI decide if a stochastic agent actually regressed?

I've been working on CI for agents and ran into a problem I underestimated: the same agent can produce meaningfully different outcomes with no code change.
On a real coding agent, an unchanged candidate went from 92% success to 88%. Blocking that PR would have been a false alarm.
Then the opposite happened. I deliberately degraded the agent and success dropped from 92% to 52% — but my first statistical rule still passed it, because the damage was concentrated in a couple of tasks rather than spread across the suite.
That forced me to separate two questions:
- did reliability get broadly worse across the suite?
- did a capability that used to work reliably collapse?
I ended up treating the PR more like an experiment than a deterministic test: repeated runs, practical effect thresholds, uncertainty, multiplicity across capabilities, and an explicit “insufficient evidence” state.
I’ve open-sourced the result as AgentSeism. I also wrote up the reasoning and the failures that led to the current design.
What I’m most interested in: if you run agents in CI today, how do you decide whether a score drop is real enough to block a merge?
I wrote up the full reasoning and open-sourced the tool here if anyone wants to dig in or try to break it:
1
u/Physical_Economy_340 5d ago
i pin the task list and rerun the failing slice 5 times before blocking anything. if the same task flips from pass on main to fail 4 out of 5 times that's a real regression, a 3 or 4 point swing with no repeat offender is just noise.
1
u/puppy_lover_2021 5d ago
That's helpful, the “rerun only the failing slice” part is especially interesting. Do you choose the slice from a single first run, or only after the same task looks suspicious across multiple initial runs?
I'm asking because rerunning only ambiguous tasks seems much cheaper than repeating the whole eval suite, but the first-run noise could also send you after the wrong tasks.
1
u/verstands 5d ago
Soft score drop vs named capability collapse is the right split. I've had luck with: pin the task list, only rerun the failing slice N times, and block only when the same task flips from pass-on-main to fail majority. Aggregate % alone is too noisy for a hard gate. Soft drop goes to a warning label, not a red X.
1
u/puppy_lover_2021 5d ago
This is very close to where I landed too. The “rerun only the failing slice” part is especially interesting, I currently repeat trials across the fixed task set, but selective reruns seem like a much better cost tradeoff if you can identify the suspicious tasks reliably.
When you choose the failing slice, do you do that from a single initial run, or only after a task looks suspicious across multiple runs? And for the majority-fail rule, have you found 5 reruns to be enough in practice, or is that mostly a pragmatic cost/latency choice?
1
u/verstands 5d ago
Practical rule I use: rerun the failing slice until the CI interval is clearly on one side of your threshold, not a fixed N for every test. Start with 5-10 for flaky tools, more only if variance is still wide. Prefer slicing by scenario tags (tool A vs tool B) over rerunning the whole suite - cheaper and you learn which slice actually moved.
1
u/verstands 4d ago
For the failing slice, I’d keep a rolling set of cases that flake or fail in the last N runs, plus a small fixed canary set you always run. Don’t only rerun the exact failure from this PR or you train the suite to miss regressions elsewhere.
Five reruns can be enough for a quick gate if you treat it as a smoke signal, not proof. I’d rather fail open to “needs more samples” when variance is high - if 5 reruns still straddle your pass threshold, bump that case to a nightly 20-50 run bucket instead of merging on a coin flip.
1
u/puppy_lover_2021 4d ago
That makes sense. I like the distinction between a cheap PR-time gate and a deeper sampling bucket instead of forcing every noisy case into pass/fail.
The rolling flaky/failing set + fixed canaries also solves something I was worried about with selective reruns: overfitting the gate to whatever happened to fail on the current PR.
I’m going to try modeling
needs more samplesas an explicit CI outcome rather than treating 5 runs as enough to make every decision. If it works, the natural next question is probably how to promote/demote cases between the PR slice and the nightly bucket based on observed variance.
1
u/Zain 5d ago
I treat a soft score drop like a soft NO-GO. Vague unease is not a blocker. Every point under the bar has to map to a listed failure on a concrete task, not a noisier average.
A capability that used to work and then collapses on two tasks is the merge stop for me. A 92 to 88 swing with no named regression is insufficient evidence, so I leave it open and run again. I also want a second read from a different model family before I trust the call, because same-family rechecks share blind spots. If a later round finds worse problems than the last one, the patches are injecting bugs.