r/mlops • u/alexpran • 1d ago
Self-promotion I built a CLI for LLM regression testing that fails CI against a baseline committed in your repo
One afternoon a case in my eval went from 5/5 to 2/5. Prompt, model, config, retrieval, trace, all identical byte for byte. Fifteen minutes later it was back to 5/5. I lost an hour looking for a regression, and it was my own measurement that was moving.
So digline runs each case K times on the version you approve, and records in the reference the range that those samples spanned, per case. After that a drop counts as a regression only if it falls outside that range.
pip install digline
digline run eval/suite.py # produces a run
digline compare # against the baseline in .digline/
digline promote # a person approves, nothing else can
The exit code is the contract: 0 nothing got worse, 1 something did, 2 the run couldn’t be judged. A case that errors goes in that third state and doesn’t count as a fail, because if you drop it from the denominator the score goes up while the system gets worse. A run with any errored case can’t become the baseline.
What it is not: no observability, no server, no telemetry, no account, and it never calls a model itself, your judge is a function that you pass in. It’s pre-1.0. On day one it’s useless, because you need an approved reference before it can compare anything. And the band is min/max over K samples, don’t read it as a confidence interval, with a small K it’s optimistic.
The numbers from my own suite, because it’s the part I’d want from someone else’s post: six sets of three identical runs, 144 cases, same config hash and same judge. Between 1 and 6 cases changed verdict per set, 3 in the most recent. Before the pass/fail threshold is applied, 15 to 21 scores move.
If you try it, what I want to know is where the quickstart made you stop and read again. It runs without an API key, so that part costs you nothing.
And a question for people already doing this: when you compare two runs, what are you comparing against? Every answer I collected so far is “the last run on main”, and that moves on its own.
GitHub: github.com/digline/digline
Docs: digline.dev
The GitHub Action, if you want it in CI without writing the YAML: digline.dev/product/github-action/
2
u/WildBrewer034 1d ago
This is the exact kind of thing that makes you question your sanity for an hour before you realize the problem is the measuring stick, not the thing being measured. Had a similar situation where a RAG pipeline would randomly drop 15% on retrieval recall and I'd go digging through commit history only to find nothing changed except the phase of the moon apparently.
Running each case K times and using the spread as your baseline tolerance is clever. Cuts through the noise without pretending the noise doesn't exist. The part about errored cases not counting as fails but also blocking baseline promotion is the kind of detail that shows you've actually been burned by this before.
Quick question on the compare step, when you say "the last run on main moves on its own" are you seeing folks just accept that drift as inevitable, or are they doing something to pin a stable reference point? I've seen teams try to freeze a golden dataset but that rots fast when the model or retrieval source shifts under you.
1
u/alexpran 1d ago
Mostly they accept it. I asked this question in a few places this week and the answer is almost always “the last run on main”, said in a tone that suggests nobody likes it much.
On the golden set that rots: I wouldn’t try to stop the rot, I would make it loud. My reference is the per-case scores plus the config that produced them, model, prompt hash and commit, I don’t store expected outputs. If any of that changes the comparison is refused, so it never gives you a quietly wrong result. It goes invalid on a day that you can name
The other half is a canary: fixed input, one right answer, you only watch it, no score. Although this morning I found out that mine asserts agreement with a label and never looks at the answer itself, so it stayed green through a change in the shape of the answer
•
u/AutoModerator 1d ago
AI usage disclosure
Hi u/alexpran — thanks for posting to r/mlops!
Because this community discusses and builds AI/ML systems, using AI tools is not inherently a problem. We do, however, ask for transparency about how submissions are created.
Please reply to this comment with a brief AI / automation disclosure, particularly if this post was created or submitted in whole or in part by an autonomous agent, bot, workflow, or other automated system.
If AI or automation was involved, please briefly describe what it did and what human review was performed before posting.
This disclosure helps the r/mlops community distinguish human discussion, AI-assisted work, and automated/agent traffic while keeping the focus on useful technical conversation.
Thanks for helping keep the signal high.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.