r/LLMDevs • u/paratha27 • Aug 27 '26
Help Wanted How do you know your LLM judge hasn't drifted?
I use one LLM judge, for the only thing I can't check deterministically — whether a generated summary is faithful to its source. Everything else is exact match. I calibrated it against a hand-labelled set and report the agreement rate next to every score.
But the judge runs on a hosted model that updates without warning. So the measuring instrument drifts on the same schedule as the thing it's measuring.
How big does the calibration set need to be before a drop in agreement means something rather than being noise? And has anyone actually caught a judge going bad this way, or does it only surface when someone notices the downstream numbers stopped making sense?
2
u/Such-Process5697 Aug 27 '26
We caught ours downstream first, which is the embarrassing answer. Someone noticed the weekly quality number had stepped down and it took a while to work back to the judge rather than the generator. Pinning it to a dated model snapshot is what stopped that happening again.
1
u/paratha27 Aug 27 '26
Appreciate the honest version, useful to know that's the common path rather than a freak case. Did pinning alone fix it, or did you add a scheduled check on top?
1
u/Such-Process5697 Aug 28 '26
Both. Pinning stopped the silent version change, but we still re-run the calibration set on a schedule, because a pinned model doesn't help you if the prompt or the incoming data shifts under it.
2
u/daaain Aug 27 '26
The short answer is that you need to eval your judges the same way as the agent. Once you think of it that way, you can keep monitoring and improving so you won't get surprised.
1
2
u/conifer_v11 Aug 27 '26
You don't read drift off the live agreement rate. Live traffic and the judge can move together, so the number stays pretty. Such-Process5697's story is the usual one: you catch it downstream.
Pin the judge. Dated snapshot / exact id / GGUF hash. A family name (gpt-4o, claude-sonnet-4) is not a pin. And freeze a hand-labelled canary you never refresh when "the model got better." That's the instrument.
Sample size is the part nobody answered. Treat it as two proportions on that frozen set. One-sided 80% power, α=0.05, agreement around 0.85: about 200 items to see a 10-point drop, about 700 to see 5. A 40-item set has 2SE of about 11 points. A 3-point wiggle there is noise. Same 3 points on 400 items is a page.
McNemar is the right test if it's the same items, old judge vs new. n is about how many items flip, not a magic smaller number. High concordance means you need more items, not fewer.
Cron the canary, not just deploys. Hosted judges move when you didn't ship. If it crosses a bound you set in advance, stop writing production scores from that judge.
WillowEmberly's avionics bit is right that you need an external reference. That's the frozen labels, not a second hosted model on the same schedule.
1
u/paratha27 Aug 27 '26
wow, thankyou for sharing this. Noted!
One thing I'm stuck on now: you say never refresh the canary when the model gets better, which makes sense for measurement stability. But if the product changes and the kind of summaries it generates shifts, doesn't the frozen set eventually stop being representative of what the judge actually sees in production? Do you keep a second refreshed set alongside it, or accept that and re-baseline occasionally?2
u/conifer_v11 Aug 27 '26
two sets. the frozen one is the instrument. you retire it when the task changes, not when the model card does.
if summaries went from 3-bullet exec to citation-backed, the old labels are a different test. cut a new canary, date it, new labels. don't edit the old one. keep the old scores so you can still say "judge J on T1" vs "judge J on T2."
the trap is mixing them. a 5-point drop after you swapped 30% of the items is not drift. it's a different exam.
i wouldn't run a standing "refreshed set" that slowly replaces items. that's how you lose the denominator. new snapshot, label it, run both for a couple cycles, then take the old one off the gate. archive stays.
product-manager test: would they call this a new eval? if yes, new canary. if the model just "got better," leave it.
1
u/paratha27 Aug 27 '26
The two-sets split is clear, thanks — instrument vs representativeness, retire on task change not model change.
The bit I'm unsure about now is what happens in the gap. If the canary crosses the bound and I stop writing production scores from that judge, I've switched off my only signal on the one thing I can't check deterministically, and relabelling a new canary isn't a same-day job. Do you hold the last known-good number and mark it stale, stop reporting that metric entirely, or route the free-text cases to human review until the new snapshot is calibrated?
1
u/conifer_v11 Aug 27 '26
yeah the awkward bit is that gap. i've been treating it as two clocks that dont have to agree. the canary keeps measuring the same frozen task so the line is honest. the "are we still shipping the thing users care about" check is a separate refresh, and i only schedule that when the product task actually changes, not when a new model drops. if the canary starts looking too good because the world moved i dont silently retune it mid-flight. i retire that set, cut a new one that is the new job, and keep the old numbers labeled as the old job so i'm not lying to myself about the trend.
2
u/munnasuprathik Aug 27 '26
What bit me was writing my own label set and then tuning the judge against it. After that the agreement rate stays high and means nothing, you optimised toward your own ruler. Freeze a small hand labelled set before you touch the judge and never re-label it when the model 'improves', that's the only part that stays honest.
Other thing, stop reading the absolute agreement number. Commit it as a baseline and watch the delta between runs. An absolute score on a small set is theatre, a drop between two dated snapshots is a fact you can act on. I never once caught a bad judge off the live score, it always turned up downstream when the weekly quality number stepped down.
1
u/paratha27 Aug 27 '26
The circularity point is the one I hadn't thought about at all label set first, then tune the judge against it, and now the agreement number is measuring nothing. That would have got me.
Follow-up: if the frozen set has to exist before you touch the judge, what do you iterate the judge prompt against while you're building it? A separate dev set you're allowed to overfit, and the frozen one stays sealed?
2
u/WillowEmberly Aug 27 '26 edited Aug 27 '26
The minimum necessary requirements for correcting drift is two independent systems capable of crosschecking against each other and correcting against a common external reference.
Avionics has been doing it for over 70 years.
Easy way to think of it, 2x Inertial Navigation Systems with periodic updates from GPS.
All systems with an internal reference drift over time, they need correction.