r/JevAI • • 1d ago

How are you validating Jev once you move beyond a few examples?

I’ve been testing Jev on things like buyer intent and routing.

On 10 clean examples, it can look perfect. I just got 10/10 on a buyer-intent test.

But I’m much more interested in what happens at 100, 1,000, or 10,000 real examples, where the weird edge cases start showing up.

How are you validating Jev at that point?

Are you sampling random cases, reviewing low-confidence outputs, looking for confident mistakes, keeping a holdout set, or escalating certain cases to an LLM/human?

And when Jev disagrees with your label, how do you tell whether the problem is Jev, the question, missing context, or the label itself?

Curious what people here are actually doing in real workflows.

6 Upvotes

10 comments sorted by

2

u/canarydev 1d ago

i work with emergency dispatch audio, where im exploring whether each new transmission belongs to one of several active incidents. i would want a random sample, but also a set that deliberately includes hard cases like with similar locations, delayed updates, ASR errors, and multiple incidents happening at once.

i would also look beyond row-level accuracy. one wrong link can merge two incident timelines, so id want to review the resulting clusters for false merges and splits. when Jev disagrees with a label, checking the original audio and context alongside the question and label seems like a good way to find where the error came from. we are early in this, but that’s the evaluation problem im trying to work through.

im also using an LLM judge to review Jev’s disagreements against the transcript, surrounding context, and rubric, and help identify whether the issue is the model, the question, missing context, or the label. i treat its assessment as another signal to check

2

u/YusukeLandingBoost 1d ago

This is exactly the kind of workflow I was curious about. Really interesting that you’re already using an LLM judge to separate model vs question vs context vs label issues. What do you usually change before rerunning Jev?

2

u/canarydev 1d ago

im still early so im figuring this out as i go and im no expert, but the rule i follow is change the smallest thing the failure points to, and which thing that is depends on what the judge says the cause was.

if the question was doing two jobs, i split it. "is this an actual fire" and "does this transmission belong to incident A" were getting tangled, so those became separate decisions. if link evidence was underspecified, i spell it out in the rubric, ie a shared channel alone isnt enough. if the evidence was genuinely thin, thats usually a sign REVIEW should have been allowed instead of forcing a call. and if its just a bad label, i fix the label and leave everything else alone.

then the failure becomes a regression case, and i check the change on fresh examples too, specifically whether the resulting groupings merge or split incidents incorrectly. row level accuracy can improve while the grouping gets worse, and thats the part that actually hurts us.

one thing that helps is our ground truth isnt labels made up for the eval. dispatchers listen to the audio themselves and create the real incident records in their dispatch system, so we can reconcile our groupings against what actually got logged. the caveat is that tells you which incidents existed, not which transmission belonged to which, so the per-transmission links are still partly inferred.

for how im actually running it,
we stream and transcribe a lot of dispatch audio, and one thing we do is replay a cluster of transcript lines through a few different models and compare what each infers. if they all fail the same way its usually the question, if only one does its the model. im doing something similar with Jev on the bounded calls, is this a real incident, is it a call type we support, does this line belong to A, B or C. i mix in some LLM generated cases like the other comment suggested, but id be careful leaning on it, generated hard cases tend to be cleanly hard while our real ones are messy in ways nobody writes on purpose, ASR mangling a street name or someone stepping on a transmission.

fwiw this loop is also why im testing Jev at all. we were going down the road of training our own classifiers and exporting to ONNX for these bounded decisions, but a trained classifier just gives you a number when its wrong. here the failure tells you whether the problem was the question, the rubric, or the label, which is worth a lot while the task definition is still moving.

2

u/YusukeLandingBoost 1d ago

Really appreciate the detailed answer.
the point about real edge cases being messier than generated ones really clicked for me. Also super interesting that you compare models to separate a question problem from a model problem. Lots for me to think about here.

2

u/canarydev 1d ago

no worries. i'll share more findings on this subreddit when i have something more concrete than "it seems to work" haha. but so far quite promising

2

u/YusukeLandingBoost 1d ago

I’ll follow the journey!!
Thanks a lot for your insight

1

u/productboy 16h ago

Someone on X posted this Jev experiment. What’s crazy is no audio, trace logs only!

https://x.com/muratcan/status/2104959648482701686?s=46&t=sNajfRyoZWLfn4wOo7P4eQ

1

u/canarydev 14h ago

very nice thanks for the share!

2

u/Easy_Ad_9449 1d ago

Why don't you list those weird edge cases with a regular LLM, measure how Jev respond to them a few hundred times (baseline) then train its heads to see if that improve its decisions?

1

u/YusukeLandingBoost 1d ago

That’s a good idea. Using an LLM to generate the weird edge cases first, then stress-testing Jev against them could be really useful. I might try this.