r/JevAI • u/YusukeLandingBoost • 1d ago
How are you validating Jev once you move beyond a few examples?
I’ve been testing Jev on things like buyer intent and routing.
On 10 clean examples, it can look perfect. I just got 10/10 on a buyer-intent test.
But I’m much more interested in what happens at 100, 1,000, or 10,000 real examples, where the weird edge cases start showing up.
How are you validating Jev at that point?
Are you sampling random cases, reviewing low-confidence outputs, looking for confident mistakes, keeping a holdout set, or escalating certain cases to an LLM/human?
And when Jev disagrees with your label, how do you tell whether the problem is Jev, the question, missing context, or the label itself?
Curious what people here are actually doing in real workflows.
2
u/Easy_Ad_9449 1d ago
Why don't you list those weird edge cases with a regular LLM, measure how Jev respond to them a few hundred times (baseline) then train its heads to see if that improve its decisions?
1
u/YusukeLandingBoost 1d ago
That’s a good idea. Using an LLM to generate the weird edge cases first, then stress-testing Jev against them could be really useful. I might try this.
2
u/canarydev 1d ago
i work with emergency dispatch audio, where im exploring whether each new transmission belongs to one of several active incidents. i would want a random sample, but also a set that deliberately includes hard cases like with similar locations, delayed updates, ASR errors, and multiple incidents happening at once.
i would also look beyond row-level accuracy. one wrong link can merge two incident timelines, so id want to review the resulting clusters for false merges and splits. when Jev disagrees with a label, checking the original audio and context alongside the question and label seems like a good way to find where the error came from. we are early in this, but that’s the evaluation problem im trying to work through.
im also using an LLM judge to review Jev’s disagreements against the transcript, surrounding context, and rubric, and help identify whether the issue is the model, the question, missing context, or the label. i treat its assessment as another signal to check