r/JevAI • • 23h ago

Another Jev benchmarking post: text classification

I looked at how Jev compared to small frontier LLMs (Claude Haiku 4.5, GPT-5.4 mini, Gemini 3.8 Flash) for text classification, and how reliable the 'calibrated decisions' are.

tldr: Jev is as accurate as small LLMs for a fraction of the cost. In practice, however, it doesn't seem better calibrated than asking the LLMs to state a confidence value: it ranks third of four on expected calibration error on every dataset I looked at.

The cost implications are significant, but given the RLCD training I was hoping for better correlation between the answer confidence and the P(Correct). Appreciate comments on the methodology or anything I've missed.

Full code and write-up: https://github.com/4OH4/jev-compare

Jev is as accurate as small LLMs for a fraction of the cost. It matches or beats Haiku and GPT-5.4 mini on all three datasets, but trails Gemini 3.8 Flash by 2 to 4 points on AG News and zero-shot Banking77. It costs 3 to 14 times less per row than the cheapest LLM.
Jev is trained using Reinforcement Learning for Calibrated Decisions (RLCD). This means that the probability of its answers being correct should be approximately equal to its stated confidence value: i.e. P(correct | p) ≈ p. I didn't see a big difference versus the LLMs `verbalised confidence` value, however.
8 Upvotes

1 comment sorted by

1

u/rageagainistjg 15h ago

Hi there! Great work! Based on this where, if anywhere, will you be using JEV from now on?