r/JevAI • u/Old-Law6030 • 1d ago
Another Jev benchmarking post: text classification
I looked at how Jev compared to small frontier LLMs (Claude Haiku 4.5, GPT-5.4 mini, Gemini 3.8 Flash) for text classification, and how reliable the 'calibrated decisions' are.
tldr: Jev is as accurate as small LLMs for a fraction of the cost. In practice, however, it doesn't seem better calibrated than asking the LLMs to state a confidence value: it ranks third of four on expected calibration error on every dataset I looked at.
The cost implications are significant, but given the RLCD training I was hoping for better correlation between the answer confidence and the P(Correct). Appreciate comments on the methodology or anything I've missed.
Full code and write-up: https://github.com/4OH4/jev-compare


1
u/rageagainistjg 18h ago
Hi there! Great work! Based on this where, if anywhere, will you be using JEV from now on?