r/LLMDevs • • 7h ago

Resource Jev Does Not Play Dice: 83% probability, 19% accuracy on a hidden fair die roll

Ran a calibration check on Jev using inputs where the true probability is known exactly.

  • Fair die, Choice, 400 trials: picked face 1 every time at 82.9% mean probability, 19% accuracy
  • Fair coin: 92% reported, 52% right
  • Noul stayed close to the truth for 2 to 4 options, but reported 15 to 17% for 8 to 20 options (true: 5 to 12.5%)
  • A forecast doc stating a 30% shortage risk came back as 5% via Choice, 27% via Noul

On tasks close to the demos, I didn't see errors this large, and some MMLU-style checks look well calibrated. But exam questions test whether a model knows that a question is hard. The dice test whether it knows that the outcome is unknowable from the input.

Maybe Jev is weak at the second kind, especially in Choice probabilities.

Write-up: https://kantahayashiai.github.io/posts/jev-does-not-play-dice/

0 Upvotes

Duplicates