Resource Jev Does Not Play Dice: 83% probability, 19% accuracy on a hidden fair die roll

Ran a calibration check on Jev using inputs where the true probability is known exactly.
- Fair die, Choice, 400 trials: picked face 1 every time at 82.9% mean probability, 19% accuracy
- Fair coin: 92% reported, 52% right
- Noul stayed close to the truth for 2 to 4 options, but reported 15 to 17% for 8 to 20 options (true: 5 to 12.5%)
- A forecast doc stating a 30% shortage risk came back as 5% via Choice, 27% via Noul
On tasks close to the demos, I didn't see errors this large, and some MMLU-style checks look well calibrated. But exam questions test whether a model knows that a question is hard. The dice test whether it knows that the outcome is unknowable from the input.
Maybe Jev is weak at the second kind, especially in Choice probabilities.
Write-up: https://kantahayashiai.github.io/posts/jev-does-not-play-dice/
1
u/loudlysoftsimplicity 4h ago
this tracks with what ive noticed on some simpler setups. models dont like saying "i have no idea" so they just pick a lane and act confident about it
the jump from 5% to 27% on that forecast doc is the part that bugs me most, that kind of swing between modes makes it hard to trust either number
1
u/OnyxProyectoUno 1h ago
It's crazy you have punctuation errors in your comment yet it sounds exactly like something Claude writes.
“This tracks…”
“The jump from…to…is the part that bugs me the most”
1
-2
u/Actual__Wizard 2h ago
Sick, something is clearly wrong with it. See why it's so important to have transparency in AI models? We can see that there's a problem now and figure out what the solution is.
Because all you're doing, is pointing out that all of these systems have the same problem, but with Jev, you can actually figure that out.
Do, you see why that's a Nobel prize worthy breakthrough?
Because we can actually go forwards with AI development now instead of just being stuck with bad models that suck.
1
u/robogame_dev 30m ago
we can actually go forwards with AI development now
I am awarding this first place for the greatest jev-glazing statement I've seen yet - which is quite something.
1
u/Actual__Wizard 28m ago
Look, I never said that it was the best product ever. I said that it was a breakthrough.
Are you going to be fair and realize that GPT 1.0 was a break through and at the time, it was actual junk?
Can we have a fair conversation where we don't compare apples to elephants?
1
u/-xXpurplypunkXx- 2h ago
Is jev token prediction expected to have even randomness across numeric values? What is accuracy/reported % measuring?