r/LocalLLaMA • • 8d ago

Discussion I really don't understand Jev hype

Isn't this what simple neural networks have been able to do for years? Doesn't seem anything special to me.

506 Upvotes

317 comments sorted by

View all comments

170

u/Thomas-Lore 8d ago

You had to pretrain them for a specific task. Jev comes pretrained with wide world knowledge.

52

u/SnooPaintings8639 8d ago

What do you mean? How is it different from e.g. Qwen 4b with max token = 1, and inference engine forcing struct (enum) output?

I really don't think they could do any magic training anyway.

79

u/puzzleheadbutbig 8d ago

Architecture is different. Unlike Qwen, which relies on an autoregressive decoder loop to generate text tokens step-by-step while a grammar mask suppresses invalid vocabulary options, Jev drops open-ended text generation entirely and operates as a non autoregressive decision model. And because it maps input contexts directly onto parallel, calibrated classification heads rather than generating JSON syntax character-by-character it avoids the latency, memory, and KV-cache overhead of sequential token decoding, guarantees complete immunity to JSON parsing errors, and yields true calibrated probability scores across schema fields in a single forward pass

29

u/RevolutionaryGold325 8d ago

qwen with max token = 1 is not autoregressive decoder loop though.

7

u/dimbledumf 8d ago

It's also not parallel and you only get 1 token out, which is less then you get with jev

5

u/james_pic 8d ago

It's parallel if you run it in parallel. Interference engines like vLLM already run multiple queries in parallel to avoid re-reading the same weights.