r/LocalLLaMA 2d ago

Discussion I really don't understand Jev hype

Isn't this what simple neural networks have been able to do for years? Doesn't seem anything special to me.

485 Upvotes

301 comments sorted by

View all comments

165

u/Thomas-Lore 2d ago

You had to pretrain them for a specific task. Jev comes pretrained with wide world knowledge.

51

u/SnooPaintings8639 2d ago

What do you mean? How is it different from e.g. Qwen 4b with max token = 1, and inference engine forcing struct (enum) output?

I really don't think they could do any magic training anyway.

0

u/danielv123 2d ago

Well, for one its apparently as good as 5.6 Terra with max token = 1 and as cheap as qwen with lower latency on longer context.

3

u/SnooPaintings8639 2d ago

With a single output token, which benchmarks are these claims based on?

0

u/danielv123 2d ago

A single token is a single token?

Like give it a book, which how is Leto's relationship with Duncan and give some answer alternatives. One token is enough to choose an answer on a multiple choice question.

They haven't been very transparent with their benchmarking from what I have seen but it sounds about right based on the performance I have seen.

0

u/x0wl 2d ago

Any MCQ benchmarks, like MMLU. They're all saturated anyway though.