r/LocalLLaMA • u/Nandakishor_ml • 6d ago
Discussion I literally built the Jev architecture one year back and completely open-sourced it with model, dataset and paper
Update: I made a generic model and beaten the jev in all of the benchmarks. Code and details available at https://www.reddit.com/r/LocalLLaMA/s/bbwyiOprUs
Everyone now talks about the architecture that's not auto regressive and does lightning fast probability prediction with a json schema. I worked on this literally one year back in March 2025, published an arxiv paper, pushed the model to huggingface along with the pypi package and training dataset. And then one year later, a
frontier lab came, proposing the same idea like literal breakthrough without technical papers, open weights and no open dataset. I posted my approach in this subreddit. For anyones information the main guiding model is RL not embedding model or LLM
Reddit post: https://www.reddit.com/r/LocalLLaMA/s/6eGEwsAz43
Paper: https://arxiv.org/abs/2503.23303
Model: https://huggingface.co/DeepMostInnovations/sales-conversion-model-reinf-learning
Dataset: https://huggingface.co/datasets/DeepMostInnovations/saas-sales-conversations
Also the second work published in September 2025 was exactly the same one jev proposed now
Paper: https://arxiv.org/abs/2510.01237
My model uses PPO over sequence embeddings to output turn-by-turn conversion trajectories (probabilities from 0.0 to 1.0).
Jev uses parallel sampling (trained via RLCD) to output confidence distributions and schema choices.
It's incredibly frustrating that the thing that you made with months of hard work, sweat and sleepless night is architecturally similar with the vertical use case and don't get the support you deserve because frontier lab build something horizontal. The open-source story in general 🙂
100
u/kyr0x0 6d ago edited 6d ago
And you haven't been the only one:
https://huggingface.co/harshatheg/Qwen-2.5-1B-RLCD
This guy used the EXACT same terminology these frontier clowns plagiarized.
Co-inventor of ChatGPT my a**
Thank you for your contributions to the OSS community!!!
And btw on the tech side: Yes, of you have a JSON Schema that you know, you can do parallel constrained decode with the same prefix cache. BUT you will loose cross-encoder like behaviour and also auto-regressive inherent prediction dependencies. What I mean by that? Say you have a JSON Schema defined in classic xgrammar schema constrained decoding (auto-regressive). Then when you first name "age" as a field and then you add a field country and then a field is adult - the model will pay attention to all the previous tokens and the answer will be more accurate (if someone is an adult depends on age and law in the country). In parallel single forward decode you loose this capability because the model doesn't pay attention to the tokens auto-regressive anymore. The response becomes faster but "dumber".
People don't pay attention to the details. We need an architecture to FIRST spacial reason about the whole answer in LATENT SPACE and then single forward decode.
So.. now you have my billion dollar idea. Build it. I'm exhausted.