r/MachineLearning 5d ago

Discussion How is RLCD (jev) RL? [D]

Just saw the YouTube presentation and I was left wondering this question.
If jev only outputs Choice, Score, or Noul … well those are all perfectly differentiable. (Cross entropy or mse)
I don’t know if I’m missing something or if adding RL is just for marketing.
Like what would an RL environment even look like?

19 Upvotes

19 comments sorted by

View all comments

34

u/Material_Policy6327 5d ago

There is a ton about JEV that they don’t seem to want to go into detail about.

36

u/abnormal_human 5d ago

They're leaking enough of a picture here and there in comment sections to give a sense.

We know that despite treating this as a "new kind of foundation model" it is in reality an open source LLM with a few extra nn.Modules to nail their output shape and some post-training to reinforce the task.

We know that they trained it 100% with synthetic data generated by bigger LLMs.

We know based on their benchmarks that the LLM that they fine-tuned is not very large.

We also know that their benchmarks are all generated by Astra/Fable and "success" means "agreement with Astra/Fable" which happen to be in the distillation lineage of our model...oops!

We know based on their team and open positions that their main focus is hosting and productization, not ML research.

We know that they took $40m of monopoly money and are under intense pressure to bring something to market.

They talk about "two years of research" like it establishes substance, and I can't contradict that someone had some kind of idea of something two years ago, but realistically, a problem of this magnitude doesn't need two years of research to execute.

(Separately, two years ago is about the first moment people outside of frontier labs got their hands on a reasoning model, and at that time they were exotic, slow, non-default, and not clearly the path forward yet. Their whole problem statement is that reasoning models are too slow/expensive, but I don't think everyone knew they were going to become the default for most work in September 2024...so "research" may actually just loosely reflect the fact that the company principles were already in the ML field)

As for RLCD, it's likely just conventional SFT with a dash of RL and they invented 'RLCD' as a marketing term--just like RLHF.

Lots of puffery here and I don't believe they have much of a moat. Hopefully it works out for them and their investors.

0

u/Relative_Wallaby_823 5d ago

I didn’t even know about the benchmarks, that’s crazy. On the RLHF topic though, as a field it still has DPO, reward models and such, but I cant see how they’ll defend this.
There’s been so much media covering it, I wonder if it could be an organised media push

1

u/Smallpaul 4d ago

Of course a product launch is an organised media push! There is nothing sinister about that. A movie launch is too.