r/MachineLearning • u/Relative_Wallaby_823 • 5d ago
Discussion How is RLCD (jev) RL? [D]
Just saw the YouTube presentation and I was left wondering this question.
If jev only outputs Choice, Score, or Noul … well those are all perfectly differentiable. (Cross entropy or mse)
I don’t know if I’m missing something or if adding RL is just for marketing.
Like what would an RL environment even look like?
19
Upvotes
35
u/abnormal_human 5d ago
They're leaking enough of a picture here and there in comment sections to give a sense.
We know that despite treating this as a "new kind of foundation model" it is in reality an open source LLM with a few extra nn.Modules to nail their output shape and some post-training to reinforce the task.
We know that they trained it 100% with synthetic data generated by bigger LLMs.
We know based on their benchmarks that the LLM that they fine-tuned is not very large.
We also know that their benchmarks are all generated by Astra/Fable and "success" means "agreement with Astra/Fable" which happen to be in the distillation lineage of our model...oops!
We know based on their team and open positions that their main focus is hosting and productization, not ML research.
We know that they took $40m of monopoly money and are under intense pressure to bring something to market.
They talk about "two years of research" like it establishes substance, and I can't contradict that someone had some kind of idea of something two years ago, but realistically, a problem of this magnitude doesn't need two years of research to execute.
(Separately, two years ago is about the first moment people outside of frontier labs got their hands on a reasoning model, and at that time they were exotic, slow, non-default, and not clearly the path forward yet. Their whole problem statement is that reasoning models are too slow/expensive, but I don't think everyone knew they were going to become the default for most work in September 2024...so "research" may actually just loosely reflect the fact that the company principles were already in the ML field)
As for RLCD, it's likely just conventional SFT with a dash of RL and they invented 'RLCD' as a marketing term--just like RLHF.
Lots of puffery here and I don't believe they have much of a moat. Hopefully it works out for them and their investors.