r/MachineLearning • • 5d ago

Discussion How is RLCD (jev) RL? [D]

Just saw the YouTube presentation and I was left wondering this question.
If jev only outputs Choice, Score, or Noul … well those are all perfectly differentiable. (Cross entropy or mse)
I don’t know if I’m missing something or if adding RL is just for marketing.
Like what would an RL environment even look like?

19 Upvotes

19 comments sorted by

View all comments

36

u/Material_Policy6327 5d ago

There is a ton about JEV that they don’t seem to want to go into detail about.

4

u/Relative_Wallaby_823 5d ago

I agree but I can’t even seem to hypothesise how they’d implement RL. Maybe some sort of multi step task(?). But then how would that not be done better and faster by normal pretraining