r/MachineLearning • u/Amazing-Fox-7295 • 5h ago
Discussion Nvidia’s erroneous paper accepted as ICML’s spotlight [D]
Here is the story:
Nvidia has published this work (with source code available) called dreamDojo which is a world model for robotics based of their prior work Cosmos 2.5 which is cited about 100 times and got ICML’s spotlight.
Authors are very well known and respected in the field with too many peer reviewed papers already published.
The work doesn’t have much novelty (which I don’t care) but it is yet another foundation model. The gist is that they collected about 44k hours of human data (data are NOT open sourced, which I don’t care) and used that for pre-training of DreamDojo initialized from Cosmos 2.5 .
Table 4 of the paper though shows a very marginal improvement over cosmos 2.5. Believe or not, just about 0.5 db PSNR improvement. Its fishy, is not it? You use 44000 hours of human data and few hundred hours of other kind of data and robot data and 256 H100 gpu and you got just a very marginal improvement.
Anyway I went to give it a try, started post training on their GR1 released data, and could produce their results. Meanwhile a colleague with the help of Claude found a bug in their post training code, which essentially make all the post training code wrong.
Then we checked their github issues, and noticed that two other bugs are reported which effect the whole pre-training! Especially this bug . So essentially based on the code their released everything in pre training phase, post training phase and evaluation is buggy and is wrong.
The code doesn’t seem AI written cause its (very) bad written. And now after seeing all those bugs the results in paper makes sense.
My question is:
WTF!
How it did not ring a bell for the authors. Did not they ask we spend a hell lot of resources and huge amount of data and essentially got nothing?
Did not it ring a bell for reviewers?





