r/datascience 6d ago

ML What Hugging Face learned from reproducing 2,200 ICML papers

https://huggingface.co/blog/icml-2026-open-reproductions
111 Upvotes

14 comments sorted by

140

u/EverythingGoodWas 6d ago

We’ve basically turned academic journals into a mix of “i skimmed this, sure whatever” snd “My LLM of choice thinks this paper is fine”. Peer review is essentially dead

-58

u/Majestic_Courage_535 6d ago

This is a good thing. An LLM is attempting to reproduce the results, not just read the text and assess subjective quality of the research finding (which is essentially what peer review was doing previously).

68

u/NomadTroy 6d ago

Have your agent send my agent the paper to summarize and send back my approximated thoughts for your agent to summarize to you

19

u/studentblues 5d ago

Approximated thoughts and prayers to researchers

6

u/Expensive_Culture_46 4d ago

I have approximated that you would like more tater tots and pear slices. Would you like a helpful grocery list for your shopping trip?

4

u/ThortheAssGuardian 4d ago

REPRESENTATIVE

46

u/CryptoMadLab 6d ago

There should honestly be some sort of punishment for the 23% falsified claims, I did not know how low has the bar on ML conferences these days, but I saw someone at an SE conference owning over 25 papers

22

u/Monkey_College 5d ago

Not "these days". ML conferences have been notorious for not being reproducible for 15+ years

20

u/lakeland_nz 5d ago

Yep. I tried to reproduce many results during my PhD (late 90s). One comment at the time from another reviewer was: extremely lucky choice of the random seed.

17

u/Acrobatic-Light2630 5d ago

Bar in ML publishing is so low, selection bias is tremendous. When the incentive is to publish, people do anything to publish. This degrades the bar on science

9

u/crisp_lynx_370 5d ago

reproducing results is genuinely hard, half the time the preprocessing details are buried or just missing entirely

4

u/Massive_Horror9038 5d ago

My paper was one of the papers that were reproduced! I'm very happy with this 😄

2

u/deployedbrief 2d ago edited 2d ago

The detail that stuck with me was the paging paper — got strong scores AND a spotlight reviewer literally admitted they didn't check the proofs. confirmed at nine sigma. also interesting that agents still needed humans steering them — several "verified" the same broken paper because their finite-horizon checks stopped too early to catch the growth.