What We Learned by Reproducing 2,200 papers from ICML
Blog post from Hugging Face
A 19-day ICML 2026 Open Reproductions hackathon enlisted 1,221 participants using coding agents to produce 6,816 auditable logbooks examining 2,226 accepted conference papers and 35,908 extracted claims. Automated judging found at least one verified claim in 51% of examined papers, including 266 fully reproduced papers, while 23% had at least one falsified or contested claim and 242 produced conflicting results across independent teams; missing artifacts and limited-scale tests accounted for many remaining inconclusive cases. Confirmed issues included a flawed robustness theorem, delayed counterexamples missed by short experimental horizons, a mismatch between theoretical and implemented loss functions, and evaluation distortion caused by padding tokens, although some claimed falsifications were themselves disproved after review. Authors contacted about findings had confirmed several results and begun corrections, while the project argues that agent-based replication can help address growing publication volumes but still requires human oversight to identify flawed assumptions, interpret long-run behavior, and assess qualitative outcomes.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.