AiPhreaks ← Back to News Feed

What We Learned by Reproducing 2,200 papers from ICML

By Jakub Antkiewicz

2026-08-14T09:06:32Z

Community Audit Reveals Reproducibility Issues in Top-Tier AI Research

A massive community hackathon organized by Abubakar Abid of abidlabs has cast a new light on the reproducibility of research from the premier ICML 2026 conference. Over 1,200 participants used AI coding agents to reproduce 2,226 published papers, finding that 23% had at least one claim that was either falsified or contested. The effort, which generated over 6,800 detailed logbooks, suggests that the AI-driven explosion in paper submissions is outpacing the capacity of traditional human peer review, creating a significant gap in scientific validation that AI itself may be required to fill.

The Verdict on 2,200 Papers

From July 15 to August 2, the ICML 2026 Open Reproductions challenge used an open process where participants brought their own coding agents to verify specific claims from indexed papers. Each attempt was documented in a Trackio logbook, hosted on Hugging Face, and assessed by an automated judge using the GLM-5.2 model. The large-scale audit produced a detailed, public dataset of the conference's scientific integrity.

  • Papers Attempted: 2,226 (34% of the entire conference)
  • Claims Verified: 51% of examined papers (1,103) had at least one claim independently verified.
  • Claims Falsified: 23% of examined papers (496) had at least one claim falsified or contested.
  • Inconclusive: The remainder had only toy-scale evidence or could not be reproduced due to missing artifacts.

The project found several high-profile issues, including an algorithm with incorrect asymptotic complexity, a theorem that failed after thousands of steps, and evaluation metrics diluted by incorrect data handling. While some falsification claims were later found to be bugs in the reproduction code itself, many were confirmed by the original authors, with several corrections already submitted to arXiv. This demonstrates a shift in the role of AI reviewers, not to fully replace humans, but to augment them. The most effective reproductions involved a human guiding an agent, setting experimental direction, and interpreting results that required perceptual judgment, much like a principal investigator managing a lab.

The future of peer review is not full automation, but a human-on-the-loop system where researchers act as principal investigators, directing fleets of AI agents to conduct scalable, adversarial verification of scientific claims.
End of Transmission
Scan All Nodes Access Archive