Evaluating the evidence trace of NeurIPS 2025 contributions with agentic code auditing

Pascal Iversen · Ferdous Nasri · Bernhard Y Renard · Katharina Baum

Video

Paper PDF

Thumbnail of paper pages

Abstract

Code serves as the primary evidence behind computational publications, yet a detailed review of an unfamiliar codebase imposes a commonly prohibitive time burden on volunteer reviewers. Consequently, self-reported reproducibility checklists at major machine learning venues are rarely empirically verified. This leaves code as a significant blind spot in peer review. To bridge this gap, we introduce AuditOwl, an autonomous, verification-centric LLM pipeline designed to make code auditing feasible for authors pre-submission and reviewers post-submission. With this framework, we conduct an audit of 100 randomly sampled empirical papers from the NeurIPS 2025 main track. For each paper, an LLM agent inspects the repository and evaluates its scientific claims against the evidence in the underlying code. Following an independent adversarial verification pass to maximize precision, the system raises 609 discrepancies across the 87 papers with retrievable code (median 7 per paper), of which 388 are high or medium severity. It produces evidence that can be quickly checked by humans: findings cite specific code and paper locations, and about half are backed by executable verification checks that the agent implements. Eleven ML researchers evaluated 120 such findings for 30 paper-code pairs and judged 1.7% incorrect and 10.0% incorrect or trivial. Our approach reveals a steep reproducibility funnel: only 87% of papers in our sample release code at all, and for just 8% AuditOwl locates producing code for every traced result. Discrepancies we find are heavily dominated by incompleteness of code and mismatches between what the paper describes and what the code does. We also detect a prevalence of technical bugs and serious methodological issues. Operating at reasonable cost, agentic code auditing can augment human peer review and help to make computational science more reproducible. Code available at: https://github.com/PascalIversen/auditowl-neurips2025.