Study Finds AI Reviewers Complement, Not Replace, Human Scientists
Summary
A large expert-annotation study examined what AI reviewers do well and where they fail in scientific peer review. Forty-five scientists from the physical, biological, and health sciences spent 469 hours rating 2,960 individual criticisms from human-written and AI-generated reviews of 82 Nature-family papers for correctness, significance, and evidentiary sufficiency. On a composite of those dimensions, a reviewing agent powered by GPT-5.2 scored above each paper’s top-rated human reviewer, 60.0% versus 48.2% (p = 0.009). Gemini 3.0 Pro and Claude Opus 4.5 were also part of the comparison, and all three AI reviewers exceeded the lowest-rated human reviewer on every dimension. Accurate AI criticisms were more often judged significant and well evidenced, and AI systems identified 26% of issues that no human reviewer raised. However, AI reviewers overlapped much more with one another than human reviewers did, 21% versus 3% for cross-reviewer pairs. The study also identified 16 recurring weaknesses absent from human reviewers, including limited subfield knowledge, difficulty managing long context across multiple files, and excessive criticism of minor issues. The authors therefore position current AI reviewers as complements to human judgment rather than substitutes.