Beyond Imitation: A Framework and Benchmark for LLM-Assisted Peer Review
Summary
The rapid growth of scientific publishing has increased reviewer workload and raised concerns about review quality, especially in machine learning. This study argues that LLM-assisted peer review should be evaluated on verification tasks, rather than mainly on how closely generated reviews resemble human-written reviews. It introduces a scalable benchmark that inserts synthetic errors into conference papers, creating clear targets for testing whether systems detect logical contradictions. The authors also propose a Multi-Layered Review framework that emphasizes detailed manuscript comprehension before generating a review and is designed to use tokens more efficiently. Across the reported evaluations, the approach aligns strongly with human review scores, detects errors effectively, and offers perspectives that complement human reviewer focus. Results depend on both the underlying LLM and the surrounding system design. The study also confirms that automated review systems remain vulnerable to adversarial manipulation, so robust, error-focused evaluation is needed before deploying them in peer review and other high-stakes scientific workflows.