Large language models are increasingly used as assistants for academic peer review, but fluent unsupported claims can reduce the reliability of reviews. HalluPeer is a benchmark designed for this setting, where verification requires grounding claims in long and technically complex papers. It provides aligned triples consisting of paper content, a human-written review, and a review injected with hallucinations. The data is annotated for hallucination detection, classification, and localization. The authors first derive a peer-review-specific hallucination taxonomy, identify relevant review contexts, and inject hallucinations through a pipeline with automated filtering. Experiments cover 12,000 papers and 38,000 reviews. Existing hallucination detectors struggle to distinguish fabricated or unsupported claims from legitimate critical commentary. Tests on authentic reviews further show that the hallucination patterns defined by HalluPeer also occur in real peer-review text. The findings point to source-aware verification as an important requirement for reliable LLM review assistance. The project is available on GitHub.
AI News
The latest AI releases, research, products, and industry updates.
Loading...