SciSlopBench Targets AI-Generated Scientific Papers Through Global Reasoning
Summary
The article examines the rapid growth of AI-assisted scientific submissions and a research collaboration from Seoul National University and the University of Minnesota that proposes measuring and mitigating “scientific slop.” arXiv received 40,363 submissions in September 2026, nearly twice its September 2024 total, and moderators reported more thin, fragmented and densely AI-written papers; arXiv responded with a two-submission-per-month limit and had previously introduced other restrictions on unchecked AI content and unsupported survey or position papers. The featured paper, “Science or Slop?: Benchmarking and Mitigating Scientific Slop in AI-Generated Papers,” argues that prose-level AI detectors miss failures in the relationships among claims, reasoning, citations, evidence and figures. It defines six measurable patterns: weak cross-section references, macro redundancy, broken argument graphs, citation isolation, poor figure exposition and evidence gaps. Its SciSlopBench pairs 390 AI-generated papers with comparable human-written papers, drawing from FARS and the Agents4Science competition across several scientific fields, with a heavy computer-science bias. On this dataset, the benchmark achieved 85.9% PairAcc, ahead of Binoculars at 68.7% and the strongest automated reviewer at 68.5%; cross-section references alone reached 90.5% PairAcc and detected 65% of AI papers at a 5% human false-positive rate. Higher slop also correlated with lower ICLR ratings and helped distinguish accepted from rejected papers across nine years, suggesting that the measure captures quality problems as well as AI authorship signals. SciSlopHarness then used the detected weaknesses to guide evidence-grounded revisions, rejecting unsupported additions and reducing the AI-human gap by 63% versus the strongest revision baseline, Claude Code. The paper’s authors also provide a live analysis demo, which gave their own paper a moderate score of 40/100; the authors explicitly describe their own AI use.