RSS feedarxiv.org
LLMs Show Promise and Limits in Disease Model Literature Reviews
Summary
Researchers evaluated an LLM pipeline for systematic literature review across 536 disease-spread agent-based modeling papers. GPT-5.0 achieved 81.67% paper-level accuracy versus 77.95% for GPT-4.1, while field-level performance ranged from 32.40% to 100%. Model agreement may help detect hallucinations, but the study highlights important limits for complex and subjective fields.