SciLitBench Benchmarks LLMs for Systematic Literature Reviews
Summary
The paper introduces SciLitBench, a multi-stage benchmark for evaluating how large language models support systematic literature reviews. It covers title and abstract screening, full-text screening, and schema-guided data extraction, using 42,981 retrieved records, 1,012 full texts, and annotations for 888 included papers. The authors evaluate 22 open-weight LLMs from six model families. Providing explicit inclusion and exclusion criteria improves title-and-abstract screening F2 by 28.8%, while researcher-authored rationales improve full-text screening by 15%. Extraction is less reliable and varies sharply by field: publication-year accuracy reaches 0.97, whereas computational-approach extraction reaches only 0.37 Jaccard overlap. Even the strongest models recover only 30% of annotated evaluation evidence and 25% of limitations. The benchmark therefore identifies a practical divide between screening that prioritizes recall and extraction that must produce complete evidence, while offering a reproducible resource for evaluating LLM-assisted evidence synthesis.