Benchmark Evaluates LLM Understanding of Middle and High School Science
Summary
Large language models are increasingly used in education, but existing evaluations provide limited evidence about their performance on content aligned with secondary school science curricula. This study introduces a comprehensive benchmark covering middle and high school science topics under the Next Generation Science Standards (NGSS). The authors construct the benchmark through a synthetic data pipeline, validate items with multiple judges, and apply item-level psychometric analysis. They evaluate nine open-weight LLMs across different science domains and question types, finding that several smaller models that can be deployed locally achieve high accuracy. Model size is not a consistent predictor of performance, so educational deployments may require deliberate model selection rather than simply choosing the largest model. A human reviewer also examined LLM-generated items for NGSS alignment and found that synthetic items were not perfectly aligned with the standards. The authors therefore argue for human involvement in item development and for closer attention to both subject-matter and pedagogical knowledge. They identify interactive, evidence-based feedback as an important extension for future benchmarks, since the current evaluation focuses on science understanding rather than the full behavior of an educational assistant.