Back to News
RSS feedarxiv.org

GPS-Bench: A Governance Policy Benchmark for Automating Policy Analysis

Summary

GPS-Bench is an evidence-grounded benchmark for studying whether large language model simulations can predict and explain governance policy outcomes. It links policies to affected actors, their actions, and downstream impacts using legislative records, lobbying disclosures, regulatory documents, corporate filings, economic data, and other public evidence. Actors are reconstructed from dated records rather than prompted as generic archetypes, making each persona an evidence object with provenance. The benchmark uses a human-annotated Gold evaluation set, while cases labelled by a separate language model from retrieved evidence provide Silver supervision and are excluded from testing. Because every inference method receives the same grounded policy state and produces the same schema, the benchmark supports controlled comparisons of joint reasoning, independent or communicating actor agents, graph-based methods, and weight-level fine-tuning. Fine-tuning on the grounded record produces the strongest actor-level impact prediction in the reported comparison, while decomposition does not improve that prediction. Multi-agent decomposition contributes a different value: actors hold private, non-identical evidence and negotiate with named partners over concrete proposals, offers, requested returns, and reasons for cooperation. These interactions make coalition formation and its underlying commitments available for checking against the documentary record. The authors present GPS-Bench as a common empirical setting for testing when evidence, actor modelling, and multi-agent interaction improve both prediction and interpretation of policy outcomes.