Competence-Gated Pooling Improves Selective Use of Language Models in Event Forecasting
Summary
The paper examines a practical problem in hybrid forecasting: when a language model is one of several signals, should it influence an existing market, crowd, or statistical forecast? It argues that standalone model accuracy is the wrong target and instead defines competence as the model’s marginal value beyond the external forecast. Under Brier loss, the authors characterize when model disagreement can improve that forecast and derive the benefit of domain-specific pooling weights over a single global weight. They introduce a competence gate that estimates source weights from resolved outcomes, shrinks uncertain domain estimates toward a global weight, and recalibrates the combined forecast. Across 2,357 resolved binary questions and five language models, the gate reduced the main external baseline’s Brier score from 0.0771 to 0.0732 and significantly outperformed global forecast combinations. The improvement remained significant with leakage controls and against a leakage-safe time-series prior on the pooled structured dataset, with separate supporting evidence from FRED. Results differed on the official ForecastBench market subset: the gate produced no significant improvement and mostly deferred to the market. Tests across four Qwen models also found that verbal confidence did not reliably indicate when a model would outperform the external forecast, whereas outcome-estimated competence supported better abstention decisions. The authors present the method as a way to decide when to use a model, rather than assuming that every available model signal should affect the forecast.