Predictive Credit Tests What Scientific Explanations Add to Experimental Forecasts
Summary
The paper introduces predictive credit, a protocol for measuring what scientific explanations contribute to forecasts of planned experiments. Paired forecasts hold the intervention, forecaster, and outcome constant while varying the description, a matched explanation, and donor context. Five checks assess commitment, delivery, predictive gain, alignment, and uptake of known signals. The authors evaluate 336 prospective states in controlled learning, 12 Tox21 endpoints, and 24 OpenML tasks, but the frozen version-5 credit decision is inconclusive. Tox21 did not meet its preregistered ROC AUC interval-score harm criterion: the reported difference was -0.0026, with a 95% interval from -0.0174 to 0.0104. OpenML also failed its joint requirements for formation, point equivalence, and repeatability, while matched point-accuracy gains and seed-donor intervals remained unconfirmed. Under the requested DeepSeek V4 Pro setting, matched and donor cards reduced secondary Tox21 drift by 64.5% and 59.1%, respectively. A DeepSeek V4 Flash replay instead increased matched point MAE from 0.01823 to 0.02020 and failed matched-donor interval-score equivalence. On OpenML, full-card assignment widened nominal 80% intervals by 21%, with 49.3% coverage versus 51.4% for description, and content appeared in 66 of 144 cards. Direct-text Flash delivered all 144 notes without detectable matched point-accuracy improvement. A researcher-authored mechanism positive control reduced point MAE by 2.60 percentage points, showing that the protocol can detect a specified signal even though natural-explanation credit was not confirmed at the tested donor resolutions.