SAUCE Estimates Uncertainty in Parallel Multi-Agent Reasoning
Summary
Parallel multi-agent reasoning systems have several language-model agents solve the same problem over multiple rounds and aggregate their answers, but estimating the reliability of the final system remains difficult. The paper introduces SAUCE, or Sequential Agent Uncertainty through Consensus Evolution, a lightweight, training-free estimator that treats system uncertainty as sequential inference over a latent system-level belief. It updates that belief by combining agreement between agents at each round with signals about uncertainty in their generations. Across five model backbones, five benchmarks, and two multi-agent protocols, SAUCE improves misclassification detection, selective prediction, and calibration compared with a broad range of baselines, including standard log-likelihood methods and estimators designed specifically for multi-agent systems.