Back to News
RSS feedarxiv.org

New AI Debate Protocol Proves Instance-Optimal for Sensitive Problems

Summary

As increasingly capable AI systems approach or exceed human expert performance on difficult tasks, reliable oversight becomes harder. This paper studies AI debate, in which two AI systems argue over a complex problem and expose simpler claims for limited human supervision. For problems that admit sufficiently stable decompositions into subproblems, the authors propose a protocol that improves on the previous best result in three ways. Its correctness guarantee holds in the worst case rather than only on average. Honesty and correctness form a dominant-strategy equilibrium for both debaters, replacing the weaker Stackelberg-equilibrium guarantee. The authors also prove black-box lower bounds showing that the protocol is instance-wise optimal: no protocol for the same problem class can do better when it can access human judgments only through black-box queries. The results are obtained by connecting stable problem decompositions with fractional block sensitivity from query complexity theory.