Back to News
RSS feedarxiv.org

When Multi-Agent Code Judges Lack Evidence, They Should Decline to Guess

Summary

When one language model judges whether another model's code is correct, its confident verdict may not reflect evidence. This paper studies multi-agent verification, which breaks a judgment into claims and checks them against evidence. The authors argue that useful evidence must be independent of the answer under review and must differ between the candidate solutions; retrieved documents generally meet both requirements, but code-judging evidence may not distinguish the candidates. They run the published MARCH framework unchanged across 80 condition-by-cell measurements on two code-judging benchmarks. MARCH declares both solutions equally good in 78% to 95% of comparisons and reaches only 4.4% accuracy, compared with 43.7% when the same model is asked directly. Easier problems and a larger judge do not resolve the failure. Two measurements from the framework's own logs explain the behavior without requiring labeled data. Gating comparisons with one measurement raises accuracy from 20.7% to 36.9%, while the system still answers about half of the comparisons. The paper's contribution is therefore not a more accurate universal judge, but a label-free way to detect when a code judge lacks a basis for deciding and should decline to guess.