When one language model judges whether another model's code is correct, its confident verdict may not reflect evidence. This paper studies multi-agent verification, which breaks a judgment into claims and checks them against evidence. The authors argue that useful evidence must be independent of the answer under review and must differ between the candidate solutions; retrieved documents generally meet both requirements, but code-judging evidence may not distinguish the candidates. They run the published MARCH framework unchanged across 80 condition-by-cell measurements on two code-judging benchmarks. MARCH declares both solutions equally good in 78% to 95% of comparisons and reaches only 4.4% accuracy, compared with 43.7% when the same model is asked directly. Easier problems and a larger judge do not resolve the failure. Two measurements from the framework's own logs explain the behavior without requiring labeled data. Gating comparisons with one measurement raises accuracy from 20.7% to 36.9%, while the system still answers about half of the comparisons. The paper's contribution is therefore not a more accurate universal judge, but a label-free way to detect when a code judge lacks a basis for deciding and should decline to guess.
AI News
The latest AI releases, research, products, and industry updates.
Loading...