Back to News
RSS feedarxiv.org

Cognitive Diversity Does Not Explain Multi-Agent Debate Gains in Small Language Models

Summary

This study tests whether cognitive diversity is the main reason multi-agent debate improves reasoning and factuality. The researchers evaluate 23 small open-weight models from 11 vendor families across five tasks and more than 5,500 debate and control runs. They vary diversity through personas, sampling temperature, and model identity, pairing each debate setup with a majority-vote control matched for generation budget. The diversity hypothesis is rejected on all three axes. Debate improves over single-agent inference by 3–7 percentage points on tasks with room for improvement, but it ties or loses to self-consistency sampling while taking 1.6 times the wall-clock time and 3.4 times the token cost. Persona prompting lowers accuracy; experiments across the full persona combinations indicate a persona tax, with redundant personas causing the greatest damage and highly diverse teams recovering only part of the loss. Mixed-model teams also underperform majority votes over the same models, with accuracy following member capability rather than heterogeneity. Most debate gains arrive after the first exchange. The study additionally finds that silent context-window overflow in debate transcripts can distort comparisons: correcting it changes one debate-versus-sampling result from a 1.8-point deficit to parity. The authors therefore characterize reported debate gains primarily as an ensemble-sampling effect and propose budget-matched, contamination-checked baselines for future work.