Back to News
RSS feedarxiv.org

Auditing Pairwise Equivalence Judgments in Multi-Agent Hypothesis Generation

Summary

This study examines two evaluation questions in LLM-based multi-agent scientific hypothesis generation: how much self-critique changes delivered hypotheses beyond ordinary run-to-run variation, and how the definition of equivalence changes measured diversity. Across four proprietary instances, the researchers fixed the opening hypotheses, reran downstream workflows with 0, 1, and 5 critique rounds, and used an LLM judge to score matched pairs. One critique round produced 34.5 percentage points of additional mechanism-level divergence relative to matched same-depth reruns, while moving from one to five rounds added only 1.3 points. The study then compared TF-IDF similarity, dense embeddings, and an LLM judge on controlled pairs that either preserved a causal explanation through wording or biological-terminology changes, or replaced one component of the causal chain. All three methods were invariant to meaning-preserving edits. When the initiating event changed, however, the LLM judge identified 83% of valid pairs as different mechanisms, compared with 0% for TF-IDF and 8% for embeddings. Changing only the rubric for what counts as the same mechanism shifted the LLM judge's result from 38% to 96%. The authors conclude that pairwise equivalence is a measurement choice that affects both estimated self-critique gains and reported hypothesis diversity.