Rationale Messages Can Distort AI Verifier Support Judgments
Summary
The study examines what rationales communicate when a role-specialized question-answering pipeline passes a reasoner’s message to a verifier. Instead of changing the evidence or candidate answer, the researchers fix both and vary only the rationale transmitted across the boundary. They test 400 examples from MuSiQue, HotpotQA, and 2WikiMultiHopQA, using DeepSeek as both generator and verifier. Faithful rationales add almost no answer accuracy compared with providing no rationale, but corrupted rationales substantially change the verifier’s assessment of whether an answer is supported. With a blind verifier prompt, harmless paraphrases change support judgments by only 0-2.5%, while corrupted rationales produce shifts of 10-22%. Asking the verifier to check the rationale amplifies the corrupted-rationale effect to 34-55%, suggesting that an explicit checking instruction can make the message channel more influential rather than more reliable. Final answers are less affected, moving by 2-30%, and only 2.9-35.3% of corrupted support flips occur together with an answer change. Human audits help explain the discrepancy: in 16 of 42 valid corruptions, the model overtrusted the corrupted rationale, whereas human reviewers rejected or marked unclear 9 of 10 such rationales accepted by the model. Cross-model and task-boundary checks further find that the channel can be active, amplified, inert, or absorbed into the task label. The authors therefore argue that rationale sharing should be evaluated as a verification-message mechanism, including its failure modes, rather than only by whether it improves answer accuracy.