Draft-verify-revise is an inference-time orchestration pattern in which one LLM writes a draft, another critiques it, and a third revises it. The study examines a failure mode that can arise as context moves between stages: different models may resolve a context-dependent expression such as “previous” to different referents, creating a deictic shift. Using a synthetic dataset of 10 base examples rendered under three conditions, the researchers varied whether the draft or verifier correctly resolved the expression and how much independent reasoning the reviser needed. Six models from three providers were evaluated across 21 reasoning-effort configurations in a primary experiment and an ablation that removed error-classification labels from grader feedback. Sequential testing used e-values, and a separate LLM analyzed the stated rationales behind incorrect verdicts. Balanced accuracy ranged from 0.156, below chance, to nearly perfect. GPT-5.2 improved from 0.156 without reasoning to 0.942 at its highest tested effort, while Gemini 3 Pro remained above 0.94 at every effort level. Gemini 3 Pro at low effort exceeded GPT-5.2 at xhigh effort for about 5% of the cost per trial. When revisers were wrong, they generally relied on surface cues instead of operational reasoning. The authors therefore advise people designing these pipelines to make the intended referent explicit at every stage.
AI News
The latest AI releases, research, products, and industry updates.
Loading...