Back to News
RSS feedarxiv.org

GAVA Helps Embodied Agents Arbitrate User Corrections with Evidence

Summary

The paper studies how a text-based embodied agent should respond when a user's correction may be wrong. It formulates grounded correction arbitration as a choice among accepting the correction, rejecting it, inspecting the environment, or asking the speaker. GAVA makes this choice using observation-bounded evidence, legal probes, and a one-step expected-loss rule. In text-only ALFWorld, the evaluation used 162 checkpoints and 972 paired true and false interventions. Complete local inspection gave GAVA and an always-verify baseline 100% correction accuracy, establishing the evidence contract rather than a comparative advantage. During same-episode execution, GAVA reduced interaction cost relative to always verifying, while tying the cost threshold when the speaker was perfect. An exploratory object-location prior, trained only on training data, reduced interaction cost and declared joint cost relative to uniform GAVA by 0.490 and 0.420 on 340 unseen scenarios. After the policy, costs, baselines, and multiplicity plan were frozen, the gains replicated on 77 non-overlapping seen checkpoints covering 308 scenarios, with reductions of 0.595 and 0.517; both 95% checkpoint-bootstrap confidence intervals excluded zero. Joint cost also improved over an identical-prior fixed policy, but a matched calibrated comparison without value-of-information selection was inconclusive. Semantic GAVA made four factual errors in each cohort, yielding 98.8% and 98.7% accuracy, and every method completed every task. The results support selective information gathering with semantic priors under explicit costs, but do not establish a general advantage for environmental value of information over clarification. The study uses normalized claims, complete symbolic observations, and controlled speakers, and does not evaluate human participants, visual input, or physical robots.