Back to News
RSS feedarxiv.org

Causal Evaluation Helps LLM Agents Recover Selectively

Summary

Large language model agents use external harnesses to pass information between the model and its environment and to recover from execution errors. The paper argues that judging recovery by average task success alone misses an important distinction: an intervention can rescue a failing trajectory, but it can also disrupt a trajectory that would otherwise succeed. It therefore formulates recovery as a causal decision problem by comparing outcomes from the same execution state with and without recovery, separating rescue from harm and examining how intervention value changes over time. The authors introduce the Causal Intervention Router, a lightweight policy that uses information available before recovery to decide whether intervention is worthwhile. On long-horizon ALFWorld tasks with Qwen3-14B, the router raises success from 70.33% to 73.33%, a gain of 3.00 percentage points. It leaves every evaluated trajectory with correct observations untouched. Additional controls indicate that the improvement cannot be explained solely by the new observation returned by the environment, supporting selective recovery as a practical evaluation and control strategy.