RSS feedarxiv.org
Study Finds Exact-Match Scores Misjudge Tool-Using Agent Reliability
Summary
The study evaluates five open-weight models on contamination-free multi-step tool-use tasks. It finds that early mistakes cause roughly 70% of clean-context capability to be lost by depth six. The authors show that fixed-trajectory exact-match scoring structurally forces extreme propagation estimates and hides recovery, then propose conditional-on-state scoring as a practical remedy.