CITA Helps Tool-Use Agents Choose Better Actions Before Execution
Summary
Large language models increasingly use long sequences of tool calls for complex tasks, but rewards based only on the final outcome provide weak feedback about which intermediate decisions were useful. The paper argues that an agent should estimate the long-horizon value of a possible next tool invocation before executing it. It introduces Comparative Inference for Tool-use Agents (CITA), which trains a Comparative Inference Model from observed tool behavior, scalable supervision generated by a Bayesian tool-graph simulator, and semantic comparisons made by another LLM. The model estimates how likely each candidate invocation is to support final task success in the current context. Across three tool-use benchmarks and multiple backbone LLMs, CITA consistently improves Tool F1 and task success. Further analysis finds that the model also learns accurate step-level value estimates for comparing tool choices.