Multi-agent LLM workflows commonly combine planning, execution, verification, and summarization, but running every component can waste computation or replace a correct intermediate result. The paper formulates component omission as a counterfactual credit-assignment problem: full-workflow logs provide the reward for the executed trajectory, while controlled interventions measure what happens when a future component is omitted. It introduces Learning What to Skip (LW2S), which learns action-specific safety models from these interventions and combines held-out calibration with guards designed for each domain. If an early skip is rejected, the controller can continue the workflow and reconsider a later component. Across mathematical reasoning, multiple-choice question answering, and code generation, evaluated with two instruction-model families, LW2S reduces recorded token cost while matching or improving aggregate full-workflow accuracy in the reported settings. Additional scale-up and second-topology experiments examine whether components are redundant. The study also finds that agreement between components is insufficient for reliable skip selection when they share the same error. Overall, the work frames efficient workflow execution as learning the conditional utility of individual components rather than always executing the full pipeline.
AI News
The latest AI releases, research, products, and industry updates.
Loading...