Plan-and-Patch Uses Diffusion Language Models to Repair Agent Plans
Summary
Long-horizon agents must coordinate subgoals, tool use, and intermediate outcomes, but assumptions can become invalid when environments or tools behave unexpectedly. Plan-and-Patch addresses this by having a diffusion language model generate a structured, program-like plan through parallel unmasking and regenerate only an affected region while preserving the plan's prefix and suffix. The study compares DreamReasoner-8B and Qwen3-8B as diffusion and autoregressive planners. On Natural Plan without task-specific training, the diffusion approach achieves a 53.7% plan-repair success rate, nearly twice the autoregressive result of 27.0%. After task-specific training on the ALFWorld and TextCraft agentic benchmarks, the two approaches achieve similar observed success in plan generation. Diffusion nevertheless reduces mean plan-generation latency by 39-46% relative to autoregressive planning. The authors present these results as evidence that local plan repair can avoid unnecessary changes while supporting faster planning for long-horizon agents.