Back to News
RSS feedarxiv.org

ProWAM Uses Progressive Visual Sub-Goals for More Robust Robot Control

Summary

World action models (WAMs) jointly predict future visual dynamics and actions for robotic control, but long-horizon tasks are difficult because dense video rollouts are expensive. ProWAM addresses this limitation by predicting actions together with an ordered sequence of sparse visual sub-goals, giving the policy progress-indexed visual guidance during execution. The model can learn sub-goal prediction from large action-free video collections, while the video backbone handles visual planning separately from the action policy. For replanning, ProWAM performs one video-backbone forward pass to cache sub-goal features and then uses lightweight action denoising instead of repeatedly generating full videos. In evaluations, it reaches 85.8% on LIBERO-Plus and 75.7% on randomized RoboTwin, with relative gains of up to 35.9% over the strongest baseline. On RoboCasa365, it records 48.1% success overall and 18.2% on the Composite-Unseen split, ranking fourth overall. In zero-shot real-world experiments on novel scenes, it achieves 70.0% success, compared with 55.0% for the strongest baseline, a 15.0 percentage-point and 27.3% relative improvement. The results support progressive visual foresight as a way to improve closed-loop robotic control and out-of-distribution robustness.