Metro-WM Uses Real States for Long-Horizon Latent Planning
Summary
Model-predictive control with Joint-Embedding Predictive Architectures can reach goals zero-shot, but its effectiveness declines over short planning horizons. Hierarchical approaches try to extend the horizon by having a macro planner predict latent sub-goals for a micro planner. The paper shows that unconstrained predictions often produce physically unrealisable states. Metro-WM addresses this by retrieving genuine observed frames from offline expert demonstrations or random-action trajectories instead of generating free-form latent vectors. It builds a graph over those frames, connects states across episodes, and searches the graph for routes to the goal. Because planning covers the full graph, the system can replan from the current state when execution drifts. In experiments, Metro-WM improved long-horizon success by up to 37.33 percentage points over the next-best hierarchical method and ran up to 10.9 times faster. It also required 13 to 56 times less offline compute and fewer tuned hyperparameters. Further analysis found shorter paths than the demonstrations, better performance than an oracle using the query's own demonstration, and robust results under extremely sparse datasets.