Back to News
RSS feedarxiv.org

Choosing the Right Backbone Layer for Vision-Language-Action Policies

Summary

Vision-language-action (VLA) policies connect a pretrained vision-language backbone to an action head through a latent interface, but the best backbone layer is not obvious. This study evaluates single-layer selection and multi-layer fusion with frozen backbones from three pretrained models on the LIBERO and CALVIN manipulation benchmarks, using three training seeds per configuration. Across three fusion mechanisms and three layer-subset strategies, 47 of 54 configurations perform worse than the best single-layer policy, indicating that fusion offers only limited benefit. The best layer also changes substantially across backbones and benchmarks, making exhaustive policy sweeps costly. The authors derive a reweighting equivalence between information-bottleneck objectives for action-conditioned InfoNCE and action prediction, supporting InfoNCE as a proxy for layer quality. Among four evaluated proxies, InfoNCE shows the most consistent positive association with policy success. Choosing the highest-InfoNCE layer uses 9-33 times less GPU compute than exhaustive sweeps and lowers mean selection regret from 17.89 percentage points for choosing the deepest layer to 3.71 points across six settings. This is close to the 3.28-3.50-point regret of retrospective fixed-layer heuristics optimized with oracle sweeps, without requiring closed-loop evaluation during selection.