Enabling Dynamic Computation in Looped Language Models
Summary
Looped language models can reuse parameters across iterations, potentially allocating less computation to easy tokens. The paper argues that current open looped models, including Ouro, fail to realize this benefit because every loop depth requires its own KV cache, forcing all computations to be retained. It proposes a “best-available” KV caching strategy that works across loop depths and expands the performance-versus-depth trade-off. On the evaluated models, the method reduces FLOPs and KV-cache memory by up to 30% while preserving full-depth performance. Training models with awareness of this caching strategy further improves performance and efficiency. The authors also modify the early-exit prior objective: instead of pushing every token toward similar lower depths, the revised objective allows tokens to exit at heterogeneous depths according to their required effort. The findings are validated on Ouro models and on smaller looped language models pretrained from scratch.