REST Improves Latent Recursive LLM Reasoning
Summary
Large language models can reason in continuous hidden-state space, either by recursively updating their own states or by passing those states between agents. The paper shows that training only the final decoded answer with cross-entropy leaves latent thoughts unconstrained, causing failures such as collapsing representations for different questions and retaining irrelevant information. It introduces REST (REpresentation-Supervised Thoughts), which adds differentiable losses for four desired properties of thought representations: causality, minimality, separability, and stability. REST is applied to both latent single-agent and multi-agent systems without architectural changes or additional inference parameters. Across seven benchmarks covering mathematics, science, medicine, and code generation, the method uses the same training data, compute, and latent budget as CE-only training. It improves accuracy by up to 7.5 percentage points across model sizes and agent settings, and speeds convergence to a final answer by up to 30%. The authors also report that REST thoughts encode more information needed for correct answers and can be decoded more effectively, making latent communication easier to interpret.