Back to News
RSS feedarxiv.org

When Structured Inference Pays Off for Language Models

Summary

A new arXiv study tests when planning, verification, and repair improve language-model reasoning enough to justify their token overhead. Across 14 budgets, verified search overtakes a single-call monolith between 1,000 and 1,500 output-equivalent tokens and reaches about 44% accuracy at the highest tier, compared with roughly 40% for the monolith.