Measuring the Sufficient Compute Needed for Each Generated Token
Summary
Large language models typically spend the same amount of computation on every generated token, even though tokens may require different levels of effort. This paper estimates each token’s sufficient compute with a Mixture-of-Agents panel of 15 language models from three families, where the smallest agent that reproduces the reference token defines the estimate and provides an upper bound on the token’s requirement. On three core benchmarks, a 0.5B agent reproduces 92–95% of reference tokens. Across Qwen, OLMo, and R1-distilled panels, the most expensive 10% of tokens account for 64–80% of estimated FLOPs. On all 500 MATH-500 problems, routing based on the resulting compute map reduces projected latency from 7.59 to 5.12 seconds while slightly improving accuracy over the best confidence-routing baseline. The map also enables drafting with 32.6% fewer draft tokens and about 20% lower projected latency than fixed-window drafting at similar accuracy. The findings indicate that token-level compute allocation retains exploitable headroom and could inform controllers that assign computation according to sufficient-compute structure.