DLFP Improves Decode Latency but Fails to Generalize Across Models and GPUs
Summary
Concurrent autoregressive inference creates interference when a newly arrived long prompt is prefilling while existing requests are decoding. The paper introduces Decode-Latency Feedback Prefill (DLFP), a model-free controller that adjusts only prefill work overlapping active decodes. After a guarded scheduling cycle, it uses the observed interval as proportional feedback to resize the next prefill chunk, while leaving isolated prefills unrestricted. Implemented in vLLM, DLFP was evaluated with open-loop Poisson arrivals, exact token accounting, raw request traces, and NVIDIA telemetry. On Qwen3-0.6B in BF16 on one NVIDIA A100 80 GB GPU, three paired 100-request trials reduced P99 inter-token latency by 24.8%, 30.1%, and 28.2%, for a 27.7% mean reduction and a paired 95% confidence interval of 21.0% to 34.3%. Outputs matched exactly, no failures were observed, and SLO compliance was unchanged. The trade-off was a 34.8% increase in mean P99 time to first token, although it remained within the declared SLO. The controller did not generalize to Qwen3-8B, Qwen3-32B, or a two-GPU tensor-parallel setup. The authors attribute this boundary to an asynchronous scheduler-call interval that only approximates completed GPU iteration time, and propose completion-timed control for future work. They present the result as a reproducible proof of concept and generalization study, not evidence of mobile-device performance.