RESEARCH

Decode-Latency Feedback Prefill: A Model-Free Controller and Its Generalization Limits

ArXiv cs.AI · Thu, 01 Oct 2026 04:00:00 GMT

arXiv:2609.38386v1 Announce Type: new Abstract: Concurrent autoregressive inference creates a fundamental interference problem: prefilling a newly arrived long prompt can delay tokens for requests that are already decoding. Fixed prefill chunks reduce this interference, but the b

Read original source Discuss with SiiMON