AI Inference & Serving
Chunked Prefill
Processing a long input in smaller scheduling pieces so its prefill work can be interleaved with other work, including ongoing decoding. It changes scheduling granularity, not the input's meaning or context limit.
Reviewed
Sources
Member lesson
The definition and sources are public. The complete practical lesson is for members.