AI Inference & Serving
KV Cache
Stored attention keys and values from processed tokens, reused during generation. It avoids repeating some prior computations; it is neither a saved final answer nor a guarantee of persistent conversation memory.
Also called: Key-Value Cache
Reviewed
Sources
Free complete lesson
This definition and the complete practical lesson are free. The interactive reader loads the lesson examples without a membership.