AI Inference & Serving
KV Cache Quantization
Storing cached attention keys and values in a lower-precision format. This targets sequence-state memory separately from weight quantization; scale choices, attention implementation and quality effects need their own checks.
Reviewed
Sources
Member lesson
The definition and sources are public. The complete practical lesson is for members.