AI Inference & Serving
Activation Quantization
Representing intermediate model activations with a restricted numerical format or discrete levels. This differs from quantizing only stored weights. Activation distributions depend on inputs, and unusual values can create clipping or rounding errors; calibration choices and supported execution paths must be evaluated rather than inferred from a smaller bit width.
Reviewed
Sources
Member lesson
The definition and sources are public. The complete practical lesson is for members.