AI Inference & Serving
Weight-Only Quantization
Reducing the precision of stored model weights while distinguishing their format from activation and computation formats. A low-bit weight label does not mean every operation or runtime tensor uses that same precision.
Also called: WOQ
Reviewed
Sources
Member lesson
The definition and sources are public. The complete practical lesson is for members.