AI Inference & Serving
Multi-Query Attention
An attention variant that retains multiple query heads while sharing one key head and one value head across them. The sharing changes the stored key/value representation, rather than reducing a conversation to one query. It is an architectural choice, not a switch guaranteed to preserve any existing model unchanged.
Reviewed
Sources
Member lesson
The definition and sources are public. The complete practical lesson is for members.