AI Inference & Serving
Mixture of Experts
In this inference context, sparse MoE routes tokens through selected expert components rather than activating every expert for every token. These experts are neural-network components, not independent people or chat agents. The broader mixture-of-experts family also includes other designs.
Also called: MoE
Reviewed
Sources
Free complete lesson
This definition and the complete practical lesson are free. The interactive reader loads the lesson examples without a membership.