AI Inference & Serving
Expert Parallelism
Distributing the expert components of a mixture-of-experts model across devices and routing token work to them. Expert placement and token routing differ from simply copying a complete dense model.
Also called: EP
Reviewed
Sources
Member lesson
The definition and sources are public. The complete practical lesson is for members.