AI Inference & Serving

Expert Parallelism

Distributing the expert components of a mixture-of-experts model across devices and routing token work to them. Expert placement and token routing differ from simply copying a complete dense model.

Also called: EP

Reviewed

Sources

Member lesson

The definition and sources are public. The complete practical lesson is for members.

Compare membership plans ยท Already a member? Sign in

Explore all dictionary definitions