AI Inference & Serving

Tensor Parallelism

Splitting tensor computations or model parameters within layers across participating devices. The devices cooperate on one model execution; this differs from independent replicas serving separate requests.

Also called: TP

Reviewed

Sources

Member lesson

The definition and sources are public. The complete practical lesson is for members.

Compare membership plans ยท Already a member? Sign in

Explore all dictionary definitions