AI Inference & Serving
Tensor Parallelism
Splitting tensor computations or model parameters within layers across participating devices. The devices cooperate on one model execution; this differs from independent replicas serving separate requests.
Also called: TP
Reviewed
Sources
Member lesson
The definition and sources are public. The complete practical lesson is for members.