Tensor-parallel all-reduce count per transformer layer in MLX shard_linear

Parent: Mac local LLMs: Clusters, RDMA, exo and ds4 · Published reference · snapshot 2026-10-05

↓ Facts as markdownall context files

A tensor-parallel all-reduce in MLX is an `mx.distributed.all_sum` on an activation. The count per transformer layer equals the number of `ShardedToAllLinear` (or quantized equivalent) calls made by layers sharded with `shard_linear`, plus the explicit `all_sum` calls that model code adds around ...

These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.

Facts

Children

← the whole tree · 3D view· how to read this page