Per-token collective count times per-call latency model

Parent: Mac local LLMs: Clusters, RDMA, exo and ds4 · Published reference · snapshot 2026-10-05

↓ Facts as markdownall context files

Per-token synchronisation time on a tensor-parallel Mac cluster is modelled as `T_sync = G x t_gate`, where `G` is the number of cross-machine exchanges per decoded token and `t_gate` is the time one exchange adds to the critical path.

These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.

Facts

Children

← the whole tree · 3D view· how to read this page