Small-message JACCL all_sum latency microbenchmark protocol
Parent: Mac local LLMs: Clusters, RDMA, exo and ds4 · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Frames are 4096 bytes and pipeline depth is 2, so an 8 KB message is two frames and a 32 KB message eight; latency steps with frame count, not with bytes.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Frames are 4096 bytes and pipeline depth is 2, so an 8 KB message is two frames and a 32 KB message eight; latency steps with frame count, not with bytes. [source]
- `all_sum` on an MLX array is lazy and goes through a distributed primitive on a stream. A timing loop must call `mx.eval` on the result inside the timed region, otherwise it measures graph construction. [source]
- `barrier()` is itself an `all_sum` of one byte, so a barrier between iterations is a latency probe of the same path and can be used to align ranks. [source]
- Sizes to sweep: 1 B (barrier), 4 KB (one frame), 8 KB, 16 KB, 32 KB, 32 KB plus 4 B, 64 KB, 512 KB minus 4 B, 512 KB. This crosses the mesh algorithm switch and the multi-wire switch and the ring 32768-byte switch. [source]
- Warm up past the first call: the first call after init builds lazily allocated state in MLX and the first UC exchange follows the warm-up barrier; discard at least tens of calls. [source]
- A 2-rank mesh all_reduce of 512 KiB or less runs one wire and one-phase [source]
- A mesh above 2 ranks switches to reduce-scatter plus all-gather above 32 KiB per wire [source]
- JACCL frames are 4096 bytes with pipeline depth 2 [source]
- `MeshGroup::barrier` is a one-byte `all_sum` and can be used as a one-frame latency probe [source]
- A latency sweep should straddle 32 KiB, 32768 B and 512 KiB, measure each of mesh and ring and each node count, call `mx.eval` inside the timed loop and verify values on the last iteration [source]
- The MLX docs say JACCL gives communication latency an order of magnitude lower than the ring backend, with no number [source]
Children
- No children recorded.