<!-- llms-explorer concept facts · https://llms-explorer.com/tree/small-message-jaccl-all-sum-latency-microbenchma/ · pack 2026-10-05 · ~677 tokens -->

# Small-message JACCL all_sum latency microbenchmark protocol

> Frames are 4096 bytes and pipeline depth is 2, so an 8 KB message is two frames and a 32 KB message eight; latency steps with frame count, not with bytes.

Parent: [Mac local LLMs: Clusters, RDMA, exo and ds4](https://llms-explorer.com/tree/mac-local-llms-clusters-rdma-exo-ds4/) · 1 facets · 11 facts · page: https://llms-explorer.com/tree/small-message-jaccl-all-sum-latency-microbenchma/

## Facts

- Frames are 4096 bytes and pipeline depth is 2, so an 8 KB message is two frames and a 32 KB message eight; latency steps with frame count, not with bytes. — source: `asserted`
- `all_sum` on an MLX array is lazy and goes through a distributed primitive on a stream. A timing loop must call `mx.eval` on the result inside the timed region, otherwise it measures graph construction. — source: `asserted`
- `barrier()` is itself an `all_sum` of one byte, so a barrier between iterations is a latency probe of the same path and can be used to align ranks. — [source](https://raw.githubusercontent.com/ml-explore/mlx/main/mlx/distributed/jaccl/lib/jaccl/mesh.cpp)
- Sizes to sweep: 1 B (barrier), 4 KB (one frame), 8 KB, 16 KB, 32 KB, 32 KB plus 4 B, 64 KB, 512 KB minus 4 B, 512 KB. This crosses the mesh algorithm switch and the multi-wire switch and the ring 32768-byte switch. — source: `asserted`
- Warm up past the first call: the first call after init builds lazily allocated state in MLX and the first UC exchange follows the warm-up barrier; discard at least tens of calls. — source: `asserted`
- A 2-rank mesh all_reduce of 512 KiB or less runs one wire and one-phase — [source](https://raw.githubusercontent.com/ml-explore/mlx/main/mlx/distributed/jaccl/lib/jaccl/mesh.cpp)
- A mesh above 2 ranks switches to reduce-scatter plus all-gather above 32 KiB per wire — [source](https://raw.githubusercontent.com/ml-explore/mlx/main/mlx/distributed/jaccl/lib/jaccl/mesh.cpp)
- JACCL frames are 4096 bytes with pipeline depth 2 — [source](https://raw.githubusercontent.com/ml-explore/mlx/main/mlx/distributed/jaccl/lib/jaccl/rdma.h)
- `MeshGroup::barrier` is a one-byte `all_sum` and can be used as a one-frame latency probe — [source](https://raw.githubusercontent.com/ml-explore/mlx/main/mlx/distributed/jaccl/lib/jaccl/mesh.cpp)
- A latency sweep should straddle 32 KiB, 32768 B and 512 KiB, measure each of mesh and ring and each node count, call `mx.eval` inside the timed loop and verify values on the last iteration — source: `asserted`
- The MLX docs say JACCL gives communication latency an order of magnitude lower than the ring backend, with no number — [source](https://ml-explore.github.io/mlx/build/html/usage/distributed.html)
