<!-- llms-explorer concept facts · https://llms-explorer.com/tree/jaccl-ring-all-reduce-wire-partitioning-n-wires/ · pack 2026-10-05 · ~859 tokens -->

# JACCL ring all_reduce wire partitioning (n_wires, directions, threads)

> Ring all_reduce uses one direction and one wire when the message is 32768 bytes or less, or has fewer than `size * 2 * n_conns` elements; otherwise both directions and all wires

Parent: [Mac local LLMs: Clusters, RDMA, exo and ds4](https://llms-explorer.com/tree/mac-local-llms-clusters-rdma-exo-ds4/) · 1 facets · 13 facts · page: https://llms-explorer.com/tree/jaccl-ring-all-reduce-wire-partitioning-n-wires/

## Facts

- Ring all_reduce uses one direction and one wire when the message is 32768 bytes or less, or has fewer than `size * 2 * n_conns` elements; otherwise both directions and all wires — [source](https://raw.githubusercontent.com/ml-explore/mlx/main/mlx/distributed/jaccl/lib/jaccl/ring.cpp)
- The per-wire slice is `wire_offset[lr] = lr * n_wires * size_per_wire + lw * size_per_wire`, clamped to the wire's own end so frames cannot spill into the next wire — [source](https://raw.githubusercontent.com/ml-explore/mlx/main/mlx/distributed/jaccl/lib/jaccl/ring_impl.h)
- Ring all_gather and send/recv always use all `n_conns` wires with no small-message threshold — [source](https://raw.githubusercontent.com/ml-explore/mlx/main/mlx/distributed/jaccl/lib/jaccl/ring.cpp)
- Ring sum_scatter uses one wire when `n_bytes <= 65536` or the element count is below `size * 2 * n_conns` — [source](https://raw.githubusercontent.com/ml-explore/mlx/main/mlx/distributed/jaccl/lib/jaccl/ring.cpp)
- `RingGroup` allocates a `ThreadPool` of `n_conns - 1` workers and `dispatch_wires` runs the last wire inline on the calling thread — [source](https://raw.githubusercontent.com/ml-explore/mlx/main/mlx/distributed/jaccl/lib/jaccl/threadpool.h)
- `RING_MAX_CONNS` is 4 per direction and the constructor throws above it — [source](https://raw.githubusercontent.com/ml-explore/mlx/main/mlx/distributed/jaccl/lib/jaccl/ring_impl.h)
- JACCL frames are 4096 bytes with 8 size classes, 2 buffers per class, 32 send and 32 receive work requests, pipeline depth 2 — [source](https://raw.githubusercontent.com/ml-explore/mlx/main/mlx/distributed/jaccl/lib/jaccl/rdma.h)
- Ring buffer j maps to wire `j % n_conns` and direction `j / n_conns`, and is registered to that wire's own protection domain — [source](https://raw.githubusercontent.com/ml-explore/mlx/main/mlx/distributed/jaccl/lib/jaccl/ring.cpp)
- `is_valid_ring` requires every neighbour pair to list as many devices as `devices_[0][1]` — [source](https://raw.githubusercontent.com/ml-explore/mlx/main/mlx/distributed/jaccl/lib/jaccl/jaccl.cpp)
- Mesh `n_wires` is the minimum device count over all pairs, multi-wire starts at 512 KiB, and `MESH_MAX_PEERS` is 8 — [source](https://raw.githubusercontent.com/ml-explore/mlx/main/mlx/distributed/jaccl/lib/jaccl/mesh_impl.h)
- Mesh all_reduce uses reduce-scatter plus all-gather when size is above 2 and the per-wire bytes exceed 32 KiB, else a one-phase fully connected all_reduce — [source](https://raw.githubusercontent.com/ml-explore/mlx/main/mlx/distributed/jaccl/lib/jaccl/mesh.cpp)
- `init` prefers the ring only when `MLX_JACCL_RING` is set and the ring is valid, then falls back mesh, then ring — [source](https://raw.githubusercontent.com/ml-explore/mlx/main/mlx/distributed/jaccl/lib/jaccl/jaccl.cpp)
- PR 3900 says the single-wire path has zero overhead because it runs in the calling thread — [source](https://github.com/ml-explore/mlx/pull/3900)
