<!-- llms-explorer concept facts · https://llms-explorer.com/tree/mlx-expert-parallelism-all-to-all-pr-3158/ · pack 2026-10-05 · ~776 tokens -->

# MLX expert parallelism all_to_all PR 3158

> Existing coverage: tensor-vs-pipeline-vs-expert-parallelism-for-moe.md already records the PR scope, the 3.1x decode and 1.07 prefill numbers, #3164, the Aug 15 2026 close and the 'not on our roadmap' reason; only the items below are absent

Parent: [Mac local LLMs: Clusters, RDMA, exo and ds4](https://llms-explorer.com/tree/mac-local-llms-clusters-rdma-exo-ds4/) · 1 facets · 11 facts · page: https://llms-explorer.com/tree/mlx-expert-parallelism-all-to-all-pr-3158/

## Facts

- Existing coverage: tensor-vs-pipeline-vs-expert-parallelism-for-moe.md already records the PR scope, the 3.1x decode and 1.07 prefill numbers, #3164, the Aug 15 2026 close and the 'not on our roadmap' reason; only the items below are absent — source: `asserted`
- The umbrella PR listed 7 sub-PRs and only PR1A (all_to_all, #3164) was ever checked; PR1B MoE dispatch/combine, PR1C Metal runtime, PR1D Python MixtureOfExperts, PR2A production infra, PR2B performance, PR2C benchmarks were never opened or merged — [source](https://github.com/ml-explore/mlx/pull/3158)
- angeloskath asked on Feb 23 2026 for the work to be split, starting with all_to_all, and the author converted the PR to draft the same day — [source](https://github.com/ml-explore/mlx/pull/3158)
- The planned PR2B would add a batched expert FFN worth about 1.2x decode and a zero-copy combine Metal kernel — [source](https://github.com/ml-explore/mlx/pull/3158)
- The implementation added 7 Metal kernels for O(N*D) data movement (dispatch_local, dispatch_scatter_remote, combine_gather_remote, combine_weighted_sum, packet_gather, packet_scatter) and automatic CPU fallback when world size exceeds 2 — [source](https://github.com/ml-explore/mlx/pull/3158)
- Blocking comm infrastructure used GroupImpl blocking_send/recv/sendrecv and exchange_v with rank-parity ordering to avoid deadlock in variable-size exchange — [source](https://github.com/ml-explore/mlx/pull/3158)
- The Python layer exposed an ep_impl switch ('python' pure-Python or 'cpp' fused) and the tests had more than 40 cases plus 2-rank JACCL dispatch/combine round-trip tests over RDMA — [source](https://github.com/ml-explore/mlx/pull/3158)
- A fix 'Re-acquire Metal command buffer after mid-primitive GPU sync' was needed because gpu::synchronize inside the MoE primitives commits the current command buffer — [source](https://github.com/ml-explore/mlx/pull/3158)
- The PR text is internally inconsistent on the prefill threshold: the table says prefill N at least 256 while the auto backend policy switches to Metal at N at least 320 — [source](https://github.com/ml-explore/mlx/pull/3158)
- Most commits are co-authored by Claude Opus 4.6, consistent with maintainers' stated review-cost objection to large machine-assisted PRs, though the closing comment cites review time only — [source](https://github.com/ml-explore/mlx/pull/3158)
- Because the benchmark covered only 2 ranks and world size above 2 falls back to CPU, no source shows EP behaviour on the 4 to 5 node M3 Ultra meshes used by the benchmark repos — source: `asserted`
