MLX expert parallelism all_to_all PR 3158
Parent: Mac local LLMs: Clusters, RDMA, exo and ds4 · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Existing coverage: tensor-vs-pipeline-vs-expert-parallelism-for-moe.md already records the PR scope, the 3.1x decode and 1.07 prefill numbers, #3164, the Aug 15 2026 close and the 'not on our roadmap' reason; only the items below are absent
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Existing coverage: tensor-vs-pipeline-vs-expert-parallelism-for-moe.md already records the PR scope, the 3.1x decode and 1.07 prefill numbers, #3164, the Aug 15 2026 close and the 'not on our roadmap' reason; only the items below are absent [source]
- The umbrella PR listed 7 sub-PRs and only PR1A (all_to_all, #3164) was ever checked; PR1B MoE dispatch/combine, PR1C Metal runtime, PR1D Python MixtureOfExperts, PR2A production infra, PR2B performance, PR2C benchmarks were never opened or merged [source]
- angeloskath asked on Feb 23 2026 for the work to be split, starting with all_to_all, and the author converted the PR to draft the same day [source]
- The planned PR2B would add a batched expert FFN worth about 1.2x decode and a zero-copy combine Metal kernel [source]
- The implementation added 7 Metal kernels for O(N*D) data movement (dispatch_local, dispatch_scatter_remote, combine_gather_remote, combine_weighted_sum, packet_gather, packet_scatter) and automatic CPU fallback when world size exceeds 2 [source]
- Blocking comm infrastructure used GroupImpl blocking_send/recv/sendrecv and exchange_v with rank-parity ordering to avoid deadlock in variable-size exchange [source]
- The Python layer exposed an ep_impl switch ('python' pure-Python or 'cpp' fused) and the tests had more than 40 cases plus 2-rank JACCL dispatch/combine round-trip tests over RDMA [source]
- A fix 'Re-acquire Metal command buffer after mid-primitive GPU sync' was needed because gpu::synchronize inside the MoE primitives commits the current command buffer [source]
- The PR text is internally inconsistent on the prefill threshold: the table says prefill N at least 256 while the auto backend policy switches to Metal at N at least 320 [source]
- Most commits are co-authored by Claude Opus 4.6, consistent with maintainers' stated review-cost objection to large machine-assisted PRs, though the closing comment cites review time only [source]
- Because the benchmark covered only 2 ranks and world size above 2 falls back to CPU, no source shows EP behaviour on the 4 to 5 node M3 Ultra meshes used by the benchmark repos [source]
Children
- No children recorded.