<!-- llms-explorer concept facts · https://llms-explorer.com/tree/mlx-discussion-2990-five-node-m3-ultra-rdma-benc/ · pack 2026-10-05 · ~1104 tokens -->

# MLX discussion 2990 five-node M3 Ultra RDMA benchmark thread

> Discussion 2990 ("TB5 RDMA Benchmarks: Pipeline Parallelism Nearly Matches Tensor Parallelism on Kimi-K2 (1T)") was opened by guruswami-ai on 2026-01-12 in the General category with two comments, both by the same author.

Parent: [Mac local LLMs: Benchmarking and comparisons](https://llms-explorer.com/tree/mac-local-llms-benchmarking-and-comparisons/) · 1 facets · 17 facts · page: https://llms-explorer.com/tree/mlx-discussion-2990-five-node-m3-ultra-rdma-benc/

## Facts

- Discussion 2990 ("TB5 RDMA Benchmarks: Pipeline Parallelism Nearly Matches Tensor Parallelism on Kimi-K2 (1T)") was opened by guruswami-ai on 2026-01-12 in the General category with two comments, both by the same author. — [source](https://github.com/ml-explore/mlx/discussions/2990)
- The five-node full mesh needs 10 Thunderbolt 5 cables, which is the pair count 5 x 4 / 2; the author names it a "cable spaghetti nightmare" and argues a ring of nodes with less RAM could match a full mesh. — [source](https://github.com/ml-explore/mlx/discussions/2990)
- The author's explanation of the load-time gap is that pipeline parallelism loads only each rank's layers while tensor parallelism "loads all weights then shards them at runtime". The thread gives no measurement of that mechanism. — [source](https://github.com/ml-explore/mlx/discussions/2990)
- The first post is "preliminary findings" and asks for feedback on whether the author is "doing something fundamentally wrong". The context-limit update and the kernel-panic analysis are later comments by the same author. — [source](https://github.com/ml-explore/mlx/discussions/2990)
- The "TPS" values in the context-length table are output tokens divided by wall time including prefill: 50 tokens in 8.4 s is 5.95 (listed 5.96), 50 in 9.2 s is 5.43, 50 in 9.8 s is 5.10, 50 in 10.4 s is 4.81, 50 in 20.4 s is 2.45. They are end-to-end rates, not decode rates, and are not comparable with the 14-15 tok/s headline table. — [source](https://github.com/ml-explore/mlx/discussions/2990)
- The headline table gives no prompt length, repeat count or variance. PP5 (14.45) and PP4 (14.49) differ by 0.3% and PP4 against TP4 by 2.3%, which are single-run differences. — [source](https://github.com/ml-explore/mlx/discussions/2990)
- The thread has no maintainer reply. The quote attributed to Awni Hannun about the GPU timeout is relayed by the author and is not linked. — [source](https://github.com/ml-explore/mlx/discussions/2990)
- The per-node peak memory values do not add up to one model size: summed over nodes they are 512 GB (PP5), 584 GB (PP4), 690 GB (TP2) and 743 GB (TP4) for the same Kimi-K2-Thinking Q4 model, so the column is a peak (including load-time transients), not a weight share. — source: `asserted`
- The author's reason for TP using more memory, "sharding metadata", is not measured; a load-time transient is an equally consistent cause. — source: `asserted`
- Thread claim versus later data. The post's key insight is that RDMA makes TP all-reduce "nearly free". The existing dossiers record exo reaching 28.3 tok/s on a 4-node TP run, almost double the 14.82 here, so the near-free reading applies to stock mlx-lm TP at 14-15 tok/s and not to the best TP path. — source: `asserted`
- Whether the batch-1 equality of PP and TP holds at larger batch sizes and long context; the author lists those tests as upcoming and the follow-on repository publishes other data. — source: `asserted`
- The thread is already covered in tensor-vs-pipeline-vs-expert-parallelism-for-moe.md and exo-vs-mlx-launch-tensor-parallel-throughput-comparison.md; this file lists deltas only. — source: `asserted`
- Discussion 2990 has one participant, two comments (both by the author) and no maintainer reply. — [source](https://github.com/ml-explore/mlx/discussions/2990)
- The context-length table's TPS equals output tokens over total wall time, so it includes prefill. — [source](https://github.com/ml-explore/mlx/discussions/2990)
- The headline PP and TP numbers carry no prompt length, run count or variance. — [source](https://github.com/ml-explore/mlx/discussions/2990)
- A five-node full mesh needs 10 Thunderbolt 5 cables. — [source](https://github.com/ml-explore/mlx/discussions/2990)
- Summed per-node peak memory differs across the four working configurations (512, 584, 690 and 743 GB). — source: `asserted`
