MLX discussion 2990 five-node M3 Ultra RDMA benchmark thread
Parent: Mac local LLMs: Benchmarking and comparisons · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Discussion 2990 ("TB5 RDMA Benchmarks: Pipeline Parallelism Nearly Matches Tensor Parallelism on Kimi-K2 (1T)") was opened by guruswami-ai on 2026-01-12 in the General category with two comments, both by the same author.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Discussion 2990 ("TB5 RDMA Benchmarks: Pipeline Parallelism Nearly Matches Tensor Parallelism on Kimi-K2 (1T)") was opened by guruswami-ai on 2026-01-12 in the General category with two comments, both by the same author. [source]
- The five-node full mesh needs 10 Thunderbolt 5 cables, which is the pair count 5 x 4 / 2; the author names it a "cable spaghetti nightmare" and argues a ring of nodes with less RAM could match a full mesh. [source]
- The author's explanation of the load-time gap is that pipeline parallelism loads only each rank's layers while tensor parallelism "loads all weights then shards them at runtime". The thread gives no measurement of that mechanism. [source]
- The first post is "preliminary findings" and asks for feedback on whether the author is "doing something fundamentally wrong". The context-limit update and the kernel-panic analysis are later comments by the same author. [source]
- The "TPS" values in the context-length table are output tokens divided by wall time including prefill: 50 tokens in 8.4 s is 5.95 (listed 5.96), 50 in 9.2 s is 5.43, 50 in 9.8 s is 5.10, 50 in 10.4 s is 4.81, 50 in 20.4 s is 2.45. They are end-to-end rates, not decode rates, and are not comparable with the 14-15 tok/s headline table. [source]
- The headline table gives no prompt length, repeat count or variance. PP5 (14.45) and PP4 (14.49) differ by 0.3% and PP4 against TP4 by 2.3%, which are single-run differences. [source]
- The thread has no maintainer reply. The quote attributed to Awni Hannun about the GPU timeout is relayed by the author and is not linked. [source]
- The per-node peak memory values do not add up to one model size: summed over nodes they are 512 GB (PP5), 584 GB (PP4), 690 GB (TP2) and 743 GB (TP4) for the same Kimi-K2-Thinking Q4 model, so the column is a peak (including load-time transients), not a weight share. [source]
- The author's reason for TP using more memory, "sharding metadata", is not measured; a load-time transient is an equally consistent cause. [source]
- Thread claim versus later data. The post's key insight is that RDMA makes TP all-reduce "nearly free". The existing dossiers record exo reaching 28.3 tok/s on a 4-node TP run, almost double the 14.82 here, so the near-free reading applies to stock mlx-lm TP at 14-15 tok/s and not to the best TP path. [source]
- Whether the batch-1 equality of PP and TP holds at larger batch sizes and long context; the author lists those tests as upcoming and the follow-on repository publishes other data. [source]
- The thread is already covered in tensor-vs-pipeline-vs-expert-parallelism-for-moe.md and exo-vs-mlx-launch-tensor-parallel-throughput-comparison.md; this file lists deltas only. [source]
- Discussion 2990 has one participant, two comments (both by the author) and no maintainer reply. [source]
- The context-length table's TPS equals output tokens over total wall time, so it includes prefill. [source]
- The headline PP and TP numbers carry no prompt length, run count or variance. [source]
- A five-node full mesh needs 10 Thunderbolt 5 cables. [source]
- Summed per-node peak memory differs across the four working configurations (512, 584, 690 and 743 GB). [source]
Children
- No children recorded.