guruswami-ai mlx-benchmarks cluster dataset
Parent: Mac local LLMs: Benchmarking and comparisons · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Existing coverage: distributed-inference-across-macs.md and tensor-vs-pipeline-vs-expert-parallelism-for-moe.md already carry the benchmark figures, the EP terminology trap, and the RDMA failure modes; only absent claims follow
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Existing coverage: distributed-inference-across-macs.md and tensor-vs-pipeline-vs-expert-parallelism-for-moe.md already carry the benchmark figures, the EP terminology trap, and the RDMA failure modes; only absent claims follow [source]
- The repository dataset has 290 Apple points on a 5-node M3 Ultra cluster plus 1,353 NVIDIA points: RTX 3080 342, RTX 4090 510, RTX 5090 501 [source]
- The CSV schema in results/ is model, topology, nodes, quant, context_tokens, prompt_tps, generation_tps, peak_memory_gb, ttft_seconds, ttft_minutes, feasibility, node, plus a 42-record perplexity file [source]
- The README lists GLM-5.3 (744B, 40B active) as a single-node row with mixed 4/8-bit quant plus an MTP head, covering accuracy, prefill, decode, 230K context and perplexity [source]
- The repo's latest commit is dated Sep 3 2026 (accuracy, speed, context and perplexity results) while the docs pages carry a Mar 24 2026 commit, so the docs and the newest data are out of step [source]
- Licences differ: data and charts are CC BY-ND 4.0 (no derivatives) and the patches in patches/ are Apache 2.0, so reusing the figures in a derived dataset is restricted [source]
- The patches are patches/llama.py and qwen2.py (pipeline parallelism) and mixtral.py (tensor plus pipeline), submitted as PRs to ml-explore/mlx-lm [source]
- The repo had 3 stars and 0 forks when fetched, so it is a single-vendor dataset with no community replication [source]
- The authors only publish TP2 and TP4 because TP5 fails head divisibility: Llama 405B has 128 query and 8 KV heads, Qwen 32B has 40 and 8, so 8/5 fails the KV check; Qwen 32B also passes only the query check, making its failure framework-dependent [source]
- MLA models (DeepSeek V3, Kimi K2.5) are listed as failing TP5 on 128 and 64 query heads alone, with no KV-head test [source]
- Headline: 4 nodes made Qwen 32B generation 42 percent slower, and Mixtral was predicted at 23 tok/s but measured 69 tok/s because active parameters differ from total [source]
- Planned additions named in the README: a cluster simulator, an NVIDIA 'LLM Space Heater' benchmark tool, Qwen 3.5, Llama 4 and M4 Pro/Max results; none had shipped at fetch [source]
Children
- No children recorded.