<!-- llms-explorer concept facts · https://llms-explorer.com/tree/guruswami-ai-mlx-benchmarks-cluster-dataset/ · pack 2026-10-05 · ~840 tokens -->

# guruswami-ai mlx-benchmarks cluster dataset

> Existing coverage: distributed-inference-across-macs.md and tensor-vs-pipeline-vs-expert-parallelism-for-moe.md already carry the benchmark figures, the EP terminology trap, and the RDMA failure modes; only absent claims follow

Parent: [Mac local LLMs: Benchmarking and comparisons](https://llms-explorer.com/tree/mac-local-llms-benchmarking-and-comparisons/) · 1 facets · 12 facts · page: https://llms-explorer.com/tree/guruswami-ai-mlx-benchmarks-cluster-dataset/

## Facts

- Existing coverage: distributed-inference-across-macs.md and tensor-vs-pipeline-vs-expert-parallelism-for-moe.md already carry the benchmark figures, the EP terminology trap, and the RDMA failure modes; only absent claims follow — source: `asserted`
- The repository dataset has 290 Apple points on a 5-node M3 Ultra cluster plus 1,353 NVIDIA points: RTX 3080 342, RTX 4090 510, RTX 5090 501 — [source](https://github.com/guruswami-ai/mlx-benchmarks)
- The CSV schema in results/ is model, topology, nodes, quant, context_tokens, prompt_tps, generation_tps, peak_memory_gb, ttft_seconds, ttft_minutes, feasibility, node, plus a 42-record perplexity file — [source](https://github.com/guruswami-ai/mlx-benchmarks)
- The README lists GLM-5.3 (744B, 40B active) as a single-node row with mixed 4/8-bit quant plus an MTP head, covering accuracy, prefill, decode, 230K context and perplexity — [source](https://github.com/guruswami-ai/mlx-benchmarks)
- The repo's latest commit is dated Sep 3 2026 (accuracy, speed, context and perplexity results) while the docs pages carry a Mar 24 2026 commit, so the docs and the newest data are out of step — [source](https://github.com/guruswami-ai/mlx-benchmarks)
- Licences differ: data and charts are CC BY-ND 4.0 (no derivatives) and the patches in patches/ are Apache 2.0, so reusing the figures in a derived dataset is restricted — [source](https://github.com/guruswami-ai/mlx-benchmarks)
- The patches are patches/llama.py and qwen2.py (pipeline parallelism) and mixtral.py (tensor plus pipeline), submitted as PRs to ml-explore/mlx-lm — [source](https://github.com/guruswami-ai/mlx-benchmarks)
- The repo had 3 stars and 0 forks when fetched, so it is a single-vendor dataset with no community replication — [source](https://github.com/guruswami-ai/mlx-benchmarks)
- The authors only publish TP2 and TP4 because TP5 fails head divisibility: Llama 405B has 128 query and 8 KV heads, Qwen 32B has 40 and 8, so 8/5 fails the KV check; Qwen 32B also passes only the query check, making its failure framework-dependent — [source](https://github.com/guruswami-ai/mlx-benchmarks/blob/main/docs/DISTRIBUTED_INFERENCE.md)
- MLA models (DeepSeek V3, Kimi K2.5) are listed as failing TP5 on 128 and 64 query heads alone, with no KV-head test — [source](https://github.com/guruswami-ai/mlx-benchmarks/blob/main/docs/DISTRIBUTED_INFERENCE.md)
- Headline: 4 nodes made Qwen 32B generation 42 percent slower, and Mixtral was predicted at 23 tok/s but measured 69 tok/s because active parameters differ from total — [source](https://github.com/guruswami-ai/mlx-benchmarks)
- Planned additions named in the README: a cluster simulator, an NVIDIA 'LLM Space Heater' benchmark tool, Qwen 3.5, Llama 4 and M4 Pro/Max results; none had shipped at fetch — [source](https://github.com/guruswami-ai/mlx-benchmarks)
