<!-- llms-explorer concept facts · https://llms-explorer.com/tree/apple-silicon-memory-bandwidth-and-decode-speed/ · pack 2026-10-05 · ~4013 tokens -->

# Apple silicon memory bandwidth and decode speed by chip

> M1 14.19 (79%), M1 Pro 36.41 (69%), M1 Max 61.19 (58%), M1 Ultra 83.73 (40%)

Parent: [Mac local LLMs: Speed, bandwidth and prefill](https://llms-explorer.com/tree/mac-local-llms-speed-bandwidth-and-prefill/) · 1 facets · 79 facts · page: https://llms-explorer.com/tree/apple-silicon-memory-bandwidth-and-decode-speed/

## Facts

- M1 14.19 (79%), M1 Pro 36.41 (69%), M1 Max 61.19 (58%), M1 Ultra 83.73 (40%) — source: `asserted`
- M2 21.91 (83%), M2 Pro 38.86 (74%), M2 Max 65.95 (63%), M2 Ultra 94.27 (45%) — source: `asserted`
- M3 Pro 30.74 (78%), M3 Max-300 56.58 (72%), M3 Max-400 66.31 (63%), M3 Ultra 92.14 (44%) — source: `asserted`
- M4 24.11 (76%), M4 Pro 50.74 (71%), M4 Max-410 69.95 (65%), M4 Max-546 83.06 (58%) — source: `asserted`
- M5 31.88 (79%), M5 Pro 67.59 (84%), M5 Max-614 119.92 (74%), M5 Ultra 179.10 (55%), M6 (170 GB/s) 35.74 (80%) — source: `asserted`
- M3 Pro and the binned M3 Max cut bandwidth versus M2 (150 vs 200; 300 vs 400). M4 Pro restored 273; M4 Max binned 410 / 546. — source: `asserted`
- M5 Pro 307 and M5 Max 460/614 are only 12% above M4 Pro and M4 Max, so decode gains there are small; Apple's "4x AI" claims are prefill (prompt processing). — source: `asserted`
- M5 Ultra (announced 2026-08-25, shipped 2026-09-22; 512GB config late October) raises the Ultra tier from 819 to ~1.2 TB/s. First measured llama.cpp result (Sept 30 2026): Q4_0 tg128 179.10, F16 73.60, pp512 5207. — source: `asserted`
- M6 (Mac mini, 2 nm) has 170 GB/s and decodes 7B Q4_0 at 35.74 tok/s, about 12% above M5 base. — source: `asserted`
- Context length erodes decode: M3 Ultra Qwen 32B Q4 goes 31.2 (1K) -> 19.0 (32K) -> 8.5 (128K); F16 goes 10.4 -> 5.5. At 128K Q2 is only 1.7x F16 (4.6x at 1K) because the FP16 KV cache dominates bandwidth. — source: `asserted`
- Prefill is compute-bound and quant-independent: Llama 405B TTFT is ~37-42 s at 1K and ~10 min at 16K regardless of Q2/Q4/Q8 on M3 Ultra (~54 TFLOPS FP16). MLX prefill throughput dropped 345 -> 154 tok/s from 1K to 128K. — source: `asserted`
- Batching converts the workload from bandwidth-bound to compute-bound: vllm-mlx on M5 Max 128GB, Qwen3-8B-4bit: single stream 87 tok/s, aggregate 215 tok/s at 15 concurrent (2.6x) while per-request decode falls to 16 tok/s. — source: `asserted`
- Binned chips: check the GPU-core bin. A 32-core M5 Max has 460 GB/s, a 40-core 614 GB/s; M4 Max 410 vs 546; M3 Max 300 vs 400. — source: `asserted`
- Power: Low Power Mode can roughly halve inference speed; High Power Mode avoids a 15-25% sag on sustained runs (single blog source). — source: `asserted`
- Wired-limit sysctl key differs by macOS: current iogpu.wired_limit_mb, older debug.iogpu.wired_limit; setting a non-existent key silently does nothing; a value of 0 means the default; the setting does not survive reboot. — source: `asserted`
- Models that exceed the working set swap to disk and decode collapses (about 2 vs 9-10 tok/s in one example). — source: `asserted`
- M5 Pro bandwidth: Apple says 307 GB/s. contracollective.com lists 350 and M5 Max 700, M5 Ultra 1400; llmcheck.net says 300/600/1200. The Apple figures win. The contracollective table (M5 Pro 61 / Max 118 / Ultra 214 tok/s on an 8B 4-bit model) is unverified and inconsistent with Apple's bandwidth. — source: `asserted`
- MLX vs llama.cpp: starmorph.com (citing Groundy) says MLX leads 20-87% under ~14B on M4 Max (Qwen3-8B 4-bit: 93.3 vs 76.9) and the gap collapses to a tie at 27B+. local-llm.net says the two tie at 1-7B and MLX leads 10-20% at 14B+. The two sources disagree on where the gap sits; neither is a primary benchmark. — source: `asserted`
- M5 Max vs M4 Max decode gain: llmcheck.net first claimed ~28% from "lab tests", then on 2026-09-23 corrected to an index estimate of 7-10% consistent with the 11% bandwidth gap (614 vs 546 GB/s). Prefer the corrected figure. — source: `asserted`
- M3 Max vs M4 Pro for LLMs (heyuan110.com): M3 Max (300-400 GB/s) beats M4 Pro (273). The llama.cpp table agrees (66.31 vs 50.74 Q4_0 tg). — source: `asserted`
- Effective-throughput framing: famstack.dev reports MLX at 8.5K context on M1 Max spent 94% of wall time in prefill (49.4 s vs GGUF 37.8 s), so a high decode tok/s number hides slow wall-clock on long prompts. — source: `asserted`
- Independent MLX (not llama.cpp) decode measurements on M5 Pro / M5 Max / M5 Ultra for dense 27-70B models. — source: `asserted`
- Why Ultra-tier decode efficiency is only 40-55% on small models (die-to-die fabric, scheduling, or software) and whether it improves with new builds as on M2 Ultra. — source: `asserted`
- Whether M5 Ultra is a dual-die or quad-die part: pinggy.io claims two bonded dual-die M5 Max chips (4.4 TB/s fabric), Apple's newsroom text does not say. — source: `asserted`
- Sustained thermal behavior of MacBook Pro M5 Max versus Mac Studio, measured. — source: `asserted`
- Exact default GPU wired fraction by RAM size (sources say ~2/3 for 32GB, ~75% on larger). — source: `asserted`
- Apple lists M5 Pro at 307 GB/s and M5 Max at up to 614 GB/s with up to 128GB unified memory. — [source](https://www.apple.com/newsroom/2026/03/apple-introduces-macbook-pro-with-all-new-m5-pro-and-m5-max/)
- M5 Max with a 32-core GPU has 460 GB/s; the 40-core GPU version has 614 GB/s. — [source](https://support.apple.com/en-us/126319)
- M5 Pro and M5 Max use Apple's Fusion Architecture, combining two dies into one SoC. — [source](https://www.apple.com/newsroom/2026/03/apple-introduces-macbook-pro-with-all-new-m5-pro-and-m5-max/)
- Apple's M5 Pro/Max "up to 4x" AI claim is for LLM prompt processing versus M4 Pro/Max, not decode. — [source](https://www.apple.com/newsroom/2026/03/apple-introduces-macbook-pro-with-all-new-m5-pro-and-m5-max/)
- Apple announced Mac Studio with M5 Max and M5 Ultra on 2026-08-25 (pre-order same day, availability September 22): M5 Ultra has up to 80 GPU cores, 36 CPU cores, up to 512GB unified memory and 1.2 TB/s bandwidth, 50 percent higher than before. — [source](https://www.apple.com/newsroom/2026/08/apple-introduces-new-mac-studio-with-m5-max-and-m5-ultra/)
- Apple claims up to 10.7x faster LLM prompt processing in LM Studio for Mac Studio M5 Ultra versus M1 Max and 3.9x versus M4 Max; these are prefill figures. — [source](https://www.apple.com/newsroom/2026/08/apple-introduces-new-mac-studio-with-m5-max-and-m5-ultra/)
- Mac Studio supports Thunderbolt 5 RDMA clustering; Apple says four systems give up to 3x faster inference than one. — [source](https://www.apple.com/newsroom/2026/08/apple-introduces-new-mac-studio-with-m5-max-and-m5-ultra/)
- The M5 Ultra 512GB configuration ships late October 2026 while other configurations ship September 22. — [source](https://www.forbes.com/sites/johnkoetsier/2026/08/25/can-apples-new-mac-ultra-replace-your-200month-ai-coding-bill/)
- M4 base is 120 GB/s with up to 32GB unified memory; M4 Pro and Max raised bandwidth up to 75% over the prior generation. — [source](https://www.apple.com/newsroom/2024/10/apple-introduces-m4-pro-and-m4-max/)
- M4 Pro has 273 GB/s; M4 Max has 546 GB/s at 40 GPU cores and about 410 GB/s at 32 GPU cores. — [source](https://modelfit.io/compare/m4-pro-vs-m4-max-llm/)
- M3 Pro has 150 GB/s versus 200 GB/s for M1 Pro and M2 Pro; the binned M3 Max has 300 GB/s versus 400 for the full part. — [source](https://forums.macrumors.com/threads/m3-pro-has-lower-memory-bandwith-than-m1-m2-pro.2409260/)
- The llama.cpp discussion #4167 table gives Q4_0 7B tg128: M1 14.19, M1 Pro 36.41, M1 Max 61.19, M1 Ultra 83.73 tok/s. — [source](https://github.com/ggml-org/llama.cpp/discussions/4167)
- Same table: M2 21.91, M2 Pro 38.86, M2 Max 65.95, M2 Ultra 94.27 tok/s (Q4_0 tg128, 2023 build). — [source](https://github.com/ggml-org/llama.cpp/discussions/4167)
- Same table: M3 Pro 30.74, M3 Max-300 56.58, M3 Max-400 66.31, M3 Ultra (80-core) 92.14 tok/s. — [source](https://github.com/ggml-org/llama.cpp/discussions/4167)
- Same table: M4 24.11, M4 Pro 50.74, M4 Max-410 69.95, M4 Max-546 83.06 tok/s. — [source](https://github.com/ggml-org/llama.cpp/discussions/4167)
- Same table, newer build: M5 31.88, M5 Pro 67.59, M5 Max-614 119.92, M5 Ultra-80-core 179.10, M6 35.74 tok/s (Q4_0 tg128). — [source](https://github.com/ggml-org/llama.cpp/discussions/4167)
- The same table gives F16 tg128: M4 Max-546 31.64, M5 Pro 21.55, M5 Max 37.11, M5 Ultra 73.60, M3 Ultra-80 39.78 tok/s. — [source](https://github.com/ggml-org/llama.cpp/discussions/4167)
- M5 Ultra prompt processing pp512 is 4944.66 (Q4_0) versus 3219.99 for M5 Max and 1620.64 for M5 Pro. — [source](https://github.com/ggml-org/llama.cpp/discussions/4167)
- Q4_0 decode efficiency (measured tg / (BW / 3.8 GB)) is 69-84% on base and Pro chips, 58-74% on Max chips, and 40-55% on Ultra chips. — source: `asserted`
- Decode does not scale linearly with bandwidth on the small test model: M5 Ultra gives 1.49x M5 Max for 2x bandwidth, M4 Max-546 gives 1.64x M4 Pro for 2x bandwidth. — source: `asserted`
- The same M2 Ultra improved from 94.27 to 125.21 Q4_0 tg128 tok/s between llama.cpp build 8e672ef (Nov 2023) and c1d0e7a (Aug 2026), a 33% gain. — [source](https://github.com/ggml-org/llama.cpp/discussions/4167)
- An M1 Ultra 128GB on build c1d0e7a measured Q4_0 tg128 113.79 tok/s, versus 83.73 on the 2023 build. — [source](https://github.com/ggml-org/llama.cpp/discussions/4167)
- Maintainer ggerganov noted the M5-family rows are missing from the old-build comparison; a contributor argued new-build M5 rows are not comparable to old-build M1-M4 rows. — [source](https://github.com/ggml-org/llama.cpp/discussions/4167)
- A 256GB Mac Studio M5 Ultra (36-core CPU, 80-core GPU) was benchmarked on 2026-09-30 with llama.cpp v0.3.0 on macOS 27.0.1. — [source](https://github.com/ggml-org/llama.cpp/discussions/4167)
- Apple's MLX test (M5 vs M4 24GB MacBook Pro) shows TTFT speedup 3.33-4.06x and generation speedup 1.19-1.27x across six models. — [source](https://machinelearning.apple.com/research/exploring-llms-mlx-m5)
- On M3 Ultra 512GB with MLX, Qwen 32B Q4 decode falls from 31.2 tok/s at 1K context to 19.0 at 32K and 8.5 at 128K. — [source](https://github.com/ml-explore/mlx/discussions/3209)
- On M3 Ultra, Qwen 32B F16 decode is 10.4 tok/s at 1K and 5.5 at 128K; at 128K Q2 is only 1.7x F16 versus 4.6x at 1K. — [source](https://github.com/ml-explore/mlx/discussions/3209)
- On M3 Ultra, Mixtral 8x7B Q4 decodes at 68.4 tok/s at 1K and 46.7 at 32K, versus 31.2 and 19.0 for dense Qwen 32B Q4. — [source](https://github.com/ml-explore/mlx/discussions/3209)
- Prefill time on M3 Ultra is compute-bound and nearly identical across Q2/Q4/Q8 (Llama 405B about 10 min at 16K context). — [source](https://github.com/ml-explore/mlx/discussions/3209)
- The benchmark author estimates M3 Ultra at about 54 TFLOPS FP16 and TTFT ~ 2 x params x context / TFLOPS. — [source](https://github.com/ml-explore/mlx/discussions/3209)
- Measured on M3 Ultra, GLM-5.2 (743B MoE) decodes at 17.7 tok/s, DeepSeek V3 just over 20, dense Llama 405B 2.9. — [source](https://pinggy.io/blog/self_hosting_llms_on_512gb_m5_ultra_mac_studio/)
- Streaming a model larger than memory from SSD through llama.cpp mmap runs at 1-2 tok/s. — [source](https://pinggy.io/blog/self_hosting_llms_on_512gb_m5_ultra_mac_studio/)
- MLX prefill throughput fell from about 345 tok/s at 1K to 154 tok/s at 128K in the M3 Ultra benchmark. — [source](https://pinggy.io/blog/self_hosting_llms_on_512gb_m5_ultra_mac_studio/)
- pinggy.io claims M5 Ultra is a quad-die part (two dual-die M5 Max bonded) with 4.4 TB/s die-to-die bandwidth and 1.2 TB/s on both 64-core and 80-core GPU configs. — [source](https://pinggy.io/blog/self_hosting_llms_on_512gb_m5_ultra_mac_studio/)
- On a MacBook Pro M5 Max 128GB, vllm-mlx with Qwen3-8B-4bit gives 87 tok/s single stream and 215 tok/s aggregate at 15 concurrent requests (per-request 16 tok/s). — [source](https://github.com/waybarrios/vllm-mlx/discussions/637)
- On the same M5 Max, plain mlx-lm with a 0.6B draft model (speculative decoding) reached 107.6 tok/s per request versus 87 for vllm-mlx without speculation. — [source](https://github.com/waybarrios/vllm-mlx/discussions/637)
- On M4 Max, Qwen3-8B 4-bit decodes at 93.3 tok/s with MLX versus 76.9 with llama.cpp; Qwen2.5-27B 4-bit is tied at about 14 tok/s. — [source](https://blog.starmorph.com/blog/apple-silicon-llm-inference-optimization-guide)
- Ollama 0.19 (2026-03-30) switched its Apple backend to MLX; Qwen3.5-35B-A3B on M4 Max 64GB went from 57.8 to 111.4 tok/s decode, and the MLX backend needs 32GB+ RAM. — [source](https://blog.starmorph.com/blog/apple-silicon-llm-inference-optimization-guide)
- Ollama's own benchmark on a Mac mini M4 Pro with Qwen3-Coder-30B-A3B showed 130 tok/s on MLX versus 43 on the old llama.cpp backend. — [source](https://pub.towardsai.net/apples-mlx-runs-local-llms-3x-faster-than-llama-cpp-until-your-context-hits-40k-715ec441afbb)
- A 40,000-token prompt on an MLX-backed local agent took about 3.5 minutes before the first token in one author's test. — [source](https://pub.towardsai.net/apples-mlx-runs-local-llms-3x-faster-than-llama-cpp-until-your-context-hits-40k-715ec441afbb)
- On M1 Max, MLX with Qwen3.5-35B-A3B at 8.5K context spent 94% of wall time in prefill (49.4 s versus 37.8 s for GGUF). — [source](https://blog.starmorph.com/blog/apple-silicon-llm-inference-optimization-guide)
- Low Power Mode can roughly halve LLM inference speed, and High Power Mode avoids a 15-25% throughput sag on sustained runs. — [source](https://www.thinkdifferent.blog/blog/the-macbook-pro-setting-that-doubles-your-ai-performance/)
- The wired-limit key is iogpu.wired_limit_mb on current macOS and debug.iogpu.wired_limit on older macOS; a key the system lacks silently does nothing and 0 restores the default. — [source](https://llmconfigurator.com/en/guides/troubleshooting/increase-metal-vram-limit-apple-silicon)
- The wired-limit sysctl does not persist across reboot; it can be set at boot by launchd or /etc/sysctl.conf, and leaving 32-64GB of headroom on a 512GB machine is advised. — [source](https://github.com/ggml-org/llama.cpp/discussions/2182)
- A 32GB Mac may expose only about 21-24GB (65-75%) to the GPU by default. — [source](https://stencel.io/posts/apple-silicon-limitations-with-usage-on-local-llm%20.html)
- The M5 Max decode gain over M4 Max is estimated at 7-10% (600 vs 546 GB/s as cited there), not the 28% first published; llmcheck.net corrected its page on 2026-09-23. — [source](https://llmcheck.net/blog/apple-silicon-m5-max-local-ai-guide/)
- llmcheck.net and contracollective.com list non-Apple bandwidth figures for M5 Pro/Max/Ultra (300/600/1200 and 350/700/1400 GB/s) that conflict with Apple's 307/614/~1228. — [source](https://contracollective.com/blog/apple-silicon-memory-bandwidth-local-llm-tokens-per-second-m5-2026)
- Per-chip tok/s tables on llmcheck.net, modelfit.io and llmconfigurator.com are labeled bandwidth-roofline estimates, not measurements. — [source](https://modelfit.io/blog/mac-studio-m5-ultra-512gb-local-llm/)
- Mac mini M6 is listed at 170 GB/s, M5 Pro Mac mini at 307 GB/s with up to 64GB. — [source](https://llmconfigurator.com/en/guides/mac-local-ai-buying-guide)
- Apple Silicon M2 Max has a 512-bit memory bus giving 400 GB/s, and M2 Ultra doubles it. — [source](https://medium.com/@andreask_75652/thoughts-on-apple-silicon-performance-for-local-llms-3ef0a50e08bd)
- Across five runtimes on M2 Ultra 192GB, MLX gave the highest sustained decode throughput and Ollama lagged in throughput and TTFT. — [source](https://www.alphaxiv.org/abs/2511.5502v1)
- The M4 Max 128GB was reported at ~100 tok/s on a 30B MoE in a from-scratch inference engine and 127 tok/s on a 24B MoE. — [source](https://news.ycombinator.com/item?id=47232730)
- The Mac Studio M5 Max and MacBook Pro 16-inch M5 Max share the same 460/614 GB/s bandwidth; the portable form factor risks thermal throttling, not bandwidth loss. — [source](https://www.promptquorum.com/local-llms/apple-silicon-m5-local-llm)
