Apple silicon memory bandwidth and decode speed by chip
Parent: Mac local LLMs: Speed, bandwidth and prefill · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
M1 14.19 (79%), M1 Pro 36.41 (69%), M1 Max 61.19 (58%), M1 Ultra 83.73 (40%)
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- M1 14.19 (79%), M1 Pro 36.41 (69%), M1 Max 61.19 (58%), M1 Ultra 83.73 (40%) [source]
- M2 21.91 (83%), M2 Pro 38.86 (74%), M2 Max 65.95 (63%), M2 Ultra 94.27 (45%) [source]
- M3 Pro 30.74 (78%), M3 Max-300 56.58 (72%), M3 Max-400 66.31 (63%), M3 Ultra 92.14 (44%) [source]
- M4 24.11 (76%), M4 Pro 50.74 (71%), M4 Max-410 69.95 (65%), M4 Max-546 83.06 (58%) [source]
- M5 31.88 (79%), M5 Pro 67.59 (84%), M5 Max-614 119.92 (74%), M5 Ultra 179.10 (55%), M6 (170 GB/s) 35.74 (80%) [source]
- M3 Pro and the binned M3 Max cut bandwidth versus M2 (150 vs 200; 300 vs 400). M4 Pro restored 273; M4 Max binned 410 / 546. [source]
- M5 Pro 307 and M5 Max 460/614 are only 12% above M4 Pro and M4 Max, so decode gains there are small; Apple's "4x AI" claims are prefill (prompt processing). [source]
- M5 Ultra (announced 2026-08-25, shipped 2026-09-22; 512GB config late October) raises the Ultra tier from 819 to ~1.2 TB/s. First measured llama.cpp result (Sept 30 2026): Q4_0 tg128 179.10, F16 73.60, pp512 5207. [source]
- M6 (Mac mini, 2 nm) has 170 GB/s and decodes 7B Q4_0 at 35.74 tok/s, about 12% above M5 base. [source]
- Context length erodes decode: M3 Ultra Qwen 32B Q4 goes 31.2 (1K) -> 19.0 (32K) -> 8.5 (128K); F16 goes 10.4 -> 5.5. At 128K Q2 is only 1.7x F16 (4.6x at 1K) because the FP16 KV cache dominates bandwidth. [source]
- Prefill is compute-bound and quant-independent: Llama 405B TTFT is ~37-42 s at 1K and ~10 min at 16K regardless of Q2/Q4/Q8 on M3 Ultra (~54 TFLOPS FP16). MLX prefill throughput dropped 345 -> 154 tok/s from 1K to 128K. [source]
- Batching converts the workload from bandwidth-bound to compute-bound: vllm-mlx on M5 Max 128GB, Qwen3-8B-4bit: single stream 87 tok/s, aggregate 215 tok/s at 15 concurrent (2.6x) while per-request decode falls to 16 tok/s. [source]
- Binned chips: check the GPU-core bin. A 32-core M5 Max has 460 GB/s, a 40-core 614 GB/s; M4 Max 410 vs 546; M3 Max 300 vs 400. [source]
- Power: Low Power Mode can roughly halve inference speed; High Power Mode avoids a 15-25% sag on sustained runs (single blog source). [source]
- Wired-limit sysctl key differs by macOS: current iogpu.wired_limit_mb, older debug.iogpu.wired_limit; setting a non-existent key silently does nothing; a value of 0 means the default; the setting does not survive reboot. [source]
- Models that exceed the working set swap to disk and decode collapses (about 2 vs 9-10 tok/s in one example). [source]
- M5 Pro bandwidth: Apple says 307 GB/s. contracollective.com lists 350 and M5 Max 700, M5 Ultra 1400; llmcheck.net says 300/600/1200. The Apple figures win. The contracollective table (M5 Pro 61 / Max 118 / Ultra 214 tok/s on an 8B 4-bit model) is unverified and inconsistent with Apple's bandwidth. [source]
- MLX vs llama.cpp: starmorph.com (citing Groundy) says MLX leads 20-87% under ~14B on M4 Max (Qwen3-8B 4-bit: 93.3 vs 76.9) and the gap collapses to a tie at 27B+. local-llm.net says the two tie at 1-7B and MLX leads 10-20% at 14B+. The two sources disagree on where the gap sits; neither is a primary benchmark. [source]
- M5 Max vs M4 Max decode gain: llmcheck.net first claimed ~28% from "lab tests", then on 2026-09-23 corrected to an index estimate of 7-10% consistent with the 11% bandwidth gap (614 vs 546 GB/s). Prefer the corrected figure. [source]
- M3 Max vs M4 Pro for LLMs (heyuan110.com): M3 Max (300-400 GB/s) beats M4 Pro (273). The llama.cpp table agrees (66.31 vs 50.74 Q4_0 tg). [source]
- Effective-throughput framing: famstack.dev reports MLX at 8.5K context on M1 Max spent 94% of wall time in prefill (49.4 s vs GGUF 37.8 s), so a high decode tok/s number hides slow wall-clock on long prompts. [source]
- Independent MLX (not llama.cpp) decode measurements on M5 Pro / M5 Max / M5 Ultra for dense 27-70B models. [source]
- Why Ultra-tier decode efficiency is only 40-55% on small models (die-to-die fabric, scheduling, or software) and whether it improves with new builds as on M2 Ultra. [source]
- Whether M5 Ultra is a dual-die or quad-die part: pinggy.io claims two bonded dual-die M5 Max chips (4.4 TB/s fabric), Apple's newsroom text does not say. [source]
- Sustained thermal behavior of MacBook Pro M5 Max versus Mac Studio, measured. [source]
- Exact default GPU wired fraction by RAM size (sources say ~2/3 for 32GB, ~75% on larger). [source]
- Apple lists M5 Pro at 307 GB/s and M5 Max at up to 614 GB/s with up to 128GB unified memory. [source]
- M5 Max with a 32-core GPU has 460 GB/s; the 40-core GPU version has 614 GB/s. [source]
- M5 Pro and M5 Max use Apple's Fusion Architecture, combining two dies into one SoC. [source]
- Apple's M5 Pro/Max "up to 4x" AI claim is for LLM prompt processing versus M4 Pro/Max, not decode. [source]
- Apple announced Mac Studio with M5 Max and M5 Ultra on 2026-08-25 (pre-order same day, availability September 22): M5 Ultra has up to 80 GPU cores, 36 CPU cores, up to 512GB unified memory and 1.2 TB/s bandwidth, 50 percent higher than before. [source]
- Apple claims up to 10.7x faster LLM prompt processing in LM Studio for Mac Studio M5 Ultra versus M1 Max and 3.9x versus M4 Max; these are prefill figures. [source]
- Mac Studio supports Thunderbolt 5 RDMA clustering; Apple says four systems give up to 3x faster inference than one. [source]
- The M5 Ultra 512GB configuration ships late October 2026 while other configurations ship September 22. [source]
- M4 base is 120 GB/s with up to 32GB unified memory; M4 Pro and Max raised bandwidth up to 75% over the prior generation. [source]
- M4 Pro has 273 GB/s; M4 Max has 546 GB/s at 40 GPU cores and about 410 GB/s at 32 GPU cores. [source]
- M3 Pro has 150 GB/s versus 200 GB/s for M1 Pro and M2 Pro; the binned M3 Max has 300 GB/s versus 400 for the full part. [source]
- The llama.cpp discussion #4167 table gives Q4_0 7B tg128: M1 14.19, M1 Pro 36.41, M1 Max 61.19, M1 Ultra 83.73 tok/s. [source]
- Same table: M2 21.91, M2 Pro 38.86, M2 Max 65.95, M2 Ultra 94.27 tok/s (Q4_0 tg128, 2023 build). [source]
- Same table: M3 Pro 30.74, M3 Max-300 56.58, M3 Max-400 66.31, M3 Ultra (80-core) 92.14 tok/s. [source]
- Same table: M4 24.11, M4 Pro 50.74, M4 Max-410 69.95, M4 Max-546 83.06 tok/s. [source]
- Same table, newer build: M5 31.88, M5 Pro 67.59, M5 Max-614 119.92, M5 Ultra-80-core 179.10, M6 35.74 tok/s (Q4_0 tg128). [source]
- The same table gives F16 tg128: M4 Max-546 31.64, M5 Pro 21.55, M5 Max 37.11, M5 Ultra 73.60, M3 Ultra-80 39.78 tok/s. [source]
- M5 Ultra prompt processing pp512 is 4944.66 (Q4_0) versus 3219.99 for M5 Max and 1620.64 for M5 Pro. [source]
- Q4_0 decode efficiency (measured tg / (BW / 3.8 GB)) is 69-84% on base and Pro chips, 58-74% on Max chips, and 40-55% on Ultra chips. [source]
- Decode does not scale linearly with bandwidth on the small test model: M5 Ultra gives 1.49x M5 Max for 2x bandwidth, M4 Max-546 gives 1.64x M4 Pro for 2x bandwidth. [source]
- The same M2 Ultra improved from 94.27 to 125.21 Q4_0 tg128 tok/s between llama.cpp build 8e672ef (Nov 2023) and c1d0e7a (Aug 2026), a 33% gain. [source]
- An M1 Ultra 128GB on build c1d0e7a measured Q4_0 tg128 113.79 tok/s, versus 83.73 on the 2023 build. [source]
- Maintainer ggerganov noted the M5-family rows are missing from the old-build comparison; a contributor argued new-build M5 rows are not comparable to old-build M1-M4 rows. [source]
- A 256GB Mac Studio M5 Ultra (36-core CPU, 80-core GPU) was benchmarked on 2026-09-30 with llama.cpp v0.3.0 on macOS 27.0.1. [source]
- Apple's MLX test (M5 vs M4 24GB MacBook Pro) shows TTFT speedup 3.33-4.06x and generation speedup 1.19-1.27x across six models. [source]
- On M3 Ultra 512GB with MLX, Qwen 32B Q4 decode falls from 31.2 tok/s at 1K context to 19.0 at 32K and 8.5 at 128K. [source]
- On M3 Ultra, Qwen 32B F16 decode is 10.4 tok/s at 1K and 5.5 at 128K; at 128K Q2 is only 1.7x F16 versus 4.6x at 1K. [source]
- On M3 Ultra, Mixtral 8x7B Q4 decodes at 68.4 tok/s at 1K and 46.7 at 32K, versus 31.2 and 19.0 for dense Qwen 32B Q4. [source]
- Prefill time on M3 Ultra is compute-bound and nearly identical across Q2/Q4/Q8 (Llama 405B about 10 min at 16K context). [source]
- The benchmark author estimates M3 Ultra at about 54 TFLOPS FP16 and TTFT ~ 2 x params x context / TFLOPS. [source]
- Measured on M3 Ultra, GLM-5.2 (743B MoE) decodes at 17.7 tok/s, DeepSeek V3 just over 20, dense Llama 405B 2.9. [source]
- Streaming a model larger than memory from SSD through llama.cpp mmap runs at 1-2 tok/s. [source]
- MLX prefill throughput fell from about 345 tok/s at 1K to 154 tok/s at 128K in the M3 Ultra benchmark. [source]
- pinggy.io claims M5 Ultra is a quad-die part (two dual-die M5 Max bonded) with 4.4 TB/s die-to-die bandwidth and 1.2 TB/s on both 64-core and 80-core GPU configs. [source]
- On a MacBook Pro M5 Max 128GB, vllm-mlx with Qwen3-8B-4bit gives 87 tok/s single stream and 215 tok/s aggregate at 15 concurrent requests (per-request 16 tok/s). [source]
- On the same M5 Max, plain mlx-lm with a 0.6B draft model (speculative decoding) reached 107.6 tok/s per request versus 87 for vllm-mlx without speculation. [source]
- On M4 Max, Qwen3-8B 4-bit decodes at 93.3 tok/s with MLX versus 76.9 with llama.cpp; Qwen2.5-27B 4-bit is tied at about 14 tok/s. [source]
- Ollama 0.19 (2026-03-30) switched its Apple backend to MLX; Qwen3.5-35B-A3B on M4 Max 64GB went from 57.8 to 111.4 tok/s decode, and the MLX backend needs 32GB+ RAM. [source]
- Ollama's own benchmark on a Mac mini M4 Pro with Qwen3-Coder-30B-A3B showed 130 tok/s on MLX versus 43 on the old llama.cpp backend. [source]
- A 40,000-token prompt on an MLX-backed local agent took about 3.5 minutes before the first token in one author's test. [source]
- On M1 Max, MLX with Qwen3.5-35B-A3B at 8.5K context spent 94% of wall time in prefill (49.4 s versus 37.8 s for GGUF). [source]
- Low Power Mode can roughly halve LLM inference speed, and High Power Mode avoids a 15-25% throughput sag on sustained runs. [source]
- The wired-limit key is iogpu.wired_limit_mb on current macOS and debug.iogpu.wired_limit on older macOS; a key the system lacks silently does nothing and 0 restores the default. [source]
- The wired-limit sysctl does not persist across reboot; it can be set at boot by launchd or /etc/sysctl.conf, and leaving 32-64GB of headroom on a 512GB machine is advised. [source]
- A 32GB Mac may expose only about 21-24GB (65-75%) to the GPU by default. [source]
- The M5 Max decode gain over M4 Max is estimated at 7-10% (600 vs 546 GB/s as cited there), not the 28% first published; llmcheck.net corrected its page on 2026-09-23. [source]
- llmcheck.net and contracollective.com list non-Apple bandwidth figures for M5 Pro/Max/Ultra (300/600/1200 and 350/700/1400 GB/s) that conflict with Apple's 307/614/~1228. [source]
- Per-chip tok/s tables on llmcheck.net, modelfit.io and llmconfigurator.com are labeled bandwidth-roofline estimates, not measurements. [source]
- Mac mini M6 is listed at 170 GB/s, M5 Pro Mac mini at 307 GB/s with up to 64GB. [source]
- Apple Silicon M2 Max has a 512-bit memory bus giving 400 GB/s, and M2 Ultra doubles it. [source]
- Across five runtimes on M2 Ultra 192GB, MLX gave the highest sustained decode throughput and Ollama lagged in throughput and TTFT. [source]
- The M4 Max 128GB was reported at ~100 tok/s on a 30B MoE in a from-scratch inference engine and 127 tok/s on a 24B MoE. [source]
- The Mac Studio M5 Max and MacBook Pro 16-inch M5 Max share the same 460/614 GB/s bandwidth; the portable form factor risks thermal throttling, not bandwidth loss. [source]
Children
- No children recorded.