<!-- llms-explorer concept facts · https://llms-explorer.com/tree/ssd-bandwidth-growth-per-apple-chip-generation-a/ · pack 2026-10-05 · ~2052 tokens -->

# SSD bandwidth growth per Apple chip generation and 397B projection

> Apple's March 2026 MacBook Pro release says M5 Pro and M5 Max models deliver "up to 2x faster read/write performance compared to the previous generation", "reaching speeds of up to 14.5GB/s".

Parent: [Mac local LLMs: MoE streaming and offload](https://llms-explorer.com/tree/mac-local-llms-moe-streaming-and-offload/) · 2 facets · 31 facts · page: https://llms-explorer.com/tree/ssd-bandwidth-growth-per-apple-chip-generation-a/

## Facts

- Apple's March 2026 MacBook Pro release says M5 Pro and M5 Max models deliver "up to 2x faster read/write performance compared to the previous generation", "reaching speeds of up to 14.5GB/s". — [source](https://www.apple.com/newsroom/2026/03/apple-introduces-macbook-pro-with-all-new-m5-pro-and-m5-max/)
- Apple's footnote: the 14.5 GB/s figure used a preproduction 16-inch MacBook Pro with M5 Max, 128 GB, 8 TB SSD, FIO 3.41, 1024 KB request size, 10 GB test file, IO depth 8, tested January and February 2026. The comparison systems were an M4 Pro with 4 TB SSD and an M4 Max with 8 TB SSD. — [source](https://www.apple.com/newsroom/2026/03/apple-introduces-macbook-pro-with-all-new-m5-pro-and-m5-max/)
- If "up to 2x" holds at the same test, the M4 Max 8 TB baseline is about 7.2 GB/s. This is derived from Apple's ratio, not stated by Apple. — source: `asserted`
- Third-party Blackmagic Disk Speed Test readings on an M4 Max Mac Studio: a 512 GB model reads about 5.0 to 5.2 GB/s (writes 3.9 to 5.7 GB/s); a 2 TB model reads 5.2 GB/s in one post (writes 6.3 to 8.4 GB/s, fluctuating), and another 2 TB owner reports 6.78 GB/s read and 6.94 GB/s write. A 5 GB test reads 5.2 GB/s and writes 4.3 GB/s. — [source](https://forums.macrumors.com/threads/mac-studio-m4-max-and-ssd-drive-speed.2453068/page-2)
- Forum posters in that thread say the 512 GB, 1 TB and 2 TB configurations are slower than the 4 TB and 8 TB ones, which matches fewer parallel NAND packages in smaller capacities. The thread shows no Apple confirmation. — [source](https://forums.macrumors.com/threads/mac-studio-m4-max-and-ssd-drive-speed.2453068/page-2)
- Per-token read volume for Qwen3.5-397B-A17B in the Anemll fork: 27.0 MB per layer at 4-bit and 21.8 MB per layer at Q3, 60 layers. That is 1.62 GB and 1.31 GB per token (derived by multiplication). — [source](https://github.com/Anemll/flash-moe)
- Ceiling at zero cache hits: at Apple's 14.5 GB/s, 1.62 GB per token gives 8.9 tok/s (4-bit) and 1.31 GB gives 11.1 tok/s (Q3). At Flash-MoE's cold 4-thread `F_NOCACHE` rate of 5.5 GB/s (M3 Max), the ceilings are 3.4 and 4.2 tok/s. All four are derived, not measured. — source: `asserted`
- 2026-03: Apple states a 2x storage step from M4 Pro and M4 Max to M5 Pro and M5 Max. — [source](https://www.apple.com/newsroom/2026/03/apple-introduces-macbook-pro-with-all-new-m5-pro-and-m5-max/)
- The Flash-MoE paper projected about 20% SSD gain per generation, so two generations after the M3 Max would be about 1.44x. Apple's own step for the latest generation alone is up to 2x. — source: `asserted`
- Capacity changes bandwidth: a 512 GB machine can read at about 5 GB/s while a large-capacity machine of the same chip reads faster. A streaming projection must name the SSD size. — [source](https://forums.macrumors.com/threads/mac-studio-m4-max-and-ssd-drive-speed.2453068/page-2)
- Blackmagic numbers are a different test from FIO at 1024 KB and IO depth 8, and vary run to run (one owner saw 8.4 GB/s write drop to 6.3 GB/s on repeat). They are not interchangeable with Apple's 14.5 GB/s. — [source](https://forums.macrumors.com/threads/mac-studio-m4-max-and-ssd-drive-speed.2453068/page-2)
- Apple's number is "up to", from a single preproduction 8 TB unit. — [source](https://www.apple.com/newsroom/2026/03/apple-introduces-macbook-pro-with-all-new-m5-pro-and-m5-max/)
- Paper projection versus Apple. The Flash-MoE paper assumes about 20% per generation; Apple reports up to 2x for one step. The two cannot both describe M-series storage, but the Apple figure is a vendor peak and the 20% figure is an author guess. — source: `asserted`
- No first-party or independent sequential-read number for M1, M2 or M3 family SSDs, and no per-generation series, was found. Reddit pages with Blackmagic results for M1 Max and Mac mini M4 Pro were blocked at fetch. — source: `asserted`
- No measured cold-cache expert-streaming rate on an M5 Max; Apple's 14.5 GB/s is a synthetic FIO read, not a streaming run. — source: `asserted`
- Whether the M5 Pro and M5 Max 2x applies to the 1 TB and 2 TB base capacities or only to larger ones. — source: `asserted`
- Apple says M5 Pro and M5 Max MacBook Pro models reach up to 14.5 GB/s storage speed and up to 2x the read/write performance of the previous generation. — [source](https://www.apple.com/newsroom/2026/03/apple-introduces-macbook-pro-with-all-new-m5-pro-and-m5-max/)
- Apple's 14.5 GB/s test used FIO 3.41 with a 1024 KB request size, a 10 GB file and IO depth 8 on an M5 Max 128 GB 8 TB preproduction unit. — [source](https://www.apple.com/newsroom/2026/03/apple-introduces-macbook-pro-with-all-new-m5-pro-and-m5-max/)
- Apple's comparison baselines for that claim were an M4 Pro MacBook Pro with 4 TB SSD and an M4 Max MacBook Pro with 8 TB SSD. — [source](https://www.apple.com/newsroom/2026/03/apple-introduces-macbook-pro-with-all-new-m5-pro-and-m5-max/)
- Apple's ratio implies about 7.2 GB/s for an M4 Max 8 TB under the same test. — source: `asserted`
- MacRumors owners report M4 Max Mac Studio Blackmagic reads of about 5.0 to 5.2 GB/s on 512 GB and 2 TB models, and 6.78 GB/s read with 6.94 GB/s write on another 2 TB unit. — [source](https://forums.macrumors.com/threads/mac-studio-m4-max-and-ssd-drive-speed.2453068/page-2)
- Forum posters in the same thread say lower-capacity Mac Studio SSDs are slower than 4 TB and 8 TB models. — [source](https://forums.macrumors.com/threads/mac-studio-m4-max-and-ssd-drive-speed.2453068/page-2)
- The Anemll Flash-MoE fork's M5 Max decode breakdown lists expert I/O of 0.74 ms for 27.0 MB per layer at 4-bit (1.62 GB of experts per token) and 0.44 ms for 21.8 MB per layer with Q3 experts, over 60 layers. — [source](https://github.com/Anemll/flash-moe)
- The Anemll Flash-MoE README lists the original hardware as an M3 Max with a 1 TB SSD at "17.5 GB/s sequential read (measured)". — [source](https://github.com/Anemll/flash-moe)
- Qwen3.5-397B-A17B reads about 1.62 GB per token at 4-bit and 1.31 GB at Q3 from SSD, before any cache hit. — source: `asserted`
- At Apple's 14.5 GB/s the zero-cache ceiling for 397B is about 8.9 tok/s at 4-bit and 11.1 tok/s at Q3. — source: `asserted`
- At 5.5 GB/s cold read the zero-cache ceiling for 397B is about 3.4 tok/s at 4-bit and 4.2 tok/s at Q3. — source: `asserted`
- Anemll's 0.74 ms per 27 MB (36 GB/s) and 0.44 ms per 21.8 MB (50 GB/s) exceed Apple's 14.5 GB/s raw ceiling and so include page-cache hits. — source: `asserted`
- The Flash-MoE paper's 20% per generation SSD growth assumption is 5x below Apple's stated M4-to-M5 step of up to 2x. — source: `asserted`

## Corrections and disagreements

- CONTRADICTS: flash-moe-paper-flash-moe-pdf-90-plus-experiment.md and the Flash-MoE README, which describe a 1 TB M3 Max with "17.5 GB/s sequential read (measured)". Apple's own peak for the newer M5 Max with 8 TB is 14.5 GB/s, and third-party M4 Max Blackmagic reads are 5 to 7 GB/s. A 17.5 GB/s figure on an M3 Max 1 TB is above every first-party and third-party raw rate found, so it likely includes page-cache hits. The same repo's cold rate on that machine (5.5 GB/s with `F_NOCACHE`) is consistent with the raw rates. — [source](https://github.com/Anemll/flash-moe)
- CONTRADICTS: moe-expert-offload-to-ssd-on-macos.md, which reads the Anemll 0.74 ms per 27 MB layer as "about 36 GB/s effective" and treats it as an M5-generation SSD gain. The Anemll table also lists 0.44 ms for 21.8 MB at Q3, which is about 50 GB/s. Both exceed Apple's 14.5 GB/s raw ceiling by 2.5x to 3.4x and can only be cache-served reads. (The Anemll docs already say these runs are warm; this adds the Apple ceiling as a hard bound.) — [source](https://github.com/Anemll/flash-moe)
