<!-- llms-explorer concept facts · https://llms-explorer.com/tree/pre-m5-vs-m5-apple-gpu-decode-behavior-for-compr/ · pack 2026-10-05 · ~280 tokens -->

# Pre-M5 vs M5 Apple GPU decode behavior for compressed KV

> Apple's MLX study reports M5 generation is 19-27% faster than M4, explained by memory bandwidth (120 GB/s M4, 153 GB/s M5, +28%), while TTFT is compute-bound and up to 4x faster with the GPU Neural Accelerators.

Parent: [Mac local LLMs: KV cache sizing and quantization](https://llms-explorer.com/tree/mac-local-llms-kv-cache-sizing-and-quantization/) · 1 facets · 2 facts · page: https://llms-explorer.com/tree/pre-m5-vs-m5-apple-gpu-decode-behavior-for-compr/

## Facts

- Apple's MLX study reports M5 generation is 19-27% faster than M4, explained by memory bandwidth (120 GB/s M4, 153 GB/s M5, +28%), while TTFT is compute-bound and up to 4x faster with the GPU Neural Accelerators. — [source](https://machinelearning.apple.com/research/exploring-llms-mlx-m5)
- Inference: because M5 accelerators speed matmul but not decode bandwidth, a compressed-KV dequant penalty on pre-M5 ALU-bound GPUs (M1 L2 wall) is not expected to be rescued by the accelerators; no measurement of turbo/q4 KV on M5 vs M1 with identical models was found. — source: `asserted`
