<!-- llms-explorer concept facts · https://llms-explorer.com/tree/mtplx-flash-decoding-verify-attention-kernel-sdp/ · pack 2026-10-05 · ~1963 tokens -->

# MTPLX flash-decoding verify attention kernel (sdpa_nax_flash)

> The route runs verify attention as a TensorOps kernel with dimension-split blocks. Block defaults came from sweeps at 72.7k and 128k context.

Parent: [Mac local LLMs: Speculative decoding and MTP](https://llms-explorer.com/tree/mac-local-llms-speculative-decoding-and-mtp/) · 1 facets · 35 facts · page: https://llms-explorer.com/tree/mtplx-flash-decoding-verify-attention-kernel-sdp/

## Facts

- The route runs verify attention as a TensorOps kernel with dimension-split blocks. Block defaults came from sweeps at 72.7k and 128k context. — source: `asserted`
- It engages only once the KV buffer reaches 8,192 tokens. — source: `asserted`
- A packed-GQA verify kernel (from 2.10.2) serves every request the route does not. — source: `asserted`
- `/health` reports the route under `degradation.nax`: availability, bail counts by reason (including a GPU-family-or-OS bail) and, since 2.11.2, flash dispatch counters. — source: `asserted`
- 2.10.x: the 27B flash-decoding route and the QSA sparse-prefill kernels use cooperative-tensor operands. — source: `asserted`
- 2.11.0 (4 Sep 2026): the route turns on in Turbo with no GPU-family gate. — source: `asserted`
- 2.11.1 (3 Sep 2026 per CHANGELOG.md): the macOS 27 header fix; route switch `MTPLX_NAX_FLASH_ROUTE=1` with dimension-split block defaults. — source: `asserted`
- 2.11.2 (6 Sep 2026): the route is gated to M5-class GPUs on macOS 26.2 or newer after wrong output on M1-M4. — source: `asserted`
- 2.12.1 (2 Oct 2026): the pinned MLX 0.32.2 tensor-unit expert-matmul bug on M5 is worked around. — source: `asserted`
- Without the gate, the route returned wrong attention on M1-M4 GPUs from 8,192 tokens on: unrelated reasoning, imaginary tasks, mixed languages and no tool calls. — source: `asserted`
- On macOS 15 the kernel failed to compile because `MetalPerformancePrimitives.h` was missing. — source: `asserted`
- On macOS 27 betas, an older header rejected address-space-qualified cooperative-tensor operands and long prompts failed mid-request until the kernels used address-space-neutral types. — source: `asserted`
- A reported incoherent-output case on an M3 Ultra through Pi matches the 8,192-token engagement point but was not confirmed from that machine. — source: `asserted`
- The wrong-output class is silent: a hardware or compiler mismatch produced fluent text, not an error. — source: `asserted`
- Gain size by source: release notes quote +7% at 16k and +16% at 88k decode from the route alone; the changelog separately says the 27B on 48 GB Macs verifies past 32,768 tokens on the slower eager path. These measure different things (kernel route versus compiled verify) and are not contradictory, but a reader comparing "flash route on" across Macs can be misled. — source: `asserted`
- Per-chip speed of the flash route on M5 Pro, M5 Ultra and M3/M4 Macs is unmeasured: M1-M4 are served by the older kernel by design, and every real-model check ran on an M5 Max 128 GB. — source: `asserted`
- No source states why 8,192 tokens is the engagement point. — source: `asserted`
- MTPLX 2.11 turned the 27B flash-decoding verify route on in Turbo without a GPU-family gate; it engages once the KV buffer reaches 8,192 tokens and uses the M5 GPU's tensor units. — [source](https://github.com/youssofal/MTPLX/blob/main/CHANGELOG.md)
- On M1-M4 GPUs the ungated route returned wrong attention (unrelated reasoning, imaginary tasks, mixed languages, no tool calls), and on macOS 15 the kernel failed to build because `MetalPerformancePrimitives.h` was not found. — [source](https://github.com/youssofal/MTPLX/blob/main/CHANGELOG.md)
- MTPLX 2.11.2 (6 Sep 2026) runs the route only on an M5-class GPU on macOS 26.2 or newer; every other Mac serves the packed kernel from 2.10.2, validated at startup. — [source](https://mtplx.com/releases/2.11.2/)
- The release notes call the 2.11.2 change "27B correct again on M1 to M4" and cite issues 459, 464, 467, 461 and 469. — [source](https://mtplx.com/releases/2.11.2/)
- `/health degradation.nax` reports `available`, the `gpu_family_or_os` bail counts and, new in 2.11.2, `flash_dispatch_counters` for an engaged route; the M5 Max rehearsal with the reporters' 14k and 32k diff-summary prompts counted 66 dispatches and no bails. — [source](https://github.com/youssofal/MTPLX/blob/main/CHANGELOG.md)
- The changelog says an M3 Ultra report of incoherent output with MTP on through Pi (item 3 of issue 455) matches the route's engagement point but was not confirmed from that machine, and that issue 455's AR slowdown is not explained by the gate. — [source](https://github.com/youssofal/MTPLX/blob/main/CHANGELOG.md)
- MTPLX 2.11.1 enables the 27B route in Turbo with `MTPLX_NAX_FLASH_ROUTE=1` and dimension-split block defaults taken from the 72.7k and 128k sweeps. — [source](https://github.com/youssofal/MTPLX/blob/main/CHANGELOG.md)
- On macOS 27 betas the Metal Performance Primitives header rejected the address-space-qualified cooperative-tensor operands used by the QSA sparse prefill kernels and the 27B flash-decoding route, so prompts past the roughly 32K sparse-prefill crossover failed mid-request with "Unable to build metal library from source"; seven kernel sites moved to address-space-neutral operand types and a startup probe dispatches the real sparse-prefill pipeline once. — [source](https://github.com/youssofal/MTPLX/blob/main/CHANGELOG.md)
- The README names Turbo as the mode with NAX verify kernels plus compiled verify, picked automatically for the quantized 27B and 9B flagship models. — [source](https://github.com/youssofal/MTPLX)
- MTPLX 2.9.2 stopped the NAX turbo verify path from using padded M=5 lanes, which measured slower than stock. — [source](https://mtplx.com/releases/2.9.2/)
- An experimental `MTPLX_VK_CROSSROW` crossrow wide-verify kernel exists, off by default, in 2.9.2. — [source](https://mtplx.com/releases/2.9.2/)
- In 2.12.0 MLX 0.32.2's tensor-unit kernel for expert-sorted quantized matmul skipped rows when one call routed more than 32,767 rows and the count was not a multiple of its tile, leaving leftover memory in those rows; Flash-Next's ten experts per token let a 3,277-4,095-token prompt chunk cross the limit, affecting 794 of every 4,096 cold prompt lengths above 3,277 tokens (19.4%). — [source](https://mtplx.com/releases/2.12.1/)
- MTPLX 2.12.1 pads and runs such a call so every real row is computed exactly as in a correct call, says M1-M4 Macs and MLX 0.32.3 were never affected, and pins pip and Homebrew installs to exactly MLX 0.32.2 because 0.32.3's Flash-Next output differs. — [source](https://mtplx.com/releases/2.12.1/)
- MTPLX 2.12.1's known issues say the 27B verifies on the slower eager path past 32,768 tokens of context. — [source](https://mtplx.com/releases/2.12.1/)
- An earlier MTPLX release moved the compiled-verify window from 12k to 32k tokens of context, giving +6.9% at 20k on Qwen 3.8 Bare Speed with peak memory flat at 20k. — [source](https://mtplx.com/releases/)
- MTPLX 2.11.3 reports that at 109k tokens compiled verify calls per 1,024 output tokens rose from 11 to 376 and peak memory stayed 95.4 GB, while at 200k the warm route stays on the eager verifier by design because the 5.7 GB the compiled verifier needs does not fit the admission budget. — [source](https://mtplx.com/releases/2.11.3/)
- MTPLX 2.11.2's verifier-depth policy measures draft and verify cost per depth on the running machine, which raised the share of cycles on Flash-Next's compiled verify route from about 11% to 90-96% in agent turns. — [source](https://mtplx.com/releases/2.11.2/)
- Every MTPLX real-model check cited for the 2.12 releases ran on one M5 Max with 128 GB, and M1-M4 behaviour was rehearsed on that M5 rather than on older GPUs. — [source](https://mtplx.com/releases/2.12.1/)
