<!-- llms-explorer concept facts · https://llms-explorer.com/tree/metal-attention-multi-row-verify-cliff-at-6-15-q/ · pack 2026-10-05 · ~1685 tokens -->

# Metal attention multi-row verify cliff at 6-15 query rows

> v0.14.0: verify width adapts to context depth by calibrating a per-width depth slope.

Parent: [Mac local LLMs: Speculative decoding and MTP](https://llms-explorer.com/tree/mac-local-llms-speculative-decoding-and-mtp/) · 2 facets · 33 facts · page: https://llms-explorer.com/tree/metal-attention-multi-row-verify-cliff-at-6-15-q/

## Facts

- v0.14.0: verify width adapts to context depth by calibrating a per-width depth slope. — source: `asserted`
- v0.15.0: `--sdpa-split` splits a wide verify into calls of 5 rows or fewer. — source: `asserted`
- 2026-09 engine pass: a multi-row attention kernel replaces the split wherever a probe admits it. — source: `asserted`
- 27 Sep 2026: v0.20.1 widens the set of GPUs that keep the kernel on. — source: `asserted`
- MTPLX 2.11 ships its own flash-decoding verify attention kernel for the 27B. — source: `asserted`
- The kernel is gated off on M5 until verified there; `MLX_DSPARK_FORCE_MULTIROW=1` overrides. — source: `asserted`
- Output stays greedy-correct, but ids can differ at floating-point ties. — source: `asserted`
- Quantized KV (`--kv-bits 8`) attention was slower than bf16 at 32k at the widths that matter. — source: `asserted`
- DSpark drafters lose acceptance on contexts far longer than they were trained on, a separate depth failure. — source: `asserted`
- No cached source explains why the cliff starts near 6 rows (kernel selection, register or threadgroup limits); the 6-15 range is stated, not derived. — source: `asserted`
- The cliff location is probably hardware dependent, because mlx-dspark enables its split only where a probe finds it. — source: `asserted`
- No independent reproduction of mlx-dspark's attention numbers exists in the cached pages. — source: `asserted`
- MLX's decode attention gives every query row its own pass over the KV cache, so a width-8 verify at 32k context reads the cache about 8 times. — [source](https://github.com/ARahim3/mlx-dspark)
- Past rows times GQA heads per KV head above 32, MLX attention falls to a much slower unfused path. — [source](https://github.com/ARahim3/mlx-dspark)
- mlx-dspark v0.14.0 calibrates a per-width depth slope once per drafter-target pair; both the derived default cap and `--max-draft auto` price it, short prompts keep the chat-depth optimum, and an explicit cap is never overridden. — [source](https://github.com/ARahim3/mlx-dspark)
- Before v0.3.1 the mlx-dspark drafter had a depth-scaling bug from redundant GQA/KV tiling, fixed bit-identically; the 2026-08-20 finding about verify cost explains reports that DFlash 2 and DSpark slow down at coding-agent context sizes. — [source](https://github.com/ARahim3/mlx-dspark)
- mlx-dspark v0.15.0 added `--sdpa-split`, which splits a wide verify into calls of 5 rows or fewer so each stays on the fast path, is on only where a one-time probe finds the cliff, and is lossless. — [source](https://github.com/ARahim3/mlx-dspark)
- With the split, the adaptive cap can stay wide on high-acceptance long content, giving about 1.3x at about 14k tokens of context. — [source](https://github.com/ARahim3/mlx-dspark)
- The 2026-09 mlx-dspark multi-row attention kernel supersedes the split for every attention shape a one-time probe admits; 8 rows at 32k on Qwen3.8-27B drop from 7.5 ms to 1.6 ms per call. — [source](https://github.com/ARahim3/mlx-dspark)
- The kernel packs every GQA head and query row of a KV head onto the matrix units so each KV tile is read once; it gives 1.5-4x on those attention calls and about 1.2x on a whole 32k-deep 27B verify, and it also covers the DSpark drafter's own context attention. — [source](https://github.com/ARahim3/mlx-dspark)
- The multi-row attention kernel is enabled per attention shape by a one-time race plus numerics probe, is gated off on M5 until verified there (`MLX_DSPARK_FORCE_MULTIROW=1` overrides), and is disabled with `--no-multirow-attn` or a per-swap `multirow_attn` on `/admin/load`. — [source](https://github.com/ARahim3/mlx-dspark)
- mlx-dspark v0.20.1, dated 27 Sep 2026, is titled "the multi-row attention kernel stays on for more GPUs". — [source](https://github.com/ARahim3/mlx-dspark)
- Before the depth changes, DSpark on Qwen3.8-27B could fall below plain decoding at 32k; the plain baseline itself slows from about 15 tok/s at 2k to 12.9 at 32k. — [source](https://github.com/ARahim3/mlx-dspark)
- mlx-dspark explains lower speedup at depth structurally: attention over a long KV cache is extra verify work that a single decode step pays only once. — [source](https://github.com/ARahim3/mlx-dspark)
- mlx-dspark limits a DSpark drafter to the last 4096 context rows by default (`--drafter-window`), because a DimInfer head's acceptance fell from 2.6 at 2k to 2.0 at 32k; heads trained with their own window (DFlash 2, Red Hat's Qwen3.8 head at 2048 trained on 8k) keep it. — [source](https://github.com/ARahim3/mlx-dspark)
- At 32k, mlx-dspark's quantized-KV (`--kv-bits 8`) attention kernel is slightly slower than bf16 at the widths that matter, so it is a RAM lever, not a speed lever. — [source](https://github.com/ARahim3/mlx-dspark)
- Qwen3.8-27B's KV cache costs 0.086 GB per 1k tokens (about 11 GB at 128k) at both 4-bit and 8-bit weights, because the cache is bf16. — [source](https://github.com/ARahim3/mlx-dspark)
- mlx-dspark states that MLX's unquantized matmul paid a cost cliff of about 2x at verify width 2, and that mlx 0.32.1's `gemv_wide` removed it so bf16-native families such as LFM2.5 became viable. — [source](https://github.com/ARahim3/mlx-dspark)
- MTPLX 2.11's TensorOps flash-decoding kernel cuts verify attention per layer by 35% at 72.7k context at half the power. — [source](https://mtplx.com/releases/)
- The MTPLX 2.11 flash-decoding route alone lifts the 27B from 38.6 to 41.4 tok/s at 16k and from 20.5 to 23.8 at 88k; 30.9 tok/s at 88k needs two opt-in long-context settings. — [source](https://mtplx.com/releases/)
- MTPLX lists `sdpa_nax_flash` (flash-decoding verify attention, 2.11) beside `verify_qmv` and a fused GDN verify kernel among its custom Metal kernels. — [source](https://mtplx.com/how-it-works/)
- MTPLX 2.12.1 keeps Flash-Next's compiled verifier active past 140K tokens, and all 50,769 verify steps of a 204,981-token Pi session ran compiled. — [source](https://mtplx.com/releases/)

## Corrections and disagreements

- CONTRADICTS: mlx-quantized-matmul-small-m-verify-kernels.md ("bf16 matmul is flat from M=2"): mlx-dspark reports MLX's unquantized matmul paid a cost cliff of about 2x at verify width 2 until mlx 0.32.1 added `gemv_wide`. Issue 4265's bf16 control and mlx-dspark's sweep may differ by shape or mlx version; neither is retracted. — source: `asserted`
