Metal attention multi-row verify cliff at 6-15 query rows
Parent: Mac local LLMs: Speculative decoding and MTP · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
v0.14.0: verify width adapts to context depth by calibrating a per-width depth slope.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- v0.14.0: verify width adapts to context depth by calibrating a per-width depth slope. [source]
- v0.15.0: `--sdpa-split` splits a wide verify into calls of 5 rows or fewer. [source]
- 2026-09 engine pass: a multi-row attention kernel replaces the split wherever a probe admits it. [source]
- 27 Sep 2026: v0.20.1 widens the set of GPUs that keep the kernel on. [source]
- MTPLX 2.11 ships its own flash-decoding verify attention kernel for the 27B. [source]
- The kernel is gated off on M5 until verified there; `MLX_DSPARK_FORCE_MULTIROW=1` overrides. [source]
- Output stays greedy-correct, but ids can differ at floating-point ties. [source]
- Quantized KV (`--kv-bits 8`) attention was slower than bf16 at 32k at the widths that matter. [source]
- DSpark drafters lose acceptance on contexts far longer than they were trained on, a separate depth failure. [source]
- No cached source explains why the cliff starts near 6 rows (kernel selection, register or threadgroup limits); the 6-15 range is stated, not derived. [source]
- The cliff location is probably hardware dependent, because mlx-dspark enables its split only where a probe finds it. [source]
- No independent reproduction of mlx-dspark's attention numbers exists in the cached pages. [source]
- MLX's decode attention gives every query row its own pass over the KV cache, so a width-8 verify at 32k context reads the cache about 8 times. [source]
- Past rows times GQA heads per KV head above 32, MLX attention falls to a much slower unfused path. [source]
- mlx-dspark v0.14.0 calibrates a per-width depth slope once per drafter-target pair; both the derived default cap and `--max-draft auto` price it, short prompts keep the chat-depth optimum, and an explicit cap is never overridden. [source]
- Before v0.3.1 the mlx-dspark drafter had a depth-scaling bug from redundant GQA/KV tiling, fixed bit-identically; the 2026-08-20 finding about verify cost explains reports that DFlash 2 and DSpark slow down at coding-agent context sizes. [source]
- mlx-dspark v0.15.0 added `--sdpa-split`, which splits a wide verify into calls of 5 rows or fewer so each stays on the fast path, is on only where a one-time probe finds the cliff, and is lossless. [source]
- With the split, the adaptive cap can stay wide on high-acceptance long content, giving about 1.3x at about 14k tokens of context. [source]
- The 2026-09 mlx-dspark multi-row attention kernel supersedes the split for every attention shape a one-time probe admits; 8 rows at 32k on Qwen3.8-27B drop from 7.5 ms to 1.6 ms per call. [source]
- The kernel packs every GQA head and query row of a KV head onto the matrix units so each KV tile is read once; it gives 1.5-4x on those attention calls and about 1.2x on a whole 32k-deep 27B verify, and it also covers the DSpark drafter's own context attention. [source]
- The multi-row attention kernel is enabled per attention shape by a one-time race plus numerics probe, is gated off on M5 until verified there (`MLX_DSPARK_FORCE_MULTIROW=1` overrides), and is disabled with `--no-multirow-attn` or a per-swap `multirow_attn` on `/admin/load`. [source]
- mlx-dspark v0.20.1, dated 27 Sep 2026, is titled "the multi-row attention kernel stays on for more GPUs". [source]
- Before the depth changes, DSpark on Qwen3.8-27B could fall below plain decoding at 32k; the plain baseline itself slows from about 15 tok/s at 2k to 12.9 at 32k. [source]
- mlx-dspark explains lower speedup at depth structurally: attention over a long KV cache is extra verify work that a single decode step pays only once. [source]
- mlx-dspark limits a DSpark drafter to the last 4096 context rows by default (`--drafter-window`), because a DimInfer head's acceptance fell from 2.6 at 2k to 2.0 at 32k; heads trained with their own window (DFlash 2, Red Hat's Qwen3.8 head at 2048 trained on 8k) keep it. [source]
- At 32k, mlx-dspark's quantized-KV (`--kv-bits 8`) attention kernel is slightly slower than bf16 at the widths that matter, so it is a RAM lever, not a speed lever. [source]
- Qwen3.8-27B's KV cache costs 0.086 GB per 1k tokens (about 11 GB at 128k) at both 4-bit and 8-bit weights, because the cache is bf16. [source]
- mlx-dspark states that MLX's unquantized matmul paid a cost cliff of about 2x at verify width 2, and that mlx 0.32.1's `gemv_wide` removed it so bf16-native families such as LFM2.5 became viable. [source]
- MTPLX 2.11's TensorOps flash-decoding kernel cuts verify attention per layer by 35% at 72.7k context at half the power. [source]
- The MTPLX 2.11 flash-decoding route alone lifts the 27B from 38.6 to 41.4 tok/s at 16k and from 20.5 to 23.8 at 88k; 30.9 tok/s at 88k needs two opt-in long-context settings. [source]
- MTPLX lists `sdpa_nax_flash` (flash-decoding verify attention, 2.11) beside `verify_qmv` and a fused GDN verify kernel among its custom Metal kernels. [source]
- MTPLX 2.12.1 keeps Flash-Next's compiled verifier active past 140K tokens, and all 50,769 verify steps of a 204,981-token Pi session ran compiled. [source]
Corrections and disagreements
- CONTRADICTS: mlx-quantized-matmul-small-m-verify-kernels.md ("bf16 matmul is flat from M=2"): mlx-dspark reports MLX's unquantized matmul paid a cost cliff of about 2x at verify width 2 until mlx 0.32.1 added `gemv_wide`. Issue 4265's bf16 control and mlx-dspark's sweep may differ by shape or mlx version; neither is retracted. [source]
Children
- No children recorded.