MTPLX compiled verifier (GraphBank) windows and eager fallback by context length
Parent: Mac local LLMs: Speculative decoding and MTP · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
The compiled route needs preallocated cache buffers. When an answer or conversation outgrows them, the verifier either grows them (checked against memory) or hands generation to the eager path.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- The compiled route needs preallocated cache buffers. When an answer or conversation outgrows them, the verifier either grows them (checked against memory) or hands generation to the eager path. [source]
- Reasons a request runs eager: context beyond the compiled window (32,768 tokens for the dense models), a buffer-growth fallback that memory refuses, warm-restore rounds before compilation, eager copy, repair or final-save forwards inside a compiled request, prompt shapes the compiled route cannot represent exactly, and (before 2.12.0) any image request. [source]
- `/health` reports a `degradation` block with the compiled-verify state and the reason for a permanent-eager fallback; `mtplx doctor` prints the compiled-verify fence with its mode, threshold and source. [source]
- Eager and compiled rounds are kept numerically aligned: eager BF16 verification uses the same attention-gate rounding as the compiled graph. [source]
- 2.0.0 (6 Jul 2026): turbo profile defaults to compiled verify with a per-model gate; the q8 Quality pack stays on the eager path because it measures best there. [source]
- 2.5.2: a handoff slowdown bug is fixed (below). 2.5.3: an opt-in eager round after a very large restore is added. [source]
- 2.7.0 (15 Aug): window moves from 12,288 to 32,768 tokens. 2.8.0: `/health` degradation block and `doctor` fence line. [source]
- 2.11.3: depth-policy calibration bug fixed. 2.12.0: image requests get a compiled route. 2.12.1: growth reserve fixed for quantized KV; Flash-Next resizes buffers layer by layer. [source]
- Compiled-to-eager handoff once the growth reserve ran out carried unfinished GPU work into the rest of the response, so decode fell steadily on long answers. [source]
- Stale operator exports of `MTPLX_LAZY_TARGET_DISTRIBUTIONS=1` and `MTPLX_LAZY_BONUS_VERIFY=1` switched off the batched compiled verifier on Flash-Next. [source]
- A one-off 125 ms first-call timing made the depth policy choose the slower eager path for 453 of 464 cycles. [source]
- A resize that memory refuses sends only that turn to the eager verifier. [source]
- Do not confuse the verify window with the 12,288-token fence on chained greedy drafting; they are separate settings. [source]
- Image requests: before 2.12.0 notes say image-bearing turns decode eagerly at 33 to 50 tok/s against 54 to 61 text-only; 2.12.0 says dense 27B image requests now take the compiled route with zero logit, hidden-state or cache differences on 762 compiled rounds. The 2.12.0 text also reports a seeded image reply that diverged from the eager route after 61 tokens at a rounding difference, so "exact" there means "differs only by rounding". [source]
- Whether the "Sustained" presets (Qwen 3.5 4B, MiMo 9B, Gemma 4) use the compiled verifier at all; the README names compiled verify only for the Turbo preset. [source]
- Compiled-route behavior on M1 to M3: the 2.7.0 measurements all ran on one M5 Max, and the notes say nothing was measured on M1 or M2. [source]
- The GraphBank's own shape set and memory cost; only the how-it-works page names it. [source]
- The how-it-works page says the target evaluates all K positions in one forward "via GraphBank-compiled verify shapes". [source]
- The README says the Turbo preset is "NAX verify kernels + compiled verify" and is picked automatically for the quantized 27B and 9B flagship models. [source]
- The 2.0.0 notes say the default agent daemon profile is turbo with the compiled-verify per-model gate and that the q8 Quality pack stays on the eager verify path it measures best on. [source]
- In 2.5.2 the compiled verifier handed generation to the eager path after its growth reserve was exhausted and the handoff carried unfinished GPU work, so a 12,000-token response opened at 59 tok/s and ended near 30. [source]
- The 2.5.2 fix settles all cache and recurrent state once at the ownership boundary and releases compiled references before eager decoding continues, with handoff telemetry in the compiled verifier stats. [source]
- In 2.5.3 an opt-in `MTPLX_COMPILED_VERIFY_POST_RESTORE_EAGER_ROUNDS` can route the first verify round after a very large restore through the eager path and ships off. [source]
- In 2.8.0 `/health` gained a `degradation` block that reports compiled-verify state with the reason for any permanent-eager fallback, and `mtplx doctor` prints the compiled-verify fence with its mode, threshold and the profile or env that supplied it (#255). [source]
- `MTPLX_COMPILED_VERIFY` can be set by hand, and `MTPLX_COMPILED_VERIFY=parity2` compares every compiled round against the eager forward and records per-leaf differences. [source]
- Chained greedy drafting is on by default for temperature 0 below 12,288 prompt tokens because A/B runs gave +2.5 to +9.8% decode at 0.5k to 8k and -2.9% and -2.7% at 16k and 32k, tunable with `MTPLX_GREEDY_TRIO_MAX_CONTEXT`; this fence is separate from the compiled-verify window. [source]
- In 2.11.3 a depth-policy cost estimate seeded with a one-off 125 ms first call (the next three took 30 to 31 ms) chose the eager path for 453 of 464 cycles, and calibrating on the minimum of four warm-up samples took a 108,919-token OpenCode turn from 48.83 to 61.77 tok/s. [source]
- In 2.11.3, stale `MTPLX_LAZY_TARGET_DISTRIBUTIONS=1` and `MTPLX_LAZY_BONUS_VERIFY=1` exports from older tuning switched off the batched compiled verifier on Flash-Next, and both exports were removed. [source]
- Before 2.12.0, an image request on Flash-Next fell back to eager verification for the whole conversation, and image-bearing turns decoded on the eager verifier at 33 to 50 tok/s against 54 to 61 text-only. [source]
- 2.12.0 gives the compiled verifier a rotary origin per request with the image position delta as an int32 graph input, so requests with different images share one compiled program, and six prompt shapes the route cannot represent exactly stay eager by name. [source]
- `MTPLX_QWEN4_VISION_COMPILED_VERIFY=0` turns the compiled image route off. [source]
- On 440 image rounds and 429 text rounds the compiled and eager routes differ only by rounding, with mean KL per round of 0.0015 for images and 0.0013 for text. [source]
- 2.12.0 makes eager BF16 verification use the compiled graph's attention-gate rounding for dense image requests, contexts beyond the compiled route's 32,768-token limit, buffer-growth fallbacks, warm-restore rounds before compilation, and eager copy, repair or final-save forwards inside a compiled request. [source]
- With `MTPLX_ONE_COPY` (default on for Flash-Next), the compiled verifier adopts the buffers the prompt was read into instead of copying them into a padded buffer, and buffers that must grow are resized one layer at a time (0.36 to 0.57 GB at the 147K edge against 4.2 GB for a padded copy). [source]
- A Flash-Next turn that cannot afford the growth verifies eagerly for that turn only, and the next turn resizes the rest. [source]
- The 2.12.1 #526 fix gives the compiled verifier's promoted paged cache a window reserve checked against Mac memory and refused with a 507 before a row moves, replacing a cache sized once at prompt plus 16,384 tokens. [source]
- The `MTPLX_FIXED_M4_COPY_WINDOWS=1` option replays Flash-Next's context-copy rounds on the compiled verifier, is bit-identical to the eager window in tests, and ships off because it was not checked on the full model. [source]
- The 2.0.0 long-context wave added a packed-GQA verify kernel and commit-first KV donation in compiled verify, measured on an M5 Max at 64k decode +12%, 128k decode from 17 to above 20 tok/s, and peak memory down 8 GB at 64k and 16 GB at 128k. [source]
Children
- No children recorded.