Adaptive block-size scheduling for DFlash at inference
Parent: Mac local LLMs: Speculative decoding and MTP · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
D-cut does not change the drafter's block size; it truncates the verified prefix, so the drafter still runs at full width and only verification shrinks.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- D-cut does not change the drafter's block size; it truncates the verified prefix, so the drafter still runs at full width and only verification shrinks. [source]
- That matches the DFlash paper's finding that draft time is nearly flat in block size for moderate blocks, so truncation saves verify cost and not draft cost. [source]
- Feb 2026: DFlash paper (v2 cached) names dynamic block-size scheduling as enabled by large-block training. [source]
- 30 Apr 2026: DFlash accepted to ICML 2026 (OpenReview forum Oz335dV48X, last modified 22 Sep 2026). [source]
- 16 Jul 2026: D-cut posted. [source]
- Block 16 unpruned is worse than block 8 under load on dense models: Qwen3.5-27B at concurrency 32 averages 0.94x for block 16 and 1.21x for block 8. [source]
- The SGLang DFlash v2 post's sample command and benchmark block size mismatch (already recorded) is explained by this effect: a static choice is concurrency-dependent. [source]
- A published adaptive scheduler for DFlash verification depth exists (D-cut, vLLM PR 47131). [source]
- D-cut keeps the drafter at full block width and prunes only the verified prefix. [source]
- DFlash block 16 averaged 0.94x and block 8 averaged 1.21x of autoregressive throughput on Qwen3.5-27B at concurrency 32. [source]
- DFlash's ICML 2026 OpenReview entry lists Jian Chen, Yesheng Liang and Zhijian Liu. [source]
Corrections and disagreements
- CONTRADICTS: draft-length-tuning-spec-draft-n-max-for-block-d.md Open questions line "the sources show no PR" for adaptive scheduling. D-cut (arXiv 2607.14647, vLLM PR 47131) schedules verified depth per request from drafter confidence and a startup cost table, using DFlash block 16 drafters with D = 15. [source]
Children
- No children recorded.