<!-- llms-explorer concept facts · https://llms-explorer.com/tree/adaptive-block-size-scheduling-for-dflash-at-inf/ · pack 2026-10-05 · ~658 tokens -->

# Adaptive block-size scheduling for DFlash at inference

> D-cut does not change the drafter's block size; it truncates the verified prefix, so the drafter still runs at full width and only verification shrinks.

Parent: [Mac local LLMs: Speculative decoding and MTP](https://llms-explorer.com/tree/mac-local-llms-speculative-decoding-and-mtp/) · 2 facets · 12 facts · page: https://llms-explorer.com/tree/adaptive-block-size-scheduling-for-dflash-at-inf/

## Facts

- D-cut does not change the drafter's block size; it truncates the verified prefix, so the drafter still runs at full width and only verification shrinks. — [source](https://arxiv.org/html/2607.14647v1)
- That matches the DFlash paper's finding that draft time is nearly flat in block size for moderate blocks, so truncation saves verify cost and not draft cost. — [source](https://arxiv.org/html/2602.06036v2)
- Feb 2026: DFlash paper (v2 cached) names dynamic block-size scheduling as enabled by large-block training. — [source](https://arxiv.org/html/2602.06036v2)
- 30 Apr 2026: DFlash accepted to ICML 2026 (OpenReview forum Oz335dV48X, last modified 22 Sep 2026). — [source](https://openreview.net/forum?id=Oz335dV48X)
- 16 Jul 2026: D-cut posted. — [source](https://arxiv.org/html/2607.14647v1)
- Block 16 unpruned is worse than block 8 under load on dense models: Qwen3.5-27B at concurrency 32 averages 0.94x for block 16 and 1.21x for block 8. — [source](https://arxiv.org/html/2607.14647v1)
- The SGLang DFlash v2 post's sample command and benchmark block size mismatch (already recorded) is explained by this effect: a static choice is concurrency-dependent. — source: `asserted`
- A published adaptive scheduler for DFlash verification depth exists (D-cut, vLLM PR 47131). — [source](https://arxiv.org/html/2607.14647v1)
- D-cut keeps the drafter at full block width and prunes only the verified prefix. — [source](https://arxiv.org/html/2607.14647v1)
- DFlash block 16 averaged 0.94x and block 8 averaged 1.21x of autoregressive throughput on Qwen3.5-27B at concurrency 32. — [source](https://arxiv.org/html/2607.14647v1)
- DFlash's ICML 2026 OpenReview entry lists Jian Chen, Yesheng Liang and Zhijian Liu. — [source](https://openreview.net/forum?id=Oz335dV48X)

## Corrections and disagreements

- CONTRADICTS: draft-length-tuning-spec-draft-n-max-for-block-d.md Open questions line "the sources show no PR" for adaptive scheduling. D-cut (arXiv 2607.14647, vLLM PR 47131) schedules verified depth per request from drafter confidence and a startup cost table, using DFlash block 16 drafters with D = 15. — [source](https://arxiv.org/html/2607.14647v1)
