Hardware-calibrated verification budget for batch-wide tree speculation
Parent: Mac local LLMs: Speculative decoding and MTP · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
The method D-cut (Tencent, arXiv 2607.14647, 16 Jul 2026) treats all draft tokens in a batch as one pool and keeps the top-K by estimated survival probability.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- The method D-cut (Tencent, arXiv 2607.14647, 16 Jul 2026) treats all draft tokens in a batch as one pool and keeps the top-K by estimated survival probability. [source]
- Survival probability of draft depth k is the prefix product of the drafter's token confidences, s_k = c_1 x ... x c_k, and the expected tokens advanced by keeping n drafts is the sum of s_0..s_n with s_0 = 1 for the bonus token. [source]
- Because s_k only falls with depth, a global top-K selection always yields a contiguous prefix per request. [source]
- Speedup is modelled as MAT(K) divided by T_spec(K)/T_AR; MAT comes from drafter confidence at runtime, T_spec(K) is a hardware property, and T_AR is constant per batch size so it drops out of the argmax. [source]
- K is restricted to ratios R = {0.25, 0.5, 0.75, 1.0} of B(D+1) positions with a floor of B (every request keeps its bonus token), so only four CUDA-graph shapes per batch size are captured. [source]
- At server start D-cut profiles a cost table C(B, rho) of end-to-end step time for each discrete batch shape; at serving time it picks the rho that maximises (sum of top-K scores) / C(B, rho). [source]
- The target's accept/reject logic is unchanged, so output matches the target distribution; only low-utility suffixes are dropped. [source]
- Prior adaptive methods changed per-request draft length or tree shape; ECHO (Hu et al. 2026) frames high-concurrency speculation as budgeted tree scheduling under a global verification cap with sparse confidence gating. [source]
- D-cut targets a different class (single linear blocks from DFlash) and also cites DSpark (trained confidence head) and Graft (refills the freed budget with retrieved candidates) as related. [source]
- The implementation is vLLM pull request 47131 and AngelSlim documentation. [source]
- Long fixed blocks lose to plain decoding at load: DFlash block 16 on Qwen3.5-27B reaches 0.57x to 1.14x of autoregressive speed at concurrency 32 across five benchmarks, average 0.94x, while D-cut (16) averages 1.25x. [source]
- No fixed ratio wins everywhere: on H20 the best fixed budget is 0.25 or 0.5; on H800 it moves to 0.5 for GSM8K and HumanEval, while MT-Bench still prefers aggressive pruning at high concurrency. [source]
- Verified depth costs differ by device: at batch 16, raising depth from 4 to 16 adds about 20 ms per step on H20 and about 1 ms on H800. [source]
- A per-step variable verification batch shape does not fit CPU-GPU overlap (Spec-V2) or full-plus-piecewise CUDA-graph capture, which assume a static shape; the authors list this as the main limitation. [source]
- On MoE targets the gain is uneven: Qwen3.5-35B-A3B changes little (2.36x vs 2.50x at block 16) while Hy3-295B-A21B gains from 1.83x to 2.48x. [source]
- For Apple silicon, the compute-bound crossover is far higher than on a datacentre GPU, so batch-wide budgets matter only for multi-user serving; a single-user Mac stays memory-bound. [source]
- The paper says pruning should be driven by profiled cost, not architecture labels, and declines to attribute Hy3's larger gain to GQA versus Qwen3.5's Gated DeltaNet hybrid layout (it calls the layout "one plausible contributor"). [source]
- DFlash's own paper says only that shrinking block size under compute-bound load "can" help and leaves scheduling to future work; D-cut chooses a different lever (pruning per-request depth) rather than changing the drafter block size. [source]
- D-cut selects one global verification budget K across a batch from drafter confidences and a startup-profiled latency table. [source]
- The budget ratio set is {0.25, 0.5, 0.75, 1.0} of B(D+1) with a floor of B. [source]
- Cost depends on runtime: depth 4 to 16 adds about 20 ms on H20 and about 1 ms on H800 at batch 16. [source]
- D-cut (16) beats DFlash (16) in 29 of 30 model-dataset configurations at concurrency 32 and lifts average speedup from 1.26x to 1.65x. [source]
- D-cut (8) beats DFlash (8) in 25 of 30 pairs, average 1.56x to 1.67x. [source]
- On H20 with Qwen3-8B at batch 32, GSM8K throughput was 2200 tok/s unpruned DFlash versus 3175 with auto budget; on H800 it was 7912 versus 9663. [source]
- D-cut (auto) stays at or near the best fixed ratio on both devices. [source]
- The evaluation used vLLM with max 64 sequences, greedy decoding, concurrencies 4 to 64, and six targets from 4B dense to 295B MoE. [source]
- At temperature 1 the relative gain of adaptive pruning is larger than under greedy decoding. [source]
- Variable per-step verify shapes conflict with static CUDA graphs and overlap scheduling. [source]
- Applicability to a single-user Mac is limited because decode there stays memory-bound at batch 1. [source]
Children
- No children recorded.