<!-- llms-explorer concept facts · https://llms-explorer.com/tree/dflash-and-ddtree-speculative-decoding-drafters/ · pack 2026-10-05 · ~1860 tokens -->

# DFlash and DDTree speculative decoding drafters

> Each round starts with a "bonus" token the target already chose but has not yet run on. The drafter sees a masked block [bonus, mask, ..., mask] and returns one marginal distribution per future position.

Parent: [Mac local LLMs: Speculative decoding and MTP](https://llms-explorer.com/tree/mac-local-llms-speculative-decoding-and-mtp/) · 1 facets · 31 facts · page: https://llms-explorer.com/tree/dflash-and-ddtree-speculative-decoding-drafters/

## Facts

- Each round starts with a "bonus" token the target already chose but has not yet run on. The drafter sees a masked block [bonus, mask, ..., mask] and returns one marginal distribution per future position. — source: `asserted`
- Those marginals are not path-conditioned: position i does not condition on the tokens chosen at positions 1..i-1 of the same block. DDTree therefore optimizes a surrogate, the factorized product of the marginals, not the target's real continuation probabilities. — source: `asserted`
- Under the surrogate, expected accepted length equals the sum of prefix probabilities over the tree's nodes (Proposition 1). The best tree under a node budget B is the B highest-probability prefixes, and those prefixes are automatically prefix-closed (Proposition 2). — source: `asserted`
- A max-heap best-first search finds them without enumerating the vocabulary power set. Each popped rank tuple spawns its next sibling and its first child, and the search is limited to the top K = min(B, vocab) tokens per depth. — source: `asserted`
- Verification flattens the tree into one sequence rooted at the bonus token, assigns position ids by tree depth, and uses an ancestor-only attention mask. The verifier then walks the tree with the target's own decoding rule (greedy or sampled) and keeps the matched path. The first unmatched target token becomes the next bonus token, and the KV cache is compacted to the accepted path. — source: `asserted`
- 14 Apr 2026: the DDTree paper appears on arXiv (Ringel and Romano, Technion). 15 Apr 2026: a DGX Spark forum thread proposes DDTree plus DFlash for the GB10 and reports a vLLM port. — source: `asserted`
- 24 Apr 2026: the llama.cpp DFlash PR author is asked about DDTree and replies it would come after the DFlash PR merges. — source: `asserted`
- Larger trees stop paying: acceptance length keeps rising with budget, but verifier cost wins above a budget of about 256 to 512 on the paper's GPU setup. — source: `asserted`
- The surrogate ignores dependence between positions. The paper states the optimality proof holds for the drafter's factorized distribution, not for the true target distribution. — source: `asserted`
- On a hybrid (recurrent) target, tree verification needs per-branch recurrent state. The ddtree-mlx Metal kernel that forks state is held in the existing dossier. — source: `asserted`
- Best budget: the paper's GPU runs peak at 256 to 512 nodes on Qwen3-8B with a 16-token block. Existing dossier speculative-decoding-and-mtp-on-apple-silicon.md reports budget 4 optimal for Qwen3.5 hybrids on a Mac M3 Ultra. The two are different hardware, targets and kernels, so neither refutes the other. — source: `asserted`
- Gain size: the paper reports +30 to 65% speedup over DFlash on H200 (for example 5.56x to 7.27x). The Mac ddtree-mlx run reports +10 to 15% over DFlash. The cause is not isolated in any source. — source: `asserted`
- Whether llama.cpp has a DDTree implementation. No cached page shows a PR; the PR 22105 author only said it would follow. — source: `asserted`
- Whether tree verification of a MoE target pays on a Mac, given that MoE verification cost grows with unique experts touched (see the expert-overlap dossier). — source: `asserted`
- The DDTree paper (arXiv 2604.12989, Ringel and Romano, 14 Apr 2026) builds a draft tree from the per-position distributions of one block-diffusion forward pass and verifies it in one target forward pass with an ancestor-only attention mask. — [source](https://arxiv.org/html/2604.12989v1)
- A DFlash drafter's output at each block position is a marginal distribution that does not condition on earlier tokens in the same block, so DDTree optimizes the factorized product of those marginals as a surrogate for the target's acceptance length. — [source](https://arxiv.org/html/2604.12989v1)
- Under that surrogate, expected acceptance length equals the sum of prefix probabilities over the tree, and the optimal tree is the B most probable prefixes, which are automatically prefix-closed. — [source](https://arxiv.org/html/2604.12989v1)
- DDTree finds the optimal tree with a max-heap best-first search over rank tuples, where a popped tuple spawns its next sibling and its first child, limited to the top min(B, vocabulary) tokens per depth. — [source](https://arxiv.org/html/2604.12989v1)
- The bonus root token does not count toward the node budget. — [source](https://arxiv.org/html/2604.12989v1)
- DDTree differs from OPT-Tree, which needs one drafter forward pass per tree depth, and from DART, which prunes with an external n-gram continuity score and an n-gram trie. — [source](https://arxiv.org/html/2604.12989v1)
- After verification the KV cache is compacted to the accepted path, and the first unmatched target token becomes the next round's bonus token. — [source](https://arxiv.org/html/2604.12989v1)
- The paper's runs use block size 16, node budgets 16 to 1024, bfloat16, 2,048 new tokens, 8 H200 GPUs and Hugging Face Transformers, with Qwen3-4B, Qwen3-8B and Qwen3-Coder-30B-A3B-Instruct. — [source](https://arxiv.org/html/2604.12989v1)
- At temperature 0 on Qwen3-4B, DDTree lifts DFlash from 5.56x to 7.27x on AIME 2024 (mean accepted length 7.54 to 10.37) and from 2.03x to 3.32x on Alpaca (3.11 to 5.35). — [source](https://arxiv.org/html/2604.12989v1)
- On the MoE target Qwen3-Coder-30B-A3B at temperature 0, DDTree lifts HumanEval from 6.09x to 8.22x and Alpaca from 1.53x to 2.46x. — [source](https://arxiv.org/html/2604.12989v1)
- At temperature 1.0 on Qwen3-4B, DDTree lifts AIME 2024 from 3.50x to 5.31x and SWE-bench Lite from 2.29x to 3.71x. — [source](https://arxiv.org/html/2604.12989v1)
- The paper reports DDTree beats vanilla DFlash in all 60 dataset, model and temperature settings. — [source](https://arxiv.org/html/2604.12989v1)
- On MATH-500 with Qwen3-8B, speedup peaks at node budgets of 256 to 512, and a budget of 1024 raises acceptance length but lowers speedup because verifier cost outweighs the gain. — [source](https://arxiv.org/html/2604.12989v1)
- The paper's acceptance-length histogram at budget 512 shows mass shifting toward full-block acceptance compared with vanilla DFlash. — [source](https://arxiv.org/html/2604.12989v1)
- A DGX Spark forum thread of 15 Apr 2026 says a community member ported DDTree plus DFlash into vLLM and reported 80+ tok/s for Qwen3.5-27B AWQ on a GB10; the thread relays this from X posts. — [source](https://forums.developer.nvidia.com/t/ddtree-plus-diffusion-drafting-dflash-to-optimize-gb10/366643)
- The DDTree cost the forum thread names is a small amount of extra memory for the drafted tree. — [source](https://forums.developer.nvidia.com/t/ddtree-plus-diffusion-drafting-dflash-to-optimize-gb10/366643)
- On 24 Apr 2026 the llama.cpp DFlash PR author answered a DDTree question with "I'd expect it to come after this PR gets merged". — [source](https://github.com/ggml-org/llama.cpp/pull/22105)
