Draft length tuning (spec-draft-n-max) for block-diffusion drafters

Parent: Mac local LLMs: Speculative decoding and MTP · Published reference · snapshot 2026-10-05

↓ Facts as markdownall context files

Drafting cost is nearly flat in length: DFlash's paper says that for moderate block sizes the draft time is largely insensitive to gamma, so shortening the block saves verify cost, not draft cost.

These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.

Facts

Children

← the whole tree · 3D view· how to read this page