llama.cpp EAGLE-3 and DSpark speculator drafters
Parent: Mac local LLMs: Speculative decoding and MTP · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
EAGLE-3 (`draft-eagle3`): extracts features from the target at fixed layers, fuses them through one linear layer, and drafts with a single-layer decoder. Drafting is autoregressive, one token per draft step.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- EAGLE-3 (`draft-eagle3`): extracts features from the target at fixed layers, fuses them through one linear layer, and drafts with a single-layer decoder. Drafting is autoregressive, one token per draft step. [source]
- DSpark (`draft-dspark`): reuses the DFlash block-diffusion drafter, which emits a whole block in one pass, and adds a low-rank Markov head. The head biases each block position's logits by the previous drafted token, so the block is sampled left to right inside one graph. [source]
- Both are lossless for greedy decoding because every drafted token is verified by the target. [source]
- 12 Jun 2026: the EAGLE-3 PR 18039 merges as commit 88a3927. [source]
- 28 Jun 2026: DFlash merges (PR 22105); DSpark layers on it in PR 25173. [source]
- Review of PR 25173 removed a separate DSpark architecture: a DSpark draft now converts to a DFlash GGUF and the Markov tensors are found by presence. [source]
- PR 25173's first phase converted the confidence head but did not use it at inference; a later commit added `--spec-draft-conf-min`. [source]
- EAGLE-3 in llama.cpp does not support image input. [source]
- GPT-OSS EAGLE-3 gains are near zero or negative on the DGX Spark tests because of the MoE architecture. [source]
- A quantized target does not by itself explain low EAGLE-3 acceptance (see Disagreements). [source]
- DSpark block drafting must draw a full n_max block for every sequence, because the Markov head views the batch as a uniform grid of sequences by block and a per-sequence clamp corrupted the strided views. [source]
- DeepSeek has not released DSpark weights for DeepSeek-V4, so llama.cpp can only run the Qwen3 and Gemma4 drafts. [source]
- PR author: EAGLE-3 on Qwen3-8B BF16 reached 1.62x to 2.17x with 57-70% acceptance on an RTX A6000. A community report on an RTX 5060 8 GB with Qwen3-8B Q4_K_M and the AngelSlim head measured 14.8-58.3% acceptance (code 58.3, JSON 41.7, math 37.9, English prose 19.0, French prose 14.8) and found a plain 0.6B draft model beat it on all five content types. The reporter ruled out target quantization (Q8_0 target did not recover), conversion error and reduced-vocabulary coverage, and left the cause open. Different hardware, quantization, builds and prompts; neither side is retracted. [source]
- No cached page reports EAGLE-3 or DSpark in llama.cpp on Apple silicon. Whether the one-token-per-step EAGLE-3 draft pays off on Metal, where decode is bandwidth-bound and verify cost rises with width, is unmeasured. [source]
- Why the community EAGLE-3 acceptance is far below the PR's numbers is unresolved. [source]
- llama.cpp's speculative docs describe the EAGLE-3 draft as a one-layer transformer trained for one target that shares the target's tokenizer and may use a reduced draft vocabulary with its own `lm_head`, mapped back through a `d2t` table. [source]
- An EAGLE-3 checkpoint is converted with `convert_hf_to_gguf.py ... --target-model-dir <target>` so it inherits the target's tokenizer and the layer indices to read, and both the SpecForge `LlamaForCausalLMEagle3` and the vLLM/AngelSlim `Eagle3LlamaForCausalLM` formats are accepted. [source]
- The run command is `llama-server -m <target>.gguf -md <eagle3>.gguf --spec-type draft-eagle3`. [source]
- The docs list these supported EAGLE-3 drafts: yuhuili LLaMA3.1-8B and LLaMA3.3-70B, RedHatAI Gemma-4 31B and 26B-A4B and gpt-oss-20b, Tengyunw Qwen3 8B and 30B MoE, AngelSlim Qwen3 1.7B/4B/8B/14B/32B, lmsys gpt-oss-120b, and nvidia gpt-oss-120b long-context. [source]
- The DFlash draft in llama.cpp uses several transformer layers and emits a whole block per draft step, in contrast to EAGLE-3's single autoregressive layer. [source]
- DSpark extends DFlash with a semi-autoregressive Markov head: each block position's logits get a low-rank term keyed on the previous token, chained in-graph, which keeps drafting at one decode per block. [source]
- A DSpark draft is a DeepSpec checkpoint such as `deepseek-ai/dspark_qwen3_4b_block7`, converted with `--target-model-dir` and run with `--spec-type draft-dspark --spec-draft-n-max 7 -fa on --jinja`; `--spec-draft-n-max` is clamped to the draft's trained block size. [source]
- `--spec-draft-conf-min P` truncates a DSpark block at the first position whose confidence-head predicted acceptance falls below P, and the default 0 disables it. [source]
- DSpark drafts exported in the vLLM `speculators` format, such as `RedHatAI/gemma-4-31B-it-speculator.dspark`, convert the same way. [source]
- PR 18039 merged into master on 12 Jun 2026 as commit 88a3927 and credits an NVIDIA and GGML collaboration. [source]
- PR 18039 describes the EAGLE-3 draft as: extract features from the target at chosen layers, compress them with a feature-fusion layer, draft with a single-layer decoder, and map draft vocabulary to target vocabulary with a `d2t` tensor. [source]
- PR 18039 reports EAGLE-3 (draft size 8, RTX A6000 48 GB) on LLaMA3.1-8B BF16 at 3.28x, 2.85x and 2.55x over baseline on three prompts, with 77-81% acceptance. [source]
- PR 18039 reports LLaMA3.1-8B Q4_K_M at 1.62x to 2.26x, LLaMA3.3-70B Q4_K_M at 1.85x to 2.41x (15.6 tok/s baseline), Qwen3-8B BF16 at 1.62x to 2.17x, Qwen3-14B BF16 at 1.26x to 1.46x, and Qwen3-32B Q4_K_M at 1.16x to 1.30x. [source]
- PR 18039 reports Qwen3-30B-A3B BF16 on a DGX Spark at 1.25x to 1.39x and GPT-OSS-20B BF16 on a DGX Spark at 1.06x and 0.95x. [source]
- PR 18039 says to apply each model's own chat template when building prompts, because it raises acceptance. [source]
- The PR 18039 author states EAGLE-3 does not support image input in llama.cpp. [source]
- A community report on PR 18039 measured EAGLE-3 acceptance of 58.3% (code), 41.7% (JSON), 37.9% (math), 19.0% (English prose) and 14.8% (French prose) for Qwen3-8B Q4_K_M with the AngelSlim head at n_max 3, and a Qwen3-0.6B Q8_0 `draft-simple` model beat it on all five. [source]
- PR 25173 builds DSpark as a `dspark` model class that inherits the DFlash class and a `draft-dspark` implementation that inherits the DFlash implementation, overriding only `draft()`. [source]
- In PR 25173 the Markov bias for a block position is `markov_w2 · markov_w1[prev]`, computed on-device, and the block is anchor-first so position 0 already predicts the first draft token. [source]
- PR 25173 keeps the draft in bf16 and quantizes only the target, because the draft is tiny and its acceptance is unaffected by target quantization. [source]
- PR 25173's per-domain RTX 4090 runs give DSpark about 1.16x geometric-mean speedup over DFlash, with the largest gains on reasoning (GSM8K, +25 points of acceptance, 1.30x) and chat (MT-Bench, 1.29x), and near parity on HumanEval where both accept about 80%. [source]
- PR 25173's review history folds DSpark into the DFlash architecture, detects Markov tensors by presence like the EAGLE-3 `d2t` table, adds `--spec-draft-conf-min`, and drafts the full n_max block for every sequence. [source]
- A maintainer reply on PR 25173 says DeepSeek has not open-sourced DSpark weights for DeepSeek-V4, only the Qwen3 and Gemma4 drafts. [source]
- When `-hfd` names a draft repo that ships mtp, dflash or eagle3 sidecars and no `--spec-type` is given, llama.cpp picks the first available sidecar by the priority mtp, then dflash, then eagle3 (PR 25989, listed in PR 25173's commit history). [source]
- In llama.cpp's argument handling the sidecar order is mtp, then dspark, then dflash, then eagle3, and DSpark outranks DFlash because its sidecar carries the extra Markov head. [source]
- An explicit draft file selection (`-md` with `-hfd`) disables sidecar resolution of the draft repo, and when no `--spec-type` is given llama.cpp infers the type from the draft GGUF metadata, reading only the first split, so a sharded draft needs an explicit `--spec-type`. [source]
- mlx-vlm exposes `--draft-kind dflash|eagle3|mtp`, auto-detects the Red Hat Speculators Gemma-4 31B EAGLE-3 checkpoint, and converts `Inferact/MiniMax-M3-EAGLE3` for MiniMax M3, which publishes no mtp or nextn tensors. [source]
- Rapid-MLX's 0.5.8 roadmap lists "EAGLE-3 feature-level draft on Metal" at an expected 3-6.5x decode and marks it not started. [source]
- oMLX issue 854's logs show the discovery step registering `RedHatAI/gemma-4-31B-it-speculator.eagle3` as an ordinary `llm` model with the batched engine, and even setting it as the default model, rather than as a drafter. [source]
- A Mac-focused write-up quotes DSpark's paper as improving macro-average accepted length by about 31% over EAGLE-3 and about 16% over DFlash on a Qwen3-4B target. [source]
- No cached page reports EAGLE-3 or DSpark acceptance or speed under llama.cpp's Metal backend. [source]
Children
- No children recorded.