llama.cpp speculative checkpointing for hybrid models (PRs 19493 and 22227)
Parent: Mac local LLMs: Speculative decoding and MTP · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
The author states the cost plainly: a checkpoint restore is slower than `seq_rm`, because after a partial accept the server returns to the checkpoint and executes a shorter batch.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- The author states the cost plainly: a checkpoint restore is slower than `seq_rm`, because after a partial accept the server returns to the checkpoint and executes a shorter batch. [source]
- At load time the server logs `common_speculative_is_compat: the target context does not support partial sequence removal` and then `speculative decoding not supported by this context without checkpoints`; with checkpoints enabled it proceeds. [source]
- Reviewer ggerganov asked on 11 Feb 2026 that the logic live in `common/speculative` so that `server` stays free of speculation-specific code; the author then moved the `spec_ckpt_` variables and logic there and added `common_speculative_session` and `common_speculative_callback` (24 Feb). [source]
- The merged commit list shows the later reshaping: a `--spec-use-checkpoints` flag was added, then removed ("avoid --spec-use-checkpoints argument"), a `leave_draft_state` rename, a callback removed from the session, `common_speculative_accept_response` removed, and mtmd (multimodal) speculative decoding enabled. [source]
- PR 22227 adds the same checkpoint logic to `speculative-simple` and avoids cloning the sampler in `llama-server` when it is not needed. [source]
- A draft-model checkpoint is a separate object: the author added optional checkpoints to the draft-model implementation of `common/speculative.cpp` on 8 Mar 2026 because a Qwen3.5 draft is itself recurrent. [source]
- 10 Feb 2026: PR 19493 opens as "follow-up to 19270" with `--spec-type ngram-map-k --draft-max 48 --spec-ckpt-num-tries 2 --ctx-checkpoints 16`, tested on Qwen3-Coder-Next. [source]
- 2 Mar 2026: sample arguments become `--spec-type ngram-mod --draft-max 48 --spec-use-checkpoints on --ctx-checkpoints 12`, tested on Qwen3.5-35B-A3B Q4_K_XL. [source]
- 6 to 8 Mar 2026: a Qwen3.5 0.8B draft model with Qwen3.5-27B as target is added; before the fix "the main model gets confused by the drafts". [source]
- 10 to 11 Mar 2026: a GGML_ASSERT when the server switches slots in a chat is traced to a dangling `server_slot &` in `server_speculative_callback` and fixed by passing the slot id. [source]
- 21 to 22 Apr 2026: PR 22227 merges; the same page lists ik_llama.cpp's "Speculative checkpoints for recurrent models" (PR 1669) as merged, mentioned on 22 Apr. [source]
- 24 Apr 2026: PR 22105 (DFlash) rebases onto speculative checkpointing, so PR 19493 had merged by then. [source]
- 29 Apr 2026: PR 22521 "spec : fix draft model checkpoints" changes when draft-model checkpoints are created and restored; the old logic discarded the checkpoint at every new completion request, which caused long draft-model recompute in large agentic sessions. [source]
- May 2026: a downstream cherry-pick lists the four PRs needed for hybrid speculation: 19493, 22114 (refactor of the "use checkpoint" logic), 22168 (reset `i_last` on a low-acceptance streak) and 22223 (`--spec-default`), and says it was smoke-tested on an M5 Max with turbo4 KV with zero regression. [source]
- Gains depend on repetition. On the quicksort prompt with `ngram-map-k` on Qwen3-Coder-Next, decode ran 96.30 tok/s with no draft hits on the first request and 161.14, 140.60, 258.38 and 267.02 tok/s on later requests at 72%, 39%, 71% and 62% draft acceptance. [source]
- A tester reported "well above 30% speedup in token generation for code and around 10% down for normal text where the drafts don't match well" (13 Mar 2026). [source]
- A recurrent draft model produces many more invalid drafts than ngram methods, so the author expected checkpoints to be less efficient there than with a non-recurrent draft such as Qwen 2.5. [source]
- A parallel test (three quicksort prompts at once) passed on master and failed on the feature branch at the text level (a diff of `print(sorted_data)` against `print("Original:", ...)`), so output equality under concurrency was still unsettled on 14 Mar 2026; the thread does not show a resolution. [source]
- A tester hit "the tokens of sequence 0 in the input batch have inconsistent sequence positions" with a Qwen3.5 draft on 6 Mar 2026 and suspected an off-by-one at the end of a batch in `common/speculative.cpp`; the author answered that the PR could not yet be used with a Qwen3.5 draft model. [source]
- The PR 22105 text calls checkpointing a fallback and names SGLang's target-side deferred commit as the more fundamental fix; the PR 19493 author frames checkpointing as good enough where drafts are mostly accepted and measures its loss only on poorly matched text. Both positions hold; no measurement compares checkpoints to a deferred-commit kernel inside llama.cpp. [source]
- The exact merge date of PR 19493 is not on the cached page (comments are hidden); only "merged before 24 Apr 2026" is known. [source]
- Whether the text-level mismatch under three concurrent requests on 14 Mar 2026 was fixed before merge. [source]
- A Metal measurement of checkpoint cost per rejected draft; every log in both threads is CUDA, ROCm or Vulkan, apart from the M5 Max smoke test without numbers. [source]
- PR 19493 was opened on 10 Feb 2026 by srogmann as a follow-up to PR 19270 and issue 19267 to support speculative decoding with recurrent modules using checkpoints. [source]
- The PR description says checkpoints are not as fast as `llama_memory_seq_rm` because a partly accepted draft means going back to the checkpoint and executing a shorter batch. [source]
- The PR included a small fix to the `ngram-map-k` implementation and listed open tasks: accept feedback to shorten drafts, separate "accepted" from "could be accepted" statistics, and factor out checkpoint creation. [source]
- On 11 Feb 2026 ggerganov called the PR "good as a prototype" and asked for the logic to move into `common/speculative`. [source]
- The first server log used `--spec-type ngram-map-k --draft-max 48 --spec-ckpt-num-tries 2 --ctx-checkpoints 16` on Qwen3-Coder-Next. [source]
- The 2 Mar 2026 test on Qwen3.5-35B-A3B-UD-Q4_K_XL used `--spec-type ngram-mod --draft-max 48 --spec-use-checkpoints on --ctx-checkpoints 12`, and the ngram_mod table was sized 4,194,304 entries. [source]
- On a DGX Spark with Qwen3.5-35B-A3B Q4_K_XL (20.7 GiB) a tester saw about 62 tok/s with `--spec-use-checkpoints on` and could not reproduce the slot-switch GGML_ASSERT. [source]
- Qwen3.5-27B Q5_K_XL ran 19.25 tok/s without a draft on the author's machine, and 29.30 tok/s with a Qwen3.5-0.8B Q8_0 draft before the draft-checkpoint fix (691 of 1,552 drafted tokens accepted). [source]
- After the fix the same pair ran 32.76, 41.41, 45.20, 53.57 and 69.59 tok/s over five successive quicksort requests, each at 100% accepted-over-generated draft rate, with `--draft-max 24 --draft-min 8 --draft-p-min 0.9 --ctx-checkpoints 4`. [source]
- The squashed commit list of PR 19493 includes "enable mtmd speculative decoding", "restore sampler in spec checkpoint and clear mem" and "avoid --spec-use-checkpoints argument". [source]
- PR 22227 states "cont 19493", adds speculative checkpoints to `speculative-simple`, and avoids cloning the sampler in `llama-server` when unnecessary; it was approved by ServeurpersoCom and merged on 22 Apr 2026 as commit bcb5eeb with a sample command using Qwen3.5-35B-A3B Q8_0 and a 0.8B base draft at `--draft 32 --temp 0 --top-k 1`. [source]
- The PR 22227 page lists ik_llama.cpp PR 1669 "Speculative checkpoints for recurrent models" as merged, mentioned on 22 Apr 2026. [source]
- PR 22521 (ggerganov, 29 Apr 2026) fixes draft-model checkpoints that were discarded on every new completion request. [source]
- A downstream fork's cherry-pick message names PRs 19493, 22114, 22168 and 22223 as the four upstream PRs that enable speculation on Qwen3.6-35B-A3B hybrid MoE+SSM. [source]
- A tester measured the feature at 30% or more faster on code and about 10% slower on ordinary text where drafts match badly. [source]
Children
- No children recorded.