llama.cpp --spec-default and low-acceptance streak reset (PRs 22223 and 22168)
Parent: Mac local LLMs: Speculative decoding and MTP · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
`--spec-default` pushes `COMMON_SPECULATIVE_TYPE_NGRAM_MOD` and sets `ngram_mod.n_match = 24`, `n_min = 48`, `n_max = 64`; a commented-out `ngram-map-k4v` block (size_n 8, size_m 24, min_hits 2) shows the author has not yet chosen to add a second type.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- `--spec-default` pushes `COMMON_SPECULATIVE_TYPE_NGRAM_MOD` and sets `ngram_mod.n_match = 24`, `n_min = 48`, `n_max = 64`; a commented-out `ngram-map-k4v` block (size_n 8, size_m 24, min_hits 2) shows the author has not yet chosen to add a second type. [source]
- The flag is registered for the server and CLI examples only. [source]
- A source comment on the flag reads "not sure if this is a good config - explore more settings". [source]
- Each sequence keeps `i_last`, the last prompt position already inserted into the shared n-gram table, plus `n_draft_last` and `n_low`. [source]
- During drafting, new n-grams are added in chunks only when `i_last + 32 < cur_len`, so the table lags the context by up to 32 tokens. [source]
- On each accept call, a round counts as low when accepted tokens divided by drafted tokens is below 0.25; five low rounds in a row trigger `mod.reset()`, `n_low = 0` and `i_last = 0`, and any round at or above 0.25 sets `n_low = 0`. [source]
- Setting `i_last = 0` is the PR 22168 change: after the reset, the next draft call re-adds every n-gram of the current context. Before the PR the table was emptied but `i_last` kept its old value, so only text generated after the reset was indexed. [source]
- The `accept` hook returns early when `is_other` is true, so rounds credited to another speculator type do not change `n_low`. [source]
- `begin()` also resets the table when its occupancy after indexing a prompt exceeds 0.25 of 4M entries (`4*1024*1024`), and warns when `n_match` is too small. [source]
- 20 Apr 2026: treo opens PR 22168 (one line, "Reset i_last when low acceptance streak occurs", AI-assisted to understand the ngram-mod flow). [source]
- 21 Apr 2026: ggerganov asks for the speculative parameters used; treo supplies `--draft-max 128 --spec-ngram-size-n 48 --draft-min 2 --spec-type ngram-mod` with `--reasoning off`. ggerganov approves and merges the same day (commit 72d693e). [source]
- 21 Apr 2026: ggerganov opens and merges PR 22223 (`arg : add --spec-default`) within hours; pwilkin and danbev approve; 46 of 49 checks passed at merge (commit 84652b8). [source]
- 27 Apr 2026: the reset is ported to ik_llama.cpp as PR 1701 and merged. [source]
- By October 2026 the source has `n_match 24, n_min 48, n_max 64`, whereas the PR 22168 benchmarks used n 48, min 2, max 128; the flag names changed from `--spec-ngram-size-n` to `--spec-ngram-mod-n-match`. [source]
- The gain is small and model dependent: treo says "the effect isn't huge but at 3% to 9% it is still measurable". [source]
- The cost is a moment of table repopulation on every reset. [source]
- Reset repopulation scans the whole context, so its cost grows with context length; no measurement at 100k tokens was published. [source]
- A low streak often precedes a long run of full acceptance (code reproduction), which is why a table that was never re-indexed lost the benefit. [source]
- Smoke tests on an M5 Max with turbo4 KV showed no regression, but a CUDA user on the same fork with a draft model (not ngram) reported about 0.4 tok/s, roughly 130 times slower, at 100% draft acceptance; the issue was closed "not planned" as stale. [source]
- The PR 22168 text says the old behaviour "skips the current context". The code comment on `i_last` says it is "the last position in the prompt that was added to the ngram container". Both agree once read together: skipping applied only to text before the reset point. [source]
- `--spec-default` enables `ngram-mod` with n_match 24, n_min 48, n_max 64. [source]
- The flag's author marks the configuration as unproven in a source comment. [source]
- A round is low when accepted/drafted is below 0.25, and five consecutive low rounds reset the `ngram-mod` table. [source]
- The reset also sets `i_last = 0` so the whole current context is re-indexed. [source]
- New n-grams are indexed only when the context is more than 32 tokens past `i_last`. [source]
- On Gemma 4 26B-A4B, PR 22168's benchmark gave 97.57 tok/s baseline, 141.08 without the reset and 155.10 with it. [source]
- On Qwen 3.6 35B-A3B, the same benchmark gave 112.54 baseline, 148.56 without and 153.65 with the reset. [source]
- The benchmark used `vllm bench serve` on the vdaita/edit_5k_char dataset at concurrency 1 with 96 prompts and `--reasoning off` for repeatability. [source]
- treo reports 3x to 7x over baseline for agentic coding with `--spec-ngram-size-n 32 --draft-max 64 --draft-min 32` at contexts above 100k tokens. [source]
- The ik_llama.cpp port measured 95.11 tok/s baseline, 130.4 without and 142.82 with the reset on 2x RTX 4000 Ada with Qwen 3.6 35B-A3B, about 9%. [source]
- PR 22223 has no description beyond "Add `--spec-default` flag for enabling default configuration for speculative decoding". [source]
- `ngram-mod` resets its table at prompt start when occupancy exceeds 0.25. [source]
Children
- No children recorded.