N-gram and prompt-lookup self-speculation for coding agents
Parent: Mac local LLMs: Speculative decoding and MTP · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Jan 2026: ggerganov opens PR 19164 adding `ngram-mod` (flags then named `--spec-ngram-size-n`, `--draft-min`, `--draft-max`).
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Jan 2026: ggerganov opens PR 19164 adding `ngram-mod` (flags then named `--spec-ngram-size-n`, `--draft-min`, `--draft-max`). [source]
- By 27 Aug 2026 the docs use `--spec-ngram-mod-n-match/-n-min/-n-max` and a `--spec-default` that enables `ngram-mod`. [source]
- mlx-dspark 0.5.0 introduced match-scaled lookup drafts, on by default. [source]
- Small n is not recommended for `ngram-mod`; MoE targets need long drafts. [source]
- Insertion-heavy edits do not benefit; mlx-dspark's acceptance gate parks the scaling there. [source]
- Hybrid and MoE targets often lose with lookup drafts; mlx-dspark ships it off for several pairs. [source]
- Draft-model acceptance and n-gram acceptance are different statistics; llama.cpp's synthetic-acceptance flags produce invalid output by design. [source]
- llama.cpp lists "summarization" and "reasoning models repeating their thinking" as applications; mlx-dspark finds chat output unchanged and hybrid lookup a wash or loss on MoE/hybrid targets. Both are single-source claims on different engines. [source]
- No Mac benchmark of llama.cpp `ngram-mod` or `--spec-default` was found. [source]
- Whether `ngram-mod`'s shared cross-slot pool helps or hurts multi-agent workloads on one Mac is untested. [source]
- llama.cpp `--spec-type` accepts a comma-separated list, so a draft model can be mixed with a draftless type, and the draftless type takes precedence. [source]
- llama.cpp `--spec-default` enables the `ngram-mod` implementation. [source]
- `ngram-simple` takes the last earlier match of the current n-gram and drafts the m tokens that followed it, with defaults n=12, m=48 and one minimum hit. [source]
- `ngram-map-k` drafts only when the key n-gram was followed by the same m-gram at least `--spec-ngram-map-k-min-hits` times (default 1), and stores the accepted-token count per n-gram. [source]
- `ngram-map-k4v` is experimental; it tracks up to four continuations per key and drafts one only when it is significantly more frequent than the others. [source]
- `ngram-mod` hashes the last n tokens with an LCG, stores the next token per hash, takes about 16 MB, uses constant memory and complexity, drafts variable lengths, and shares one hash pool across all server slots. [source]
- `ngram-mod` defaults are `--spec-ngram-mod-n-match 24`, `--spec-ngram-mod-n-min 48` and `--spec-ngram-mod-n-max 64`; the docs advise against small n and say MoEs need long drafts while dense models can lower min and max. [source]
- llama.cpp names code-block iteration (llama.vim), reasoning models repeating their thinking and summarization as `ngram-mod` applications. [source]
- `ngram-cache` keeps n-gram statistics and can load external statistics from files for better accuracy. [source]
- llama.cpp's `--spec-synth-rates` and `--spec-synth-len` replace verification with synthetic accept decisions for benchmarking, so their output is not valid model output. [source]
- llama.cpp's default `--spec-draft-n-max` is 3. [source]
- ggerganov's `ngram-mod` PR 19164 (28 Jan 2026) states constant memory, a hash pool shared across slots and the three applications, and drew 21 thumbs-up. [source]
- mlx-dspark `--mode lookup` runs n-gram speculation with no drafter on any target, and `--mode auto` falls back to it when no drafter or DFlash pair exists. [source]
- mlx-dspark's match-scaled long drafts (`--lookup-long-draft`, default 32) give a copy run that matches at least 8 tokens deep drafts up to about twice the matched length. [source]
- mlx-dspark measured verify width 16-32 as a plateau on M-series costing about 2.5 times one decode step, so verbatim spans commit about 20-30 tokens per forward. [source]
- On M4 Pro (8-bit, outputs bit-identical) match-scaled lookup lifted Gemma-12B file re-emission from 3.0x to 4.5x (75 tok/s) and Ornith-9B rename-edit from 2.8x to 3.6x (93 tok/s); chat was unchanged. [source]
- mlx-dspark's acceptance gate parks the long-draft scaling on insertion-heavy edits, where it measured neutral. [source]
- mlx-dspark's hybrid lookup is on by default for dspark mode, switched off with `--no-lookup-drafts` or per swap through `/admin/load`, and ships off for several pairs where it measured a loss. [source]
- MTPLX reports a 27B rewriting a file it just wrote at 87.6 tok/s against 65.2 tok/s on a fresh coding task, and 2.10.0 notes rewrites running 19% faster than fresh writes (the mechanism is not stated). [source]
Children
- No children recorded.