Gemma-4 assistant drafters on MLX and llama.cpp
Parent: Mac local LLMs: Speculative decoding and MTP · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
20 May 2026: llama.cpp PR 23398 opens (31B and 26B-A4B only). 7 Jun 2026: merged. Later PR 24282 adds the E2B and E4B assistants.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- 20 May 2026: llama.cpp PR 23398 opens (31B and 26B-A4B only). 7 Jun 2026: merged. Later PR 24282 adds the E2B and E4B assistants. [source]
- 12 Jun 2026: LM Studio 0.4.17-beta on macOS lists the drafter but does not offer it (issue 2044, still open in the cached page). [source]
- Sep 2026: mlx-dspark runs Gemma-4 12B with DSpark and DFlash drafters, a different drafter family from Google's `-assistant`. [source]
- Quantizing the target KV cache (`-ctk q8_0 -ctv q8_0`) gave 0% acceptance until the PR added the missing Hadamard rotation for Q. [source]
- Arch-name spelling differs across tools: upstream `gemma4-assistant` (hyphen) against `gemma4_assistant` (underscore); downstream loaders failed on one spelling. [source]
- A later llama.cpp regression made the drafter fail to load with "invalid vector subscript" (worked on b9553, broke on b9702 and b9717). [source]
- mlx-dspark `--wired-limit` corrupted verify logits on the Gemma-4 mlx-vlm route, so wrong tokens could be committed. [source]
- mlx-vlm 0.6.4 with transformers 5.12 or newer breaks loading Gemma-4 for mlx-dspark. [source]
- MoE target: the PR author saw no speedup for 26B-A4B on a DGX Spark, while a M1 Max user (already in existing dossier) measured +24% for the same model. Both stand; the setups differ (CUDA unified memory against Metal, quant, n_max). [source]
- Whether LM Studio's dropdown filters on vocabulary match is a reporter's guess; LM Studio is closed source and nobody has confirmed it. [source]
- No Mac measurement of the E2B/E4B llama.cpp assistants was found. [source]
- llama.cpp PR 23398 by am17an adds Gemma 4 MTP support using `--spec-type draft-mtp` with a separate drafter GGUF, opened 20 May 2026. [source]
- The PR initially supported Gemma-4 31B and 26B-A4B but not the E2B and E4B variants. [source]
- Support for the E2B and E4B assistants arrived in a later PR 24282, cited by a downstream bump to llama.cpp b9568. [source]
- On a DGX Spark the 31B dense target decoded 5.9-6.2 tok/s without MTP and 11.4-19.3 tok/s at `--spec-draft-n-max 4`, aggregate acceptance 58.8%, total wall time 290.01 s to 120.65 s over 1,728 tokens. [source]
- The PR author reported more than 2x average speedup on the dense model, no speedup for the MoE model on his system, and AIME-26 of about 87% matching Google's published figure. [source]
- The PR's AI-usage disclosure says AI wrote mainly the KV-cache sharing code and the test against the transformers implementation. [source]
- With the target KV cache quantized to q8_0 the drafter initialized but had 0% acceptance; the author fixed it by adding the missing Hadamard rotation for Q, after which quantized KV works. [source]
- Multi-GPU Gemma-4 MTP works but needs `--spec-draft-device` together with `-sm layer`. [source]
- Downstream projects hit two spellings of the architecture key, `gemma4-assistant` and `gemma4_assistant`, and unified them as aliases of one llama.cpp architecture; the Unsloth drafters use the hyphen spelling. [source]
- A llama.cpp regression made `gemma4-assistant` fail to load with "invalid vector subscript" on builds b9702 and b9717 after working on b9553. [source]
- llama.cpp merged the Gemma-4 drafter on 7 Jun 2026 and shipped it in LM Studio's llama.cpp runtime 2.22.0. [source]
- On LM Studio 0.4.17-beta (macOS 26.5.1, M5 Max 128 GB) the drafter `mtp-gemma-4-31B-it.gguf` is indexed by `lms ls` as architecture `gemma4-assistant` (279.95 MB) but never appears in the Draft Model dropdown, in either folder layout tried; the same pair works under llama-server. [source]
- The LM Studio reporter suspects the dropdown's vocabulary-match filter excludes drafters that carry no standalone vocabulary; a Windows 0.4.16 user could select a Gemma-4 drafter in issue 2036 but it crashed at inference. [source]
- llama.cpp's speculative docs list RedHatAI Gemma-4 31B and 26B-A4B EAGLE-3 speculators as supported `draft-eagle3` drafters, and a RedHatAI Gemma-4 31B DSpark speculator in the DSpark conversion path. [source]
- llama.cpp `draft-dspark` supports only Qwen3-backbone drafts as of 27 Aug 2026; Gemma4 backbone support is planned. [source]
- mlx-dspark's Gemma-4 12B (8-bit) row runs cap 7, accept length 4.98, 18.4 to 59.8 tok/s (3.25x) on M4 Pro; the 12B drafters are `deepseek-ai/dspark_gemma4_12b_block7` and `z-lab/gemma4-12B-it-DFlash`. [source]
- mlx-dspark measured `makora-ai/gemma4-26b-a4b-dspark` on `gemma-4-26b-a4b-it-8bit` at 1.27x (1.38x code, 1.37x math, 1.06x chat; 46.9 to 59.5 tok/s) with `--max-draft 2 --no-lookup-drafts`, and left it unregistered for auto-resolution. [source]
- mlx-dspark does not yet run the DFlash-plus-Markov community hybrid `Hikari07jp/DSpark-Gemma-4-31B-draft`. [source]
- mlx-dspark's `--wired-limit` corrupted verify logits on the Gemma-4 mlx-vlm route; mlx-lm targets did not reproduce it. [source]
- Gemma-4 12B did not converge on the pi agent's tool protocol in mlx-dspark's agent test. [source]
- The Gemma-4 model card gives sliding windows of 512 tokens (E2B, E4B) and 1024 tokens (12B, 26B A4B, 31B), with the final layer always global. [source]
- MTPLX release notes list Gemma 4 warm-turn support (PR 283), prompt scoring on Gemma 4 and a fix so Gemma 4 streamed replies finish (issue 517). [source]
Children
- No children recorded.