Mac local LLMs: Speculative decoding and MTP

Parent: Running LLM models locally on a Mac · Published reference · snapshot 2026-10-05

↓ Facts as markdownall context files

llama.cpp (lowest friction): `llama-server -m Qwen3.6-35B-A3B-Q8_0.gguf --spec-type draft-mtp --spec-draft-n-max 3 -ngl 999 -fa on -c 65536 --parallel 1 --jinja`; the MTP GGUF runs normally without the flag.

These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.

Which engine, what to run

Expected speedups (all Mac, from reports)

When it loses

Draft depth

Failure modes and fixes

Drafter families (llama.cpp)

MLX verify kernels (why MLX gains vary)

Structured output and grammars (llama.cpp)

Batched serving (not single-user Mac)

Corrections to earlier claims

Open questions

Corrections and disagreements

Concepts in this cluster

Children

← the whole tree · 3D view· how to read this page