<!-- llms-explorer concept facts · https://llms-explorer.com/tree/mtplx-mtp-self-drafting-mac-server/ · pack 2026-10-05 · ~1864 tokens -->

# MTPLX MTP self-drafting Mac server

> 2.0.0 (6 Jul 2026): prefix cache works with speculation on hybrid models.

Parent: [Mac local LLMs: oMLX, Rapid-MLX and related internals](https://llms-explorer.com/tree/mac-local-llms-omlx-and-rapid-mlx-internals/) · 1 facets · 37 facts · page: https://llms-explorer.com/tree/mtplx-mtp-self-drafting-mac-server/

## Facts

- 2.0.0 (6 Jul 2026): prefix cache works with speculation on hybrid models. — source: `asserted`
- 2.9.2: agent transcripts pass through unmodified by default. — source: `asserted`
- 2.11 (4 Sep 2026): same-origin local server, no dead time between tool turns, flash-decoding verify. — source: `asserted`
- 2.12.1 (2 Oct 2026): memory guard rewritten. 2.12.2 (3 Oct 2026): prompts priced per model for 8, 16 and 32 GB Macs. — source: `asserted`
- A request that would exceed the memory limit is refused with a structured 507 before prefill instead of swapping the Mac toward a panic. — source: `asserted`
- Codex sends a hosted `web_search` tool by default; MTPLX returns a precise 400 unless it is removed. — source: `asserted`
- Non-zero presence or frequency penalties change the sampled distribution, so exactness holds only at the default of 0 (inferred from the README's "exact no-op" wording). — source: `asserted`
- Cache keys for vision turns come from the image bytes so one image never reads another's cache. — source: `asserted`
- MTPLX says output is exact at any temperature; Gemma and Qwen tests elsewhere (existing dossier) show multi-token verify differs from single-token at near-ties. MTPLX reports max_diff = 0.0 and token-id equality per release; independent confirmation is absent. — source: `asserted`
- Whether concurrent clients batch speculative decoding or serialize (docs/concurrency.md was not fetched). — source: `asserted`
- Independent reproduction of the exactness claim. — source: `asserted`
- MTPLX's verify step evaluates all K drafted positions in one batched target forward using GraphBank-compiled verify shapes. — [source](https://mtplx.com/how-it-works/)
- MTPLX decides acceptance with the probability ratio in an fp32 path because bf16 underflows, and rejected drafts never enter committed history. — [source](https://mtplx.com/how-it-works/)
- MTPLX's committed-history KV contract was verified against a reference at cosine above 0.9998 through depth 5, and its quoted exactness check is max_diff = 0.0 against single-token autoregressive decoding. — [source](https://mtplx.com/how-it-works/)
- MTPLX's acceptance by draft position on Qwen 3.6 27B at depth 4 (192-token coding bench, temperature 0.6, 28 Apr 2026) was 97.62%, 95.24%, 88.10% and 75.61%. — [source](https://mtplx.com/how-it-works/)
- MTPLX builds on an MLX source fork (mlx-mtplx-0.31.2-qmm) with a retuned small-M qmv (BN16, 4 simdgroups, unroll_count 4) for verify shapes M=3..6, a fused GDN verify kernel built from the conv tape, and a draft-only 4- or 3-bit LM head. — [source](https://mtplx.com/how-it-works/)
- MTPLX's SessionBank keeps a warm-prefix exact state, and the quoted check is logits_max_abs_diff = 0.0 across turns. — [source](https://mtplx.com/how-it-works/)
- The MTPLX server also exposes `/v1/responses` (stateless, text-only, for Codex), `/v1/completions`, `/health` and `/metrics`, plus optional `/v1/embeddings` and `/v1/rerank`. — [source](https://github.com/youssofal/MTPLX)
- MTPLX answers Codex's default hosted `web_search` tool with a 400 unless the tool is removed, and resolves `reasoning.effort: "xhigh"` per model (Qwen 3.8 keeps it, Step 3.5 clamps to high, Qwen 3.6 has no tier). — [source](https://github.com/youssofal/MTPLX)
- `mtplx start` attaches to the app's already-running model rather than loading a second copy. — [source](https://github.com/youssofal/MTPLX)
- MTPLX serves embedding and reranker models from the same daemon with `--embedding-model` and `--reranker-model`; they bypass the MTP path, load on first request, are capped by `--retrieval-max-resident` (default 2), and checkpoints with bundled Python code get a 403 unless `--retrieval-trust-remote-code` is set. — [source](https://github.com/youssofal/MTPLX)
- MTPLX keeps `/v1/models` chat-only by default and lists retrieval models via `?capability=embedding` or `?capability=rerank`. — [source](https://github.com/youssofal/MTPLX)
- Presence and frequency penalties default to 0, an exact no-op that preserves MTP exactness; the README says to leave them at 0 for coding and agent work. — [source](https://github.com/youssofal/MTPLX)
- A default-on SSD session cache restores sessions across restarts (`--ssd-session-cache off` disables it); on 2.11.3 a 96,760-token conversation restored in 8 ms in the app. — [source](https://github.com/youssofal/MTPLX)
- MTPLX 2.11.3 reports a stop-and-resend reuse of 10,918 of 10,928 prompt tokens, and a save to SSD that yields to a waiting request in under 200 ms. — [source](https://mtplx.com/releases/)
- MTPLX 2.11 reports a tool turn after a file write re-prefills 20 tokens in 0.12 s instead of 3,535 tokens in 3.7 s, and a forced tool round in a 41k session drops from two cold re-prefills to 0.58 s. — [source](https://mtplx.com/faq/)
- Since MTPLX 2.9.2 the agent server no longer compacts tool results, trims file reads or injects steering text unless a rewrite is explicitly enabled. — [source](https://mtplx.com/releases/)
- MTPLX 2.9.2 turns chained greedy drafting on by default at temperature 0 under 12k context, measured +2.5 to +9.8% across 0.5k-8k prompts. — [source](https://mtplx.com/releases/)
- MTPLX 2.11 makes the local server same-origin by default, removes the API key from URLs and logs, and fixes prompts past 32k failing with HTTP 500 on macOS 27 betas. — [source](https://mtplx.com/releases/)
- MTPLX 2.12.1 rewrote the memory guard to price every request at its largest moment, release idle conversations before refusing, and name what holds the memory; the default memory limit became 90 GiB on 128 GB Macs and `--memory-limit max` uses everything outside macOS's reserve. — [source](https://mtplx.com/releases/)
- MTPLX refuses a request that would cross the memory limit with a structured HTTP 507 before prefill, instead of swapping the Mac toward a panic. — [source](https://mtplx.com/releases/)
- MTPLX 2.12.2 (3 Oct 2026) prices each prompt at its measured use (0.19 GiB for a 59-token prompt on the 4B instead of the 27B's 3 GiB chunk) and retries in 1,024 or 512-token chunks before refusing. — [source](https://mtplx.com/releases/)
- MTPLX 2.8 stopped long sessions silently losing prompt-cache reuse near 38k tokens by advancing the committed frontier on every turn, tool-call turns included. — [source](https://mtplx.com/releases/)
- MTPLX `/health` reports degradation such as compiled-verify fallbacks, overridden profile keys and kernel bails, and `mtplx_stats` is always populated. — [source](https://mtplx.com/releases/)
- MTPLX records a per-second flight-recorder log per request and `mtplx trace` turns a coding session into a diagnosis report, local-only at a few MB a day. — [source](https://mtplx.com/releases/)
- Fan-backed MTPLX modes restore automatic fan control through a detached watchdog even after `kill -9` or closing the terminal. — [source](https://github.com/youssofal/MTPLX)
- Using MTPLX in a shipped product requires an in-product "Powered by MTPLX" notice under the NOTICE file, beyond Apache-2.0's license text. — [source](https://github.com/youssofal/MTPLX)
