MTPLX MTP self-drafting Mac server
Parent: Mac local LLMs: oMLX, Rapid-MLX and related internals · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
2.0.0 (6 Jul 2026): prefix cache works with speculation on hybrid models.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- 2.0.0 (6 Jul 2026): prefix cache works with speculation on hybrid models. [source]
- 2.9.2: agent transcripts pass through unmodified by default. [source]
- 2.11 (4 Sep 2026): same-origin local server, no dead time between tool turns, flash-decoding verify. [source]
- 2.12.1 (2 Oct 2026): memory guard rewritten. 2.12.2 (3 Oct 2026): prompts priced per model for 8, 16 and 32 GB Macs. [source]
- A request that would exceed the memory limit is refused with a structured 507 before prefill instead of swapping the Mac toward a panic. [source]
- Codex sends a hosted `web_search` tool by default; MTPLX returns a precise 400 unless it is removed. [source]
- Non-zero presence or frequency penalties change the sampled distribution, so exactness holds only at the default of 0 (inferred from the README's "exact no-op" wording). [source]
- Cache keys for vision turns come from the image bytes so one image never reads another's cache. [source]
- MTPLX says output is exact at any temperature; Gemma and Qwen tests elsewhere (existing dossier) show multi-token verify differs from single-token at near-ties. MTPLX reports max_diff = 0.0 and token-id equality per release; independent confirmation is absent. [source]
- Whether concurrent clients batch speculative decoding or serialize (docs/concurrency.md was not fetched). [source]
- Independent reproduction of the exactness claim. [source]
- MTPLX's verify step evaluates all K drafted positions in one batched target forward using GraphBank-compiled verify shapes. [source]
- MTPLX decides acceptance with the probability ratio in an fp32 path because bf16 underflows, and rejected drafts never enter committed history. [source]
- MTPLX's committed-history KV contract was verified against a reference at cosine above 0.9998 through depth 5, and its quoted exactness check is max_diff = 0.0 against single-token autoregressive decoding. [source]
- MTPLX's acceptance by draft position on Qwen 3.6 27B at depth 4 (192-token coding bench, temperature 0.6, 28 Apr 2026) was 97.62%, 95.24%, 88.10% and 75.61%. [source]
- MTPLX builds on an MLX source fork (mlx-mtplx-0.31.2-qmm) with a retuned small-M qmv (BN16, 4 simdgroups, unroll_count 4) for verify shapes M=3..6, a fused GDN verify kernel built from the conv tape, and a draft-only 4- or 3-bit LM head. [source]
- MTPLX's SessionBank keeps a warm-prefix exact state, and the quoted check is logits_max_abs_diff = 0.0 across turns. [source]
- The MTPLX server also exposes `/v1/responses` (stateless, text-only, for Codex), `/v1/completions`, `/health` and `/metrics`, plus optional `/v1/embeddings` and `/v1/rerank`. [source]
- MTPLX answers Codex's default hosted `web_search` tool with a 400 unless the tool is removed, and resolves `reasoning.effort: "xhigh"` per model (Qwen 3.8 keeps it, Step 3.5 clamps to high, Qwen 3.6 has no tier). [source]
- `mtplx start` attaches to the app's already-running model rather than loading a second copy. [source]
- MTPLX serves embedding and reranker models from the same daemon with `--embedding-model` and `--reranker-model`; they bypass the MTP path, load on first request, are capped by `--retrieval-max-resident` (default 2), and checkpoints with bundled Python code get a 403 unless `--retrieval-trust-remote-code` is set. [source]
- MTPLX keeps `/v1/models` chat-only by default and lists retrieval models via `?capability=embedding` or `?capability=rerank`. [source]
- Presence and frequency penalties default to 0, an exact no-op that preserves MTP exactness; the README says to leave them at 0 for coding and agent work. [source]
- A default-on SSD session cache restores sessions across restarts (`--ssd-session-cache off` disables it); on 2.11.3 a 96,760-token conversation restored in 8 ms in the app. [source]
- MTPLX 2.11.3 reports a stop-and-resend reuse of 10,918 of 10,928 prompt tokens, and a save to SSD that yields to a waiting request in under 200 ms. [source]
- MTPLX 2.11 reports a tool turn after a file write re-prefills 20 tokens in 0.12 s instead of 3,535 tokens in 3.7 s, and a forced tool round in a 41k session drops from two cold re-prefills to 0.58 s. [source]
- Since MTPLX 2.9.2 the agent server no longer compacts tool results, trims file reads or injects steering text unless a rewrite is explicitly enabled. [source]
- MTPLX 2.9.2 turns chained greedy drafting on by default at temperature 0 under 12k context, measured +2.5 to +9.8% across 0.5k-8k prompts. [source]
- MTPLX 2.11 makes the local server same-origin by default, removes the API key from URLs and logs, and fixes prompts past 32k failing with HTTP 500 on macOS 27 betas. [source]
- MTPLX 2.12.1 rewrote the memory guard to price every request at its largest moment, release idle conversations before refusing, and name what holds the memory; the default memory limit became 90 GiB on 128 GB Macs and `--memory-limit max` uses everything outside macOS's reserve. [source]
- MTPLX refuses a request that would cross the memory limit with a structured HTTP 507 before prefill, instead of swapping the Mac toward a panic. [source]
- MTPLX 2.12.2 (3 Oct 2026) prices each prompt at its measured use (0.19 GiB for a 59-token prompt on the 4B instead of the 27B's 3 GiB chunk) and retries in 1,024 or 512-token chunks before refusing. [source]
- MTPLX 2.8 stopped long sessions silently losing prompt-cache reuse near 38k tokens by advancing the committed frontier on every turn, tool-call turns included. [source]
- MTPLX `/health` reports degradation such as compiled-verify fallbacks, overridden profile keys and kernel bails, and `mtplx_stats` is always populated. [source]
- MTPLX records a per-second flight-recorder log per request and `mtplx trace` turns a coding session into a diagnosis report, local-only at a few MB a day. [source]
- Fan-backed MTPLX modes restore automatic fan control through a detached watchdog even after `kill -9` or closing the terminal. [source]
- Using MTPLX in a shipped product requires an in-product "Powered by MTPLX" notice under the NOTICE file, beyond Apache-2.0's license text. [source]
Children
- No children recorded.