<!-- llms-explorer concept facts · https://llms-explorer.com/tree/mtplx-memory-guard-and-request-pricing-on-8-128/ · pack 2026-10-05 · ~3220 tokens -->

# MTPLX memory guard and request pricing on 8-128 GB Macs

> Price: every request is priced at its largest moment: the end of prompt processing (including the copy a cache restore makes), the move to paged memory, and the start of the answer.

Parent: [Mac local LLMs: oMLX, Rapid-MLX and related internals](https://llms-explorer.com/tree/mac-local-llms-omlx-and-rapid-mlx-internals/) · 1 facets · 49 facts · page: https://llms-explorer.com/tree/mtplx-memory-guard-and-request-pricing-on-8-128/

## Facts

- Price: every request is priced at its largest moment: the end of prompt processing (including the copy a cache restore makes), the move to paged memory, and the start of the answer. — source: `asserted`
- Order when memory is short: clear the engine's buffer pool, narrow the prompt chunk, release RAM copies of idle conversations (oldest first, already-on-SSD first), and only then refuse. — source: `asserted`
- Live inputs: MLX's own account (active plus cache) against the Metal limit, plus the process's real footprint from macOS beyond what MLX explains. — source: `asserted`
- A conversation that is generating is never released. — source: `asserted`
- 2.10.2: first "honest memory refusals" (issue 415): a large prefill is projected up front, superseded session-bank entries are cleared, and a request that cannot fit gets a structured 507 instead of a stream that dies mid-flight. — source: `asserted`
- 2.11.2 (6 Sep 2026): a request that would cross the memory limit is refused with a 507 before prefill instead of swapping the Mac to a panic (issues 450 and 447). — source: `asserted`
- 2.12.0: the guard checks only requests with 4,096 or more uncached tokens and forgives host memory up to 22 GiB. — source: `asserted`
- 2.12.1 (2 Oct 2026): guard rewritten; it over-priced small Macs. — source: `asserted`
- 2.12.2 (3 Oct 2026): per-model prompt pricing, chunk fallback and a smaller host-memory charge on Macs under 64 GB. — source: `asserted`
- 2.12.1 refused ordinary prompts on small seats that 2.12.0 served, because it charged every model the 27B's whole-request reserve and counted a healthy engine's own memory as a leak. — source: `asserted`
- With little free memory a prompt is still refused a hair short even in 2.12.2 (0.1-0.2 GiB), because 1 GiB is always kept free for macOS. — source: `asserted`
- A 48 GB Mac running the 27B keeps two copies of a restored conversation, so a warm turn near 100K tokens is refused. — source: `asserted`
- Compaction requests (Pi) rewrite the prompt and cannot reuse the saved conversation, so they re-read the whole session and can still be stopped under desktop memory pressure. — source: `asserted`
- None found between sources. All claims are the MTPLX maintainer's; no independent user reproduces the guard's prices, and some memory figures (a 205K-token compaction taking 410 s, a re-read estimate of 3 min 55 s against 7 min 43 s actual) show the author's own estimates running optimistic. — source: `asserted`
- No real 8, 16 or 32 GB Mac with a busy desktop has run 2.12.1 or 2.12.2 with the new prices; the arithmetic is emulated on an M5 Max. — source: `asserted`
- Reading a saved conversation back from SSD before the prompt, and the eager verifier's cache growth during an answer, are not checked by the guard yet. — source: `asserted`
- Whether Ultra-class Macs and 192 GB-plus seats behave as the 75% rule predicts is unmeasured. — source: `asserted`
- MTPLX 2.12.0 checked memory only for requests with 4,096 or more uncached tokens and charged them once plus a flat 3 GiB, so a short turn on a long conversation was never checked; 2.12.1 prices every request at the end of prompt processing (including a restore's copy), the move to paged memory and the start of the answer. — [source](https://mtplx.com/releases/2.12.1/)
- In 2.12.1 the 48 GB Mac 27B turn from issue 499 is priced at 39.47 GiB and refused before its prompt is read, where 2.12.0 admitted it and reached 38.5 GiB against a 36 GiB limit. — [source](https://mtplx.com/releases/2.12.1/)
- A 250K-token prompt on a 64 GB Mac with 4-bit KV cache (issue 525) was projected at 7.6 GiB and needed 20.1 GiB; the 2.12.1 price covers the full-width rows that exist while 4-bit pages fill. — [source](https://mtplx.com/releases/2.12.1/)
- The 2.12.1 price covers a quantized KV snapshot at full width (15.3 GiB for a 250K-token 27B snapshot, not the 4.9 GiB of its 4-bit pages), MoE routed experts, the draft head's history (4,096 bytes per token on the 27B) and Gemma 4's own prompt pass. — [source](https://mtplx.com/releases/2.12.1/)
- When memory is short the 2.12.1 engine first clears its buffer pool, then runs Flash-Next's prompt in 2,048-token chunks instead of 4,096, then releases RAM copies of conversations that are not generating (oldest first, those already saved to SSD first); a generating conversation is never touched and a request for a conversation being released waits instead of getting a 409. — [source](https://mtplx.com/releases/2.12.1/)
- The 2.12.1 guard re-chooses the chunk width after memory is freed, so a prompt that fits once idle state is released keeps its 4,096-token chunks. — [source](https://mtplx.com/releases/2.12.1/)
- The 2.12.1 engine reads free, purgeable and file-backed pages (not macOS's memory level, which counted compressible memory as available) and checks them before every prompt chunk. — [source](https://mtplx.com/releases/2.12.1/)
- The 2.12.1 stop floor is the largest of 1 GiB, 2.5% of RAM and a sixteenth of the memory macOS has wired, the shed floor is twice that (6.2-6.4 GB and 12.4-12.8 GB in the 2 Oct session), and a prompt also stops when macOS has compressed another eighth of RAM during a run of requests or holds a quarter of RAM compressed while still losing ground. — [source](https://mtplx.com/releases/2.12.1/)
- The early compression macOS 27 does on a healthy Mac does not count toward the 2.12.1 stop. — [source](https://mtplx.com/releases/2.12.1/)
- 2.12.0 forgave host memory up to the larger of 8 GiB and RAM minus system reserve minus the limit, so lowering the limit raised the allowance and the ceiling stayed at 112 GiB on a 128 GB Mac at 96, 90 or 88 GiB limits; 2.12.1 allows a sixteenth of RAM (1 to 8 GiB, 8 GiB from 128 GB), caps it at a twelfth of an explicit limit, and measured a 98 GiB process ceiling at the 90 GiB default. — [source](https://mtplx.com/releases/2.12.1/)
- `MTPLX_HOST_MEMORY_ALLOWANCE_BYTES` is `auto` (RAM/16, at most 8 GiB, at most a twelfth of an explicit limit) and `0` makes every byte above MLX's account count, which reads a full session on a 48 GB Mac as critical. — [source](https://github.com/youssofal/MTPLX/blob/main/docs/server.md)
- The default `--memory-limit` is 75% of RAM; from 128 GB up it is at most RAM minus 38 GiB (90 GiB on a 128 GB Mac) and never above 192 GiB; Macs under 128 GB and from 192 GB up keep the 75% rule, and `--memory-limit max` equals the default on Macs with 32 GB or less. — [source](https://mtplx.com/releases/2.12.1/)
- With 16 GB of other apps open on a 128 GB Mac, a 96 GiB limit made the guard refuse a 74,000-token compaction-style request and a 90 GiB limit served it. — [source](https://mtplx.com/releases/2.12.1/)
- If memory does not allow the cache to grow during an answer, 2.12.1 ends the answer between two rounds with `finish_reason` `length` and `memory_stop` in the stats, keeps every streamed token, and prices the longest answer up front (the limit applied from 112K tokens in the 2 Oct session, 28,014 tokens at 205K). — [source](https://mtplx.com/releases/2.12.1/)
- If a guard step itself raises, the request still runs and `/health` reports `memory_guard.guard_degraded` with the error. — [source](https://mtplx.com/releases/2.12.1/)
- `--allow-swap` (or the app's swap setting) turns off the prompt-processing checks and the refusals before them. — [source](https://mtplx.com/releases/2.12.1/)
- A real Pi compaction of a 204,981-token Flash-Next session on 2.12.0 was refused with a 507 thirteen times in a row because the idle conversation kept its memory; on 2.12.1 its 138,774 and 21,597-token summary requests were served with swap flat at 2.08 GB and a lowest free memory of 259 MB. — [source](https://mtplx.com/releases/2.12.1/)
- The 2.12.1 notes list an unreleased gap: Pi compaction of a 205K-token session took 410 s on a busy Mac, the app estimated a 204,817-token re-read at 3 min 55 s against 7 min 43 s actual, and a compaction can still be stopped about 20 s in with a 507 under a heavy desktop. — [source](https://mtplx.com/releases/2.12.1/)
- MTPLX 2.12.1 left a 32 GB Mac's default 27B pick with an 8,192-token window, against 192,512 tokens for Ternary Bonsai 2 27B and about 77,824 for Qwen 3.8 27B Bare Speed. — [source](https://mtplx.com/releases/2.12.1/)
- On a 48 GB Mac the 27B keeps two copies of a restored conversation (the single live copy is Flash-Next only), so Pi is told 204,800 tokens but a warm turn near 100K is refused with a 507. — [source](https://mtplx.com/releases/2.12.1/)
- 2.12.1 refused ordinary prompts on 8, 16 and 32 GB Macs because it charged every model the 27B's 3 GiB per 2,048-token chunk (about 2 GiB for a one-line prompt) and counted the engine's normal memory as a leak; GitHub's 7 GB M1 runner failed with a 507 on a 64-token request. — [source](https://mtplx.com/releases/2.12.2/)
- 2.12.2 measured per-chunk prefill memory through `mtplx serve` over 43 requests (60 to 49,468 tokens): the 4B holds 0.19 GiB at 59 rows and 1.55-1.65 GiB at 2,048, Bonsai 2 27B 0.45 GiB at 98 rows and 2.69-2.72 GiB at 2,048, and the 27B 1.52-1.54 GiB at 1,024 and 2.52-2.55 GiB at 2,048. — [source](https://github.com/youssofal/MTPLX)
- The 2.12.2 dense-family prefill bill is the lower of (0.25 GiB plus 36 MLP rows per row) and (0.75 GiB plus 10 per row), which covers every reading by 0.09 to 0.25 GiB; families with routed experts, Flash-Next and Gemma 4 keep their 2.12.1 bills. — [source](https://github.com/youssofal/MTPLX)
- 2.12.2 lets a dense family run prefill at 1,024 or 512 rows before refusing (a routed-expert family at 1,024 only), picks a narrower width only where it saves at least 128 MiB, and notes a 512-row chunk holds about a live row per token of context (the 4B: 2.22 GiB at 49K) so it suits only short and medium prompts. — [source](https://github.com/youssofal/MTPLX)
- A 1,024-token chunk uses about 1 GiB less than a 2,048-token chunk on the 27B models, and a 7,000-token Bonsai prompt took 9.5 s, 9.6 s and 9.7 s at 2,048, 1,024 and 512 rows in single runs. — [source](https://mtplx.com/releases/2.12.2/)
- 2.12.2 charges memory held outside the GPU allocator only past 4 GiB on Macs under 64 GB (a healthy engine holds 2.1-2.9 GiB with Bonsai and 2.3-4.0 GiB with the 27B), while Macs with 64 GB or more keep a sixteenth of RAM up to 8 GiB; 2.12.1 allowed RAM/16, which is 1 GiB on a 16 GB Mac. — [source](https://mtplx.com/releases/2.12.2/)
- In 2.12.2 the 99,355-token 27B turn from issue 499 on a 48 GB Mac runs in 1,024-token chunks priced at 34.87 GiB against the 36 GiB limit, in tests with the reporter's numbers. — [source](https://mtplx.com/releases/2.12.2/)
- 2.12.2's known issues: the engine keeps at least 1 GiB free for macOS, so with 2 GiB free on an 8 GB Mac a 2,500-token prompt was refused 0.2 GiB short and with 3 GiB free on a 16 GB Mac 6,300 and 6,800-token prompts were refused 0.1-0.2 GiB short. — [source](https://mtplx.com/releases/2.12.2/)
- The 2.12.2 before-and-after table on an M5 Max set up as each Mac (limits 8, 12 and 24 GiB for 8, 16 and 32 GB seats) shows 2.12.0 served the 16 and 32 GB rows, 2.12.1 refused them, and 2.12.2 served them, including a five-turn agent session on the 4B with 3 GiB free. — [source](https://mtplx.com/releases/2.12.2/)
- An earlier release reserved a static 3 GiB for generation transients while a deep chunked prefill measured up to 12.4 GiB peak over active memory, so the session bank now reserves the spike the process has observed (clamped between 3 GiB and half the post-weights memory, at most 16 GiB) and demotes idle entries to SSD before the next spike. — [source](https://github.com/youssofal/MTPLX/blob/main/CHANGELOG.md)
- An earlier release added `--memory-budget` to fit MTPLX inside a declared RAM envelope, a RAM-tiered cap on the MLX allocator cache, a re-clamped per-session admission gate on smaller machines, and a paged KV pool that stops growing past the context window. — [source](https://github.com/youssofal/MTPLX/blob/main/CHANGELOG.md)
- MTPLX 2.12.1's known issues list host memory still growing in very long sessions (issue 546, partly fixed by capping compiled verify programs at 16 versions each via `MTPLX_COMPILED_VERIFY_TRACES_PER_PROGRAM`) and a Metal shared-event leak behind "Failed to create Metal shared event" fixed in issue 544. — [source](https://mtplx.com/releases/2.12.1/)
