oMLX memory guard tiers and prefill ceiling on 16-24 GB Macs
Parent: Mac local LLMs: oMLX, Rapid-MLX and related internals · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
PR 3933 redefines the tiers as memory kept free for other apps: `safe` 20% of RAM (6-16 GB), `balanced` 8% (3-8 GB), `aggressive` 2% (1.5-4 GB), and `aggressive` may count half of other apps' active memory as compressible.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- PR 3933 redefines the tiers as memory kept free for other apps: `safe` 20% of RAM (6-16 GB), `balanced` 8% (3-8 GB), `aggressive` 2% (1.5-4 GB), and `aggressive` may count half of other apps' active memory as compressible. [source]
- PR 3933 reads usage from the `task_vm_info` graphics ledger and MLX counters instead of `phys_footprint` deltas, which lag behind freed Metal buffers and skewed per-chunk rates. [source]
- PR 3933's admission line is min(hard limit, hard watermark, hard x headroom), with headroom 90% for safe, 92% for balanced, 97% for aggressive and 95% for a custom ceiling; a request resumed after an eviction pause is not re-admitted and `speed` priority charges the widest scheduled step. [source]
- PR 3933 grows each KVCache (and PoolingCache through a hook) once to the full prompt after the first chunk, so mlx-lm's 256-token concat no longer strands full KV copies in the Metal pool during prefill. [source]
- PR 3933 sends the Qwen4 PLE table and the DeepSeek V4.1 Engram to SSD when a resident load would leave no room for a prompt, and re-reads a dynamic ceiling before refusing a load right after an unload. [source]
- On an M2 MacBook Air 16 GB (macOS 26.6.2, Qwen3.5-9B-oQ8 at 10.5 GB, Metal cap 11.84 GB) PR 3933 changed a 17.6K prompt from aborted at 7,808 tokens after 43 s to completing at 11.36 GB peak, and a 23.4K prompt from rejected to completed. [source]
- On that 16 GB Air, loading under `safe` or `balanced` is now refused with less than the tier reserve left free, while main loaded at every tier with 1.1-1.8 GB of swap; the context bench reached 32,768 tokens in one attempt. [source]
- On the 16 GB Air three 11.7K prompts sent 2 s apart gave 2 of 3 completions (the third rejected at admission) against 0 of 3 before, and cancelling a 17.5K prefill then sending 23.4K completed at 11.51 GB peak instead of aborting at 8,288 after 47 s. [source]
- On an M5 Max 128 GB PR 3933 raised the Qwen3.8-27B 44 GB-custom context bench from 143,360 tokens (two attempts) to 262,144 on the first attempt, and Flash-Next (99 GB) under `aggressive` + `speed` from 51,200 to 223,232. [source]
- PR 3933's behavior change: the final KV must fit under hard x headroom, so prompts in that band that were admitted and then aborted are now rejected at admission, and with other apps open `safe` and `balanced` admit less than before. [source]
- PR 3933 treats saved soft/hard thresholds equal to the old defaults as 'use the tier default'. [source]
- PR 3933 lists known limits: decode still grows KV by 256-token concat so a full 16 GB cache reads 32 GB right after the step; the bootstrap chunk estimate for Qwen3.5-family hybrids is about 2.5x high on the first request after load; embedded GDN snapshots and vision encoding are not priced. [source]
- Issue 4213 (M5 Max 128 GB, `balanced`, about 139K context) shows `static cap is 120.00 GB but only 33.73 GB is reclaimable right now`, and the log text tells the user to raise `memory_guard_tier` safe -> balanced -> aggressive. [source]
- A second 4213 reporter had a 26,764-token request rejected at 73.88 GB peak against a 73.54 GB dynamic ceiling and saw the 0.7.0rc1 app report a balanced process ceiling of 101.0 GB after a rollback. [source]
- Two further 4213 reporters (M5 128 GB with a 200K server; headless M5 Max Studio 64 GB with Qwen3.8-27B oQ5e) say rc1 worked and 0.7.0 aborts near 140K-156K tokens, the latter with a 20 GB Metal pool that was not clearing. [source]
- Issue 4213 had no maintainer reply on the cached page 2026-10-04. [source]
- Issue 3284 measured in-flight growth of about 237 KB per token on an ArraysCache hybrid against the 27.28 KB per token the guard estimated, so doomed 200K prefills were admitted and died at about 160K after 10-15 minutes. [source]
- Issue 3284 notes `aggressive` cannot help when the Metal cap binds (115.90 GB = 0.95 x `iogpu.wired_limit_mb` 122 GB). [source]
- Issue 3737 logged `ProcessMemoryEnforcer: could not resolve scheduler for engine type BatchedEngine - prefill memory guard will not propagate to this engine` at unload, and `Released ANE prefill banks ... freed 11.94GB` when a 16K prefill plus a VLM MTP drafter exceeded headroom on a 64 GB Mac. [source]
- After the ANE banks are released the model serves GPU-only prefill until its next load; the 16K prefill rate fell from 302 to 239 tok/s in that report. [source]
- Issue 3956 reports that with `prefill_memory_guard=false` a second 181 GB-class GLM load was attempted on top of a resident one (`projected 197.54GB > ceiling 124.00GB`) and ended in a Metal Insufficient Memory error. [source]
- Issue 3683 reports a Metal `kIOGPUCommandBufferCallbackErrorOutOfMemory` when an image is attached at about 160K context (80 GB steady, 85 GB peak) with `iogpu.wired_limit_mb=92000`; PR 3933 states vision encoding is not priced. [source]
- Issue 4175's hot-cache regression (reconstruct 18-73x slower) is suspected to trace to issue 1833, where `ProcessMemoryEnforcer` still counts resident hot-cache bytes; the scheduler-side double count (issue 1796) was fixed in commit e52aeb9 in 0.5.2.dev1. [source]
Corrections and disagreements
- CONTRADICTS: disk-and-ssd-tiered-kv-cache-servers.md (lines 12 and 73, default ceiling is total RAM minus 8 GB): from PR 3933 numbers `balanced` reserves 8% of RAM clamped to 3-8 GB, so RAM minus 8 GB holds only at about 100 GB or more, and on 16 GB and 24 GB Macs the reserve is the 3 GB floor. [source]
Children
- No children recorded.