Mac local LLMs: oMLX, Rapid-MLX and related internals
Parent: Running LLM models locally on a Mac · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
oMLX (Apache-2.0, macOS 15+, Python 3.11-3.13, M1-M5): multi-model EnginePool, menu-bar app, /admin dashboard, persistent SSD cache, embeddings and rerank on one endpoint. Install: signed DMG, `brew install jundot/omlx/omlx` (CLI only; `brew tap jundot/omlx` first).
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Choose and install
- oMLX (Apache-2.0, macOS 15+, Python 3.11-3.13, M1-M5): multi-model EnginePool, menu-bar app, /admin dashboard, persistent SSD cache, embeddings and rerank on one endpoint. Install: signed DMG, `brew install jundot/omlx/omlx` (CLI only; `brew tap jundot/omlx` first). [source]
- Rapid-MLX: one model served headless on port 8000, Ollama-style verbs. `brew install rapid-mlx`, `uv tool install rapid-mlx@latest` or `pip install rapid-mlx`; extras `[vision]` `[audio]` `[all]`. `rapid-mlx launch claude-code` patches ~/.claude/settings.json. [source]
- Both ship several releases a week with fresh regressions (oMLX 0.7.0 memory guard and MiniMax-M3; Rapid-MLX 0.15.5 wedge): pin a version, keep the previous one. No independent oMLX vs Rapid-MLX benchmark exists. [source]
- oMLX custom kernels (GLM-5.2, MiniMax M3, Qwen3.5) are precompiled only in the DMG; source builds need full Xcode, else GLM-5.2 prefill is ~29 vs 845 tok/s. [source]
- oMLX 0.7.0 (2026-09-30): `server.gpu_keep_warm_interval` (saves ~1-1.5 s TTFT), live `max_concurrent_requests`, MoE expert SSD offload. [source]
oMLX memory guard and pressure levels
- Tiers keep memory free for other apps: safe 20% of RAM (6-16 GB), balanced 8% (3-8 GB), aggressive 2% (1.5-4 GB); admission = min(hard limit, watermark, hard x headroom 90/92/97%). [source]
- Failure: #4213 prefill rejected "~78.05 GB peak ... dynamic ceiling is 77.82 GB". Raise `memory_guard_tier` safe -> balanced -> aggressive. [source]
- Guard does not price vision encoding (#3683 OOM); `aggressive` cannot beat the Metal cap (0.95 x `iogpu.wired_limit_mb`). [source]
- ProcessMemoryEnforcer: soft pressure evicts idle LRU models only while >1 loaded and pauses admission; hard unloads even the last idle model, or aborts a sole busy model's requests (unload only on emergency: over ceiling by 2 GiB or 2 polls). [source]
oMLX store-cache gate and stalled admission
- Stores run on one worker thread (`omlx-store-cache`, max_workers=1), gated by `_StoreCacheGate`; its cap steps up toward `max_num_seqs` at ok pressure and down to 1 at soft/hard, one step per poll tick. A slot frees when cache references are released, not when the write finishes. [source]
- Gate full: admission defers (debug `Admission deferred: store-cache pipeline full`). After 60 s the request fails with `finish_reason="error"`, `error_code="store_cache_admission_stalled"` ("store-cache cleanup stayed full ... retry after the cache writer drains"); at the memory limit the code is `memory_admission_stalled`. Fix: retry or shrink `ssd_cache_max_size`. [source]
oMLX SSD and hot cache
- Eviction is global LRU, size-driven; no per-model quota (#2663, #3612 unanswered). [source]
- Set `ssd_cache_max_size` explicitly on small disks: `auto` = 50% of (free disk + existing cache) since 0faa19c. Clear via `POST /admin/api/ssd-cache/clear` only on 0.7.0rc1+. [source]
- `hot_cache_max_size 4GB` made reconstruct 18x-73x slower (#4175, M4 Max 36 GB). For single-user chats use `--hot-cache-write-through` / `OMLX_HOT_CACHE_WRITE_THROUGH` (default off). [source]
Hybrid GDN, MTP and boundary snapshots
- Non-sliceable state is stored only at block boundaries. Log `reason=boundary_snapshot_unavailable ... available_boundaries=0` counts only the current request; judge reuse by `cached_tokens`. On 0.6.4 qwen3_5 with Lightning MTP, #3317 re-prefilled 76K prompts (496 s reply); 0.6.3 is clean. Use 0.7.0+. [source]
oMLX engines and Metal stream clears
- Glm5_next is VLM-only: forced-LM reload fails `Model type glm5_next not supported` (#3959). MiniMax-M3 0.7.0 fails `'KVCache' object has no attribute 'keys_and_values'` (#4228). [source]
- Clears must drain the engine's own stream: PR 2412/2413 pass `self._stream` to `_sync_and_clear_cache` (bare `mx.synchronize()` drains Stream 0 and left ~70 ms in flight). Still bare on main: `engine/base.py` after every non-streaming request, engine_pool, dflash, vlm, oq. [source]
oMLX pressure reclaim and executors
- Hard pressure sets `_pending_pressure_clear`: it drains at the next step boundary even under load (idle reclaim waits for no requests), and prefill clears after each chunk since one step can last minutes. [source]
- `OMLX_DISABLE_PRESSURE_RECLAIM=1` turns off the hard-level reclaim request; grace gate is 5 polls. Unknown: whether 5 polls cover a slow prefill chunk. [source]
- Two executors: global `mlx-global` thread (keep-warm, default stream) and per-engine `mlx-engine-<id8>` thread with its own stream; they can overlap on GPU, so "serialized" holds only among global users. Decode burst: `OMLX_DECODE_BURST_MAX_STEPS=64`, `_BUDGET_SINGLE_S=0.1`, `_BUDGET_S=0.03` (~80 vs ~74 tok/s sync vs async). Engine close() reclaims on its own thread; timeout means `fatal_exit`. [source]
oMLX health, teardown and wedges
- `/health` is unauthenticated and returns 200 `healthy`, or 503 `loading` until pinned preload finishes; it reads no scheduler state, so it stayed green ~14 h in a 0.5.3 wedge. Probe a tiny real completion. [source]
- Unload must finish in 60 s or the process exits with code 70. A 420 GB GLM-5.2 unload hit it (#2334, open); a 40 GB model took ~26 s. [source]
- Killing a wedged server strands wired memory (#2184, open): 404 GiB after `bootout` of a wedged server vs 5 GiB after a healthy one;. Use `POST /v1/models/{id}/unload` before `launchctl bootout`, launchd `ExitTimeOut` 180, never `kickstart -k`. Since 2026-07-11 the port binds before preload. [source]
Rapid-MLX, MTPLX, Osaurus
- Rapid-MLX 0.15.5 wedge (#4108): Metal memory climbs ~25-36 GB over 19 Claude Code requests, then HTTP 503 "max concurrent requests" until restart. PFlash `always` silently drops ~80% of the middle of no-tools prompts above ~11.5K tokens (#4092). [source]
- MTPLX memory guard: default 75% of RAM; over-limit requests get HTTP 507 before prefill. 2.12.1 over-refused small Macs; use 2.12.2. Codex's hosted `web_search` gets a 400 unless removed. Osaurus: 127.0.0.1:1337, `brew install --cask osaurus`. [source]
Corrections to earlier claims
- Default oMLX ceiling is not RAM minus 8 GB: balanced reserves 8% clamped to 3-8 GB, so 16 and 24 GB Macs get the 3 GB floor. [source]
Open
- Whether #4213 is tier calibration or a bug; whether any release closed #3317, #2624, #2334, #2184; MTPLX exactness unverified. [source]
Corrections and disagreements
- CONTRADICTS: disk-and-ssd-tiered-kv-cache-servers.md (lines 12 and 73, default ceiling is total RAM minus 8 GB): from PR 3933 numbers `balanced` reserves 8% of RAM clamped to 3-8 GB, so RAM minus 8 GB holds only at about 100 GB or more, and on 16 GB and 24 GB Macs the reserve is the 3 GB floor. [source]
- CONTRADICTS prompt-cache-survival-across-model-unload-swap-and-restart.md (lines 13, 76): incompatible blocks are not "not evicted"; they are LRU-evicted when over cap at startup and on writes needing space, but not deleted on scan. [source]
- CONTRADICTS: omlx-per-model-ssd-cache-isolation-quotas-and-stale-block-eviction.md (claim that issue 1578 records a maintainer position against partial-tail snapshots for hybrid GDN): the maintainer later closed 1578 as completed via PR 3835 (commit 8288884, 2026-09-22), which stores a prefix tail block. [source]
- CONTRADICTS: omlx-store-cache-admission-gate-and-60-s-admissi.md only in timing granularity: the cap walk runs per enforcement tick (1 s under pressure, but 10 s or 30 s when idle and ok), not per pressure transition. [source]
- CONTRADICTS: omlx-metal-cache-clears-wait-for-the-owning-stre.md open question implies the sites may have been converted; on main only the `vlm_mtp.py` and GLM decode sites are, and the engine, engine_pool, dflash, vlm and oq sites are not. [source]
Concepts in this cluster
- oMLX and Rapid-MLX runtimes [source]
- MTPLX MTP self-drafting Mac server [source]
- oMLX VLM versus text engine selection and per-model settings [source]
- oMLX memory guard tiers and prefill ceiling on 16-24 GB Macs [source]
- Hybrid and recurrent model checkpoint handling in oMLX [source]
- oMLX per-model SSD cache isolation quotas and stale-block eviction [source]
- oMLX GDN sidecar and rotating-window snapshot size [source]
- oMLX hybrid GDN boundary snapshot capture regression [source]
- oMLX trailing partial block and restore-latency limits [source]
- MTPLX memory guard and request pricing on 8-128 GB Macs [source]
- osaurus vmlx-swift-lm runtime and vmlxctl [source]
- Osaurus inference scheduler, model leases and single-residency handoff [source]
- oMLX issue 2624: engine stops dispatching for 6 hours with every thread idle [source]
- oMLX boundary snapshot store serialization and async cache-store worker locking [source]
- oMLX issue 2330: scheduler never dispatches after two identical cached prefills [source]
- oMLX store-cache admission gate and 60 s admission stall error [source]
- oMLX fatal teardown watchdog and wired memory stranding (issues 2334 and 2184) [source]
- oMLX Metal cache clears wait for the owning stream (PRs 2412 and 2413) [source]
- oMLX ProcessMemoryEnforcer pressure levels and adjust_store_cache_cap [source]
- oMLX /health liveness versus scheduler wedge [source]
- oMLX remaining ad-hoc mx.clear_cache sites needing stream analysis [source]
- oMLX OMLX_DISABLE_PRESSURE_RECLAIM and request_pressure_reclaim scheduler path [source]
- oMLX shared MLX executor thread versus per-engine streams [source]
Children
- MTPLX memory guard and request pricing on 8-128 GB Macs
- MTPLX MTP self-drafting Mac server
- oMLX and Rapid-MLX runtimes
- oMLX boundary snapshot store serialization and async cache-store worker locking
- oMLX fatal teardown watchdog and wired memory stranding (issues 2334 and 2184)
- oMLX GDN sidecar and rotating-window snapshot size
- oMLX /health liveness versus scheduler wedge
- oMLX hybrid GDN boundary snapshot capture regression
- oMLX issue 2330: scheduler never dispatches after two identical cached prefills
- oMLX issue 2624: engine stops dispatching for 6 hours with every thread idle
- oMLX memory guard tiers and prefill ceiling on 16-24 GB Macs
- oMLX Metal cache clears wait for the owning stream (PRs 2412 and 2413)
- oMLX OMLX_DISABLE_PRESSURE_RECLAIM and request_pressure_reclaim scheduler path
- oMLX per-model SSD cache isolation quotas and stale-block eviction
- oMLX ProcessMemoryEnforcer pressure levels and adjust_store_cache_cap
- oMLX remaining ad-hoc mx.clear_cache sites needing stream analysis
- oMLX shared MLX executor thread versus per-engine streams
- oMLX store-cache admission gate and 60 s admission stall error
- oMLX trailing partial block and restore-latency limits
- oMLX VLM versus text engine selection and per-model settings
- Osaurus inference scheduler, model leases and single-residency handoff
- osaurus vmlx-swift-lm runtime and vmlxctl
- Hybrid and recurrent model checkpoint handling in oMLX