mlx-optiq residency ceiling near 73 percent of RAM
Parent: Mac local LLMs: Memory and wired limits · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
The vendor gives no mechanism for the cliff. It says only that system memory is free, nothing swaps and no error is raised, and that raising `iogpu.wired_limit_mb` from 20480 to 22528 does not move it (18.1 GiB reads 0.1 tok/s at both).
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- The vendor gives no mechanism for the cliff. It says only that system memory is free, nothing swaps and no error is raised, and that raising `iogpu.wired_limit_mb` from 20480 to 22528 does not move it (18.1 GiB reads 0.1 tok/s at both). [source]
- optiq's preflight prices the model against each node's currently free memory minus a per-rank run-reserve for activations, KV and the MLX buffer pool, and refuses to load rather than swap. Its refusal text shows rank 0 "free 25.9 GiB - 0.9 GiB to run" and rank 1 "free 21.5 GiB - 3.2 GiB to run". [source]
- The per-rank reserve is an environment variable, `OPTIQ_CLUSTER_HEADROOM_GB`, taking a comma list in rank order. The example `"0.4,4.1"` places 30 layers (26.2 GiB) on the 36 GB node and 18 layers (16.6 GiB) on the 24 GB node and yields 20.5 tok/s. [source]
- The load log prints a per-rank "cap" and a "GPU limit": rank 1 (24 GB) shows resident 16.6 GiB, cap 17.3 GiB, free 21.4, GPU limit 22.0; rank 0 (36 GB) shows resident 26.2 GiB, cap 26.6 GiB, free 27.0, GPU limit 30.0. [source]
- The cluster page and the mlx-optiq PyPI page carry the same 4.2x claim (20.5 tok/s two-Mac resident versus 4.9 tok/s one-Mac streaming); no dated change log for the ceiling statement was found. [source]
- Internal inconsistency in the page: the prose says "nothing swaps", but its own table labels the 19.6 GiB row "0.05 tok/s (swapping)". The cliff at 17.5-18.1 GiB and the swapping at 19.6 GiB may be two regimes, with only the first free of swap. [source]
- The two tested rows near the cliff differ in `iogpu.wired_limit_mb` (20480 for 16.6, 17.3 and the first 18.1 row; 22528 for the second 18.1 and 19.6 rows), so only the 18.1 GiB pair isolates the wired limit; the other rows confound it. [source]
- The page says a Mac whose swap has filled "stays degraded until it reboots. macOS never shrinks the swap file", which would make a later run on a node that once swapped look like a ceiling effect. [source]
- Throughput is reported flat in context for the cluster path, so the ceiling cannot be a KV-growth effect on that path. [source]
- 73% versus the kernel wire table. The XNU table in xnu-vm-per-task-user-wire-limit-and-global-wire.md gives 76% for RAM from 16 to under 32 GiB, which is 18.2 GiB on a 24 GiB Mac; the observed collapse sits between 17.3 GiB (72%) and 18.1 GiB (75%). The 73% entry in that table belongs to an 8 GiB machine, so the match is a coincidence unless the cliff is not the wire limit. [source]
- 73% versus Metal's own limit. The log's "GPU limit 22.0" on the 24 GB node is 92% of RAM, so the collapse occurs well below the Metal working-set limit the allocator reports. The ceiling is therefore not the MLX or Metal cap. [source]
- The two nodes do not sit at the same fraction: the 24 GB node's cap is 72% of RAM and the 36 GB node's cap is 74%, so "73%" is a rounded figure from one 24 GB machine, not a measured constant for both. [source]
- Whether the cliff exists on a 16 GB, 32 GB or 64 GB Mac, or tracks absolute free memory instead of a RAM fraction. No other machine was tested in the page. [source]
- Whether `vm_stat` shows rising decompressions or pageins at the cliff. [source]
- Which macOS and MLX versions the table used; the page states none. [source]
- The mlx-optiq cluster page states the ceiling as "roughly 73% of its RAM" and measured it on a 24 GB M4 with cliff near 17.5 GiB. [source]
- The page's 18.1 GiB row reads 0.1 tok/s at both `iogpu.wired_limit_mb` 20480 and 22528. [source]
- The page's prose says "nothing swaps" at the cliff while its table labels the 19.6 GiB row "swapping". [source]
- `OPTIQ_CLUSTER_HEADROOM_GB="0.4,4.1"` assigns 30 layers (26.2 GiB) to a 36 GB node and 18 layers (16.6 GiB) to a 24 GB node for 20.5 tok/s. [source]
- The load log shows rank 1 with cap 17.3 GiB and GPU limit 22.0 on 24 GB, and rank 0 with cap 26.6 GiB and GPU limit 30.0 on 36 GB. [source]
- optiq's preflight refusal subtracts a per-rank run-reserve (0.9 GiB and 3.2 GiB in the shown example) from free memory. [source]
- The page says a Mac whose swap is full stays degraded until it reboots because macOS never shrinks the swap file. [source]
- The XNU wire-limit table's 24 GiB entry is 76%, not 73%, so the optiq cliff is not explained by `vm.global_user_wire_limit`. [source]
- The observed 73% ceiling sits about 19 percentage points below the 92% "GPU limit" the same log prints for the node. [source]
- The ceiling is one vendor measurement on one 24 GB M4 and is not shown to be a constant fraction. [source]
Children
- No children recorded.