<!-- llms-explorer concept facts · https://llms-explorer.com/tree/mlx-optiq-residency-ceiling-near-73-percent-of-r/ · pack 2026-10-05 · ~1440 tokens -->

# mlx-optiq residency ceiling near 73 percent of RAM

> The vendor gives no mechanism for the cliff. It says only that system memory is free, nothing swaps and no error is raised, and that raising `iogpu.wired_limit_mb` from 20480 to 22528 does not move it (18.1 GiB reads 0.1 tok/s at both).

Parent: [Mac local LLMs: Memory and wired limits](https://llms-explorer.com/tree/mac-local-llms-memory-and-wired-limits/) · 1 facets · 25 facts · page: https://llms-explorer.com/tree/mlx-optiq-residency-ceiling-near-73-percent-of-r/

## Facts

- The vendor gives no mechanism for the cliff. It says only that system memory is free, nothing swaps and no error is raised, and that raising `iogpu.wired_limit_mb` from 20480 to 22528 does not move it (18.1 GiB reads 0.1 tok/s at both). — [source](https://mlx-optiq.com/docs/cluster)
- optiq's preflight prices the model against each node's currently free memory minus a per-rank run-reserve for activations, KV and the MLX buffer pool, and refuses to load rather than swap. Its refusal text shows rank 0 "free 25.9 GiB - 0.9 GiB to run" and rank 1 "free 21.5 GiB - 3.2 GiB to run". — [source](https://mlx-optiq.com/docs/cluster)
- The per-rank reserve is an environment variable, `OPTIQ_CLUSTER_HEADROOM_GB`, taking a comma list in rank order. The example `"0.4,4.1"` places 30 layers (26.2 GiB) on the 36 GB node and 18 layers (16.6 GiB) on the 24 GB node and yields 20.5 tok/s. — [source](https://mlx-optiq.com/docs/cluster)
- The load log prints a per-rank "cap" and a "GPU limit": rank 1 (24 GB) shows resident 16.6 GiB, cap 17.3 GiB, free 21.4, GPU limit 22.0; rank 0 (36 GB) shows resident 26.2 GiB, cap 26.6 GiB, free 27.0, GPU limit 30.0. — [source](https://mlx-optiq.com/docs/cluster)
- The cluster page and the mlx-optiq PyPI page carry the same 4.2x claim (20.5 tok/s two-Mac resident versus 4.9 tok/s one-Mac streaming); no dated change log for the ceiling statement was found. — source: `asserted`
- Internal inconsistency in the page: the prose says "nothing swaps", but its own table labels the 19.6 GiB row "0.05 tok/s (swapping)". The cliff at 17.5-18.1 GiB and the swapping at 19.6 GiB may be two regimes, with only the first free of swap. — [source](https://mlx-optiq.com/docs/cluster)
- The two tested rows near the cliff differ in `iogpu.wired_limit_mb` (20480 for 16.6, 17.3 and the first 18.1 row; 22528 for the second 18.1 and 19.6 rows), so only the 18.1 GiB pair isolates the wired limit; the other rows confound it. — [source](https://mlx-optiq.com/docs/cluster)
- The page says a Mac whose swap has filled "stays degraded until it reboots. macOS never shrinks the swap file", which would make a later run on a node that once swapped look like a ceiling effect. — [source](https://mlx-optiq.com/docs/cluster)
- Throughput is reported flat in context for the cluster path, so the ceiling cannot be a KV-growth effect on that path. — [source](https://mlx-optiq.com/docs/cluster)
- 73% versus the kernel wire table. The XNU table in xnu-vm-per-task-user-wire-limit-and-global-wire.md gives 76% for RAM from 16 to under 32 GiB, which is 18.2 GiB on a 24 GiB Mac; the observed collapse sits between 17.3 GiB (72%) and 18.1 GiB (75%). The 73% entry in that table belongs to an 8 GiB machine, so the match is a coincidence unless the cliff is not the wire limit. — source: `asserted`
- 73% versus Metal's own limit. The log's "GPU limit 22.0" on the 24 GB node is 92% of RAM, so the collapse occurs well below the Metal working-set limit the allocator reports. The ceiling is therefore not the MLX or Metal cap. — source: `asserted`
- The two nodes do not sit at the same fraction: the 24 GB node's cap is 72% of RAM and the 36 GB node's cap is 74%, so "73%" is a rounded figure from one 24 GB machine, not a measured constant for both. — source: `asserted`
- Whether the cliff exists on a 16 GB, 32 GB or 64 GB Mac, or tracks absolute free memory instead of a RAM fraction. No other machine was tested in the page. — source: `asserted`
- Whether `vm_stat` shows rising decompressions or pageins at the cliff. — source: `asserted`
- Which macOS and MLX versions the table used; the page states none. — source: `asserted`
- The mlx-optiq cluster page states the ceiling as "roughly 73% of its RAM" and measured it on a 24 GB M4 with cliff near 17.5 GiB. — [source](https://mlx-optiq.com/docs/cluster)
- The page's 18.1 GiB row reads 0.1 tok/s at both `iogpu.wired_limit_mb` 20480 and 22528. — [source](https://mlx-optiq.com/docs/cluster)
- The page's prose says "nothing swaps" at the cliff while its table labels the 19.6 GiB row "swapping". — [source](https://mlx-optiq.com/docs/cluster)
- `OPTIQ_CLUSTER_HEADROOM_GB="0.4,4.1"` assigns 30 layers (26.2 GiB) to a 36 GB node and 18 layers (16.6 GiB) to a 24 GB node for 20.5 tok/s. — [source](https://mlx-optiq.com/docs/cluster)
- The load log shows rank 1 with cap 17.3 GiB and GPU limit 22.0 on 24 GB, and rank 0 with cap 26.6 GiB and GPU limit 30.0 on 36 GB. — [source](https://mlx-optiq.com/docs/cluster)
- optiq's preflight refusal subtracts a per-rank run-reserve (0.9 GiB and 3.2 GiB in the shown example) from free memory. — [source](https://mlx-optiq.com/docs/cluster)
- The page says a Mac whose swap is full stays degraded until it reboots because macOS never shrinks the swap file. — [source](https://mlx-optiq.com/docs/cluster)
- The XNU wire-limit table's 24 GiB entry is 76%, not 73%, so the optiq cliff is not explained by `vm.global_user_wire_limit`. — source: `asserted`
- The observed 73% ceiling sits about 19 percentage points below the 92% "GPU limit" the same log prints for the node. — source: `asserted`
- The ceiling is one vendor measurement on one 24 GB M4 and is not shown to be a constant fraction. — source: `asserted`
