Metal residency sets and llama.cpp Metal internals
Parent: Mac local LLMs: llama.cpp internals · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
The cost was first measured with the `llama-idle` tool (PR 10119): sleep t, decode one token, repeat. On M2 Ultra with Llama-3.1-8B, F16 decode was 29 ms for pauses up to 1000 ms, then 227 ms at 1200 ms and 413-472 ms from 1400 ms up. Q8_0 went 19 ms to 106 ms (1200 ms) to 225-299 ms. Q4_0 went 1...
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- The cost was first measured with the `llama-idle` tool (PR 10119): sleep t, decode one token, repeat. On M2 Ultra with Llama-3.1-8B, F16 decode was 29 ms for pauses up to 1000 ms, then 227 ms at 1200 ms and 413-472 ms from 1400 ms up. Q8_0 went 19 ms to 106 ms (1200 ms) to 225-299 ms. Q4_0 went 14 ms to 24 ms (1000 ms) to 86 ms (1200 ms) to 158-216 ms. The extra time grows with model size. It occurred on M1 Pro and M2 Ultra, only on the GPU, not on CUDA. [source]
- ggerganov first guessed a power-saving throttle and proposed a "heartbeat" dummy kernel every 0.9 s. He then found the real cause: Awni Hannun (MLX) said memory is unwired after a while, re-wiring is expensive and scales with model size, and suggested `sudo sysctl iogpu.disable_wired_collector=1` (macOS 15+) to test it. That sysctl fixed the idle slowdown in the llama-idle test. [source]
- The programmatic alternative was residency sets, merged as llama.cpp PR 11427 (Jan 26, 2025). Reported gain: about 250 ms faster per request on M2 Ultra with a 7B Q8_0 model. Steady-state pp512/tg128 on 3B models was unchanged (speedup 1.00). [source]
- PR 11427 text: without residency sets the OS collects GPU memory "after 1 second of inactivity"; `GGML_METAL_NO_RESIDENCY` restores that and "hurts the performance of the application". The 3-minute keep-alive was added later; the original PR kept sets alive permanently. [source]
- A Metal engineer's comment in the PyTorch MPS thread (pytorch issue 124056) says the same: with commit-and-continue, the memory wired for the processed command buffer is released after a sleep ("I think there is a 1s or a fixed counter"), and mapping it back slows the next call. [source]
- One `MTLResidencySet` per `ggml_metal_buffer`, labelled `ggml_metal`, `initialCapacity = n_buffers`. At buffer init: `addAllocation` for every MTLBuffer view, then `commit`, then `requestResidency`. At buffer free: `endResidency`, `removeAllAllocations`, `commit`, release. The sets are not attached to the command queue; PR 11427 says attaching is not necessary. [source]
- Each set is also registered in a device-level collection (`ggml_metal_rsets`, an `NSMutableArray` guarded by an `NSLock`). Only that collection is touched by the heartbeat thread. [source]
- The heartbeat thread runs for the whole life of the device and wakes every 5 ms even when the counter is 0 (the `usleep` is outside the `if`). Only the `requestResidency` calls are gated by the counter. The counter starts at `loops_per_s * keep_alive_s` = 200 x 180 = 36,000 loops and is decremented once per loop while positive. Each graph compute stores the full value back (`ggml_metal_device_rsets_keep_alive`). So the heartbeat re-requests residency on every set 200 times a second for the 180 s after the last compute, then stops re-requesting but keeps polling. [source]
- Startup logs: `creating a residency set collection (keep_alive = 180 s)` and `use residency sets = true|false`. [source]
- `GGML_METAL_RESIDENCY_KEEP_ALIVE_S` is read with `atoi`. Non-numeric text parses to 0, which falls back to 180 s, so it cannot disable the timer. A very large value keeps memory wired until the process exits. [source]
- The model weights are not copied. Mmap'd weights are wrapped with `newBufferWithBytesNoCopy` in `MTLResourceStorageModeShared`, page-aligned. If the data is larger than `maxBufferLength`, the code creates overlapping views of `max_buffer_size` bytes stepped by `max_buffer_size - overlap`. The overlap is the largest tensor rounded up plus two pages, so every tensor fits fully in one view. Each view is added to the same residency set. [source]
- A dedicated llama.cpp debug log (`GGML_LOG_DEBUG`, compiled out when `GGML_METAL_NDEBUG` is set) prints per allocation `allocated buffer, size = ... MiB, (currentAllocatedSize / recommendedMaxWorkingSetSize)` and warns `current allocated size is greater than the recommended max working set size` when over. [source]
- `ggml_metal_device_get_memory` returns total = `recommendedMaxWorkingSetSize` and free = total - `currentAllocatedSize` (0 if over). `max_working_set_size` falls back to `maxBufferLength` only on macOS < 10.12. [source]
- `ggml_metal_rsets_free` asserts the collection is empty ("you haven't deallocated all Metal resources before exiting"), so freeing the device before the buffers aborts. [source]
- MLX `ResidencySets` (mlx/backend/metal/resident.cpp, resident.h) is off on devices without `MTLGPUFamilyMetal3` or on macOS < 15. It is a no-op until `mx.set_wired_limit` is called (default 0). [source]
- MLX uses several residency sets, not one. Each set holds a standing `requestResidency()`. The per-set cap is `MLX_RESIDENCY_SET_MAX_PCT` percent of `recommendedMaxWorkingSetSize`, default 5, with a floor of 64 MiB; a value <= 0 or >= 100 gives one set. Rationale in the header: macOS decides residency per set, so when one set loses residency under GPU memory pressure only the allocations in that set must be re-made resident. At most 32 sets exist (`kMaxSets`) because a command queue accepts a limited number of sets and every set is attached to every queue. At the cap, allocations go to the emptiest set and it grows past the cap. [source]
- Unlike llama.cpp, MLX attaches the sets to each command queue (`attach_new_sets`) before every command-buffer commit, because Metal locks a command buffer's residency at commit time. [source]
- `MLX_RESIDENCY_DEBUG=1` prints `[residency] created residency set N (max_bytes_per_set=... MB)` to stderr. This is the MLX counterpart of llama.cpp's startup log. [source]
- Budget logic: allocations beyond the wired limit are tracked but left out of any set; `set_wired_limit` (via `resize`) adds the ones that now fit, or removes allocations until under the new limit, then commits each touched set once. [source]
- Each allocation insert and erase commits its set. This is the "add/remove plus commit traffic" that the panic dossier blames, so more sets means each commit touches a smaller set. [source]
- MLX's memory limit (separate knob): default `min(1.5 x max_recommended_working_set_size, 0.95 x physical memory)`; the allocator's GC limit is `min(0.95 x max_recommended, block_limit)`. Past the memory limit with no free RAM and swap, allocations raise an exception. The set_memory_limit doc says the default is 1.5 times the recommended size. [source]
- Awni Hannun (Jan 2025): "In MLX we use residency sets to keep memory wired when needed, though we usually allow the system to unwire it between queries to the LM so that other applications can use the RAM." This is the opposite default to llama.cpp's keep-warm design. [source]
- mlx-lm: `wired_limit(model)` is a context manager that calls `mx.set_wired_limit(max_recommended_working_set_size)`, warns if model bytes > 0.9 x recommended (`[WARNING] Generating with a model that requires N MB which is close to the maximum recommended size ...`), and on exit synchronizes the streams then restores the old limit (0). So a one-shot `generate()` wires only for the call. `BatchGenerator` and the server wire once and keep it (the BatchGenerator restores in `close()`; `server.py` calls `maybe_set_recommended_wired_limit()` once at startup and never restores). A synchronize before restoring is required because "the wired limit should not be changed during an async eval". [source]
- Idle behavior differs by engine. llama.cpp (macOS 15+): wired for 180 s after the last compute, then eligible for unwiring; next request pays the re-wire cost, which scales with model size (the 250-470 ms per decode seen in llama-idle). mlx_lm one-shot: unwired as soon as the call ends. mlx_lm.server: wired until exit. [source]
- Mmap'd weights: residency requests wire the pages. After the heartbeat stops and macOS unwires, the pages are still in the page cache but not wired; the cost is re-wiring, not re-reading from disk (inference, [asserted]). [source]
- Quick check that a process still holds wired memory: `top -l 1 | grep PhysMem` prints `PhysMem: 44G used (16G wired, ...)` (used in issue 25937, where wired fell from 16G to 3000M only after exit). [source]
- The 25937 bug is Metal-wide: reproduced outside llama.cpp with a minimal Objective-C program (`requestResidency`, `endResidency`, release, no GPU work); any GPU operation anywhere in the process before or after fixes it. Apple Feedback FB23959296. Fixed in llama.cpp by PR 26082 (the issue shows it as the closing PR). [source]
- `iogpu.disable_wired_collector=1` is a system-wide switch that stops the unwiring. The panic dossier already records it did NOT prevent the IOGPUFamily panic. As an idle-latency fix it works but needs root and macOS 15+; residency sets are the per-process alternative. [source]
- Setting `GGML_METAL_NO_RESIDENCY` also makes the 25937 leak invisible, and re-introduces the ~1 s unwire cost. [source]
- A server wired by a long keep-alive also holds that memory against other apps, which MLX avoids by design. [source]
- llama.cpp keeps memory warm (180 s heartbeat, wired indefinitely during load) while MLX by default lets the OS unwire between queries and only wires while generating. Neither side calls the other wrong; the trade is first-token latency versus RAM given back to other apps. [source]
- Cause of the idle slowdown: ggerganov initially attributed it to GPU power saving; the MLX author and later a Metal engineer comment attributed it to unwiring. The sysctl test supports unwiring. [source]
- Whether residency-set wiring counts against `iogpu.wired_limit_mb` or against the `recommendedMaxWorkingSetSize` budget is not stated in any fetched source. [source]
- Whether llama.cpp's heartbeat polling every 5 ms forever has a measurable idle power cost on laptops was not measured in any source. [source]
- What exactly the "GPU wired memory collector" timeout is (1 s versus a counter) is only a Metal engineer's recollection; Apple documents neither the collector nor `iogpu.disable_wired_collector`. [source]
- Whether `-ngl`/fit logic uses `recommendedMaxWorkingSetSize` on Mac is still not verified (carried from the existing dossier). [source]
- Residency sets exist on macOS 15.0+, iOS 18+, tvOS 18+, visionOS 2+ and are described by Apple as groups of allocations that can move in and out of resident memory. [source]
- Apple: Metal makes the union of all residency sets' allocations resident, so removing an allocation from one set does not unmake it resident if it is in another set. [source]
- Apple: residency sets do not track hazards, so apps must use fences and events; adding allocations to a set costs less CPU than per-encoder `useResource` calls. [source]
- Apple: Metal attaches all of a command queue's residency sets to a command buffer when it is committed. [source]
- llama-idle (PR 10119) on M2 Ultra, Llama-3.1-8B F16: decode 29 ms with pauses up to 1000 ms, 227 ms at 1200 ms, 413-472 ms from 1400 ms. [source]
- llama-idle Q8_0: 19 ms up to 1000 ms pause, 106 ms at 1200 ms, 225-299 ms from 1400 ms; Q4_0: 14 ms up to 800 ms, 24 ms at 1000 ms, 86 ms at 1200 ms, 158-216 ms from 1400 ms. [source]
- The idle penalty grows with model size, appears on M1 Pro and M2 Ultra, only on GPU, not on CUDA. [source]
- ggerganov first proposed a heartbeat kernel every 0.9 s as a workaround, then found the cause was memory unwiring. [source]
- `sudo sysctl iogpu.disable_wired_collector=1` (macOS 15+) removed the idle slowdown in that test. [source]
- Awni Hannun: re-wiring memory is relatively expensive and scales with model size. [source]
- Awni Hannun: MLX uses residency sets to keep memory wired when needed but usually lets the system unwire it between LM queries so other apps can use RAM. [source]
- A Metal engineer's comment in PyTorch issue 124056 says memory wired for the processed command buffer is released after a sleep (about 1 s or a fixed counter) and mapping it back costs time. [source]
- llama.cpp PR 11427 (merged Jan 26, 2025) added residency sets; on M2 Ultra a 7B Q8_0 model's requests were about 250 ms faster. [source]
- PR 11427 benchmarks on 3B F16/Q4_0/Q8_0 show pp512 and tg128 speedup 1.00, so the gain is idle-resume latency only. [source]
- PR 11427 says that without residency sets the OS collects GPU memory after 1 second of inactivity. [source]
- PR 11427 attaches no residency sets to the command queue or buffers; it creates one set per buffer, adds the MTLBuffers, commits and requests residency. [source]
- llama.cpp creates one MTLResidencySet per ggml_metal_buffer, labelled "ggml_metal", and calls `endResidency`, `removeAllAllocations`, `commit` on free. [source]
- llama.cpp's heartbeat thread polls every 5 ms for the life of the device; only the `requestResidency` calls stop once the keep-alive counter (initial 36,000 loops = 180 s x 200/s) reaches 0. [source]
- `GGML_METAL_RESIDENCY_KEEP_ALIVE_S` is parsed with `atoi`; 0, negative or non-numeric values fall back to 180 s. [source]
- llama.cpp logs `creating a residency set collection (keep_alive = N s)` at init. [source]
- `ggml_metal_rsets_free` asserts the residency set collection is empty at device teardown. [source]
- Mmap'd model data is wrapped with `newBufferWithBytesNoCopy` in `MTLResourceStorageModeShared`, page-aligned; data above `maxBufferLength` is split into overlapping views. [source]
- Overlap between views is the largest tensor rounded up plus two pages, so every tensor fits entirely in one view. [source]
- `max_working_set_size` is `recommendedMaxWorkingSetSize` on macOS 10.12+ and `maxBufferLength` otherwise. [source]
- Issue 25937 reproduction measured `top -l 1 | grep PhysMem` showing 16G wired before exit and 3000M after, with no GPU work run. [source]
- The 25937 bug reproduces outside llama.cpp: create a residency set, `requestResidency`, `endResidency`, release, with no GPU operation; any GPU op anywhere in the process frees the memory. [source]
- MLX residency sets are enabled only on devices supporting MTLGPUFamilyMetal3 and macOS 15+; otherwise every call is a no-op. [source]
- MLX spreads wired allocations over several residency sets capped at `MLX_RESIDENCY_SET_MAX_PCT` percent (default 5) of `recommendedMaxWorkingSetSize`, with a 64 MiB floor; <= 0 or >= 100 means a single set. [source]
- MLX allows at most 32 residency sets because each set is attached to every command queue and a queue accepts a limited number. [source]
- MLX rationale for multiple sets: macOS makes residency decisions per set, so a set that loses residency under GPU memory pressure only forces re-residency of its own allocations. [source]
- `MLX_RESIDENCY_DEBUG=1` prints a line to stderr each time a residency set is created. [source]
- MLX attaches new residency sets to each command queue before every command-buffer commit because Metal locks a command buffer's residency at commit time. [source]
- MLX leaves allocations that exceed the wired limit tracked but out of any set; raising the limit later adds them, lowering it removes them and commits each touched set once. [source]
- MLX commits a residency set on every single allocation insert or erase. [source]
- MLX's allocator block limit is `min(1.5 x max_recommended_working_set_size, 0.95 x memory size)` and its GC limit is `min(0.95 x max_recommended, block_limit)`. [source]
- `mx.set_memory_limit` doc: default is 1.5 times the recommended working set size when Metal is available; past it with no RAM or swap left, allocations raise an exception. [source]
- mlx-lm `wired_limit(model)` sets the wired limit to the recommended size, warns if model bytes exceed 0.9 x recommended, synchronizes streams on exit and restores the old limit. [source]
- mlx-lm's `BatchGenerator` sets the recommended wired limit at construction and restores it in `close()`; `wired_limit` docstring says not to change the limit during an async eval. [source]
- `mlx_lm.server` calls `maybe_set_recommended_wired_limit()` once at startup and does not restore it. [source]
- `maybe_set_recommended_wired_limit()` returns None when the device reports no `max_recommended_working_set_size`, otherwise the previous limit. [source]
- MLX_RESIDENCY_DEBUG, MLX_RESIDENCY_SET_MAX_PCT, MLX_MAX_OPS_PER_BUFFER and MLX_MAX_MB_PER_BUFFER are read once through `get_var` and cached in function-local statics. [source]
- Inference: the wake-up cost after unwiring is a re-wire of already-cached pages, not a disk re-read. [source]
- Inference: with default settings, an interactive llama.cpp user pays the idle penalty only after more than 3 minutes idle, while the same user without residency sets would pay it after about 1 second. [source]
Children
- No children recorded.