MLX allocator memory limit GC limit and buffer cache
Parent: Mac local LLMs: MLX kernels, numerics and internals · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Resource count cap: malloc throws "[metal::malloc] Resource limit (N) exceeded." when the number of live Metal resources reaches the device resource limit, after trying to reclaim cache.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Resource count cap: malloc throws "[metal::malloc] Resource limit (N) exceeded." when the number of live Metal resources reaches the device resource limit, after trying to reclaim cache. [source]
- max buffer: sizes above device maxBufferLength throw a dedicated message. [source]
- set_wired_limit throws if the limit exceeds max_recommended_working_set_size. [source]
- Reported fragmentation: one user blames the buffer cache for a SIGABRT after about 14 hours (issue 1015), but this is a single unverified report. [source]
- Docs for set_cache_limit say the cache limit defaults to the memory limit. In source, max_pool_size_ is initialized to block_limit_ at construction, and set_memory_limit changes block_limit_ and gc_limit_ but not max_pool_size_. After a set_memory_limit call the two diverge. Both sides stated as read; not averaged. [source]
- Whether the issue 1015 fragmentation claim is real: MLX buffers are exact page-multiple sizes cached by size, so same-size reuse is the only reuse path. [source]
- The device resource_limit value on Apple silicon (not in the fetched pages). [source]
- MetalAllocator initializes block_limit_ = min(1.5 x max_recommended_working_set_size, 0.95 x memory_size), gc_limit_ = min(0.95 x max_recommended, block_limit_), and max_pool_size_ = block_limit_. [source]
- set_memory_limit swaps block_limit_ and recomputes gc_limit_ = min(block_limit_, 0.95 x recommendedMaxWorkingSetSize) but leaves max_pool_size_ unchanged. [source]
- set_cache_limit only swaps max_pool_size_ and returns the old value. [source]
- The set_cache_limit doc says free memory is reclaimed from the cache on the next allocation and that a limit of 0 disables the cache. [source]
- malloc rounds sizes above vm_page_size up to a whole page multiple before looking in the cache. [source]
- On a cache miss malloc computes mem_required = active + cache + size and, if it is at least gc_limit_ or the resource count is at the limit, releases cached buffers worth mem_required - gc_limit_. [source]
- Buffers smaller than 256 bytes are taken from a 1 MiB MTL heap (small_size_ = 256, heap_size_ = 1 << 20); the heap is skipped on the "Apple Paravirtual device" VM. [source]
- malloc throws "[metal::malloc] Resource limit (N) exceeded." when num_resources_ >= resource_limit_ after reclaiming. [source]
- malloc throws a dedicated error when a request exceeds device maxBufferLength. [source]
- free() recycles the buffer into the cache only while cache memory is below max_pool_size_; otherwise it erases the buffer from the residency sets and releases it. [source]
- Every non-heap buffer is inserted into the residency set on creation and erased on release, so cached buffers stay resident until released. [source]
- make_buffer wraps existing memory with device newBuffer(ptr, size, options, nullptr), inserts it into the residency set and counts it in active memory; release() frees it without caching. [source]
- All MLX Metal buffers use shared storage mode with untracked hazard tracking. [source]
- set_wired_limit throws std::invalid_argument when the limit exceeds max_recommended_working_set_size. [source]
- The global allocator is heap-allocated and never destroyed, so cached buffers are leaked at exit on purpose to save exit time. [source]
- mlx-lm issue 1015 (opened 2026-03-17) reports generate() terminating the process with "[METAL] Command buffer execution failed: Insufficient Memory (kIOGPUCommandBufferCallbackErrorOutOfMemory)" and no recovery path. [source]
- The 1015 reporter attributes a SIGABRT after about 14 hours at 50 requests/hour on Llama-3.1-8B-Instruct-4bit to cache fragmentation, and works around it with mx.clear_cache(), gc.collect(), set_memory_limit/set_cache_limit, and process recycling every 24 hours. [source]
- Because buffers are page-multiple sized and reused only by size, a long-running server with varied prompt lengths can hold many unreusable cached sizes until the cache limit or gc_limit_ forces release. [source]
Corrections and disagreements
- CONTRADICTS (docs vs source, no existing file): the set_cache_limit doc says the cache limit defaults to the memory limit; in source this holds only until set_memory_limit is called. [source]
Children
- No children recorded.