<!-- llms-explorer concept facts · https://llms-explorer.com/tree/mlx-allocator-memory-limit-gc-limit-and-buffer-c/ · pack 2026-10-05 · ~1594 tokens -->

# MLX allocator memory limit GC limit and buffer cache

> Resource count cap: malloc throws "[metal::malloc] Resource limit (N) exceeded." when the number of live Metal resources reaches the device resource limit, after trying to reclaim cache.

Parent: [Mac local LLMs: MLX kernels, numerics and internals](https://llms-explorer.com/tree/mac-local-llms-mlx-kernels-numerics-and-internals/) · 2 facets · 26 facts · page: https://llms-explorer.com/tree/mlx-allocator-memory-limit-gc-limit-and-buffer-c/

## Facts

- Resource count cap: malloc throws "[metal::malloc] Resource limit (N) exceeded." when the number of live Metal resources reaches the device resource limit, after trying to reclaim cache. — source: `asserted`
- max buffer: sizes above device maxBufferLength throw a dedicated message. — source: `asserted`
- set_wired_limit throws if the limit exceeds max_recommended_working_set_size. — source: `asserted`
- Reported fragmentation: one user blames the buffer cache for a SIGABRT after about 14 hours (issue 1015), but this is a single unverified report. — source: `asserted`
- Docs for set_cache_limit say the cache limit defaults to the memory limit. In source, max_pool_size_ is initialized to block_limit_ at construction, and set_memory_limit changes block_limit_ and gc_limit_ but not max_pool_size_. After a set_memory_limit call the two diverge. Both sides stated as read; not averaged. — source: `asserted`
- Whether the issue 1015 fragmentation claim is real: MLX buffers are exact page-multiple sizes cached by size, so same-size reuse is the only reuse path. — source: `asserted`
- The device resource_limit value on Apple silicon (not in the fetched pages). — source: `asserted`
- MetalAllocator initializes block_limit_ = min(1.5 x max_recommended_working_set_size, 0.95 x memory_size), gc_limit_ = min(0.95 x max_recommended, block_limit_), and max_pool_size_ = block_limit_. — [source](https://raw.githubusercontent.com/ml-explore/mlx/main/mlx/backend/metal/allocator.cpp)
- set_memory_limit swaps block_limit_ and recomputes gc_limit_ = min(block_limit_, 0.95 x recommendedMaxWorkingSetSize) but leaves max_pool_size_ unchanged. — [source](https://raw.githubusercontent.com/ml-explore/mlx/main/mlx/backend/metal/allocator.cpp)
- set_cache_limit only swaps max_pool_size_ and returns the old value. — [source](https://raw.githubusercontent.com/ml-explore/mlx/main/mlx/backend/metal/allocator.cpp)
- The set_cache_limit doc says free memory is reclaimed from the cache on the next allocation and that a limit of 0 disables the cache. — [source](https://ml-explore.github.io/mlx/build/html/python/_autosummary/mlx.core.set_cache_limit.html)
- malloc rounds sizes above vm_page_size up to a whole page multiple before looking in the cache. — [source](https://raw.githubusercontent.com/ml-explore/mlx/main/mlx/backend/metal/allocator.cpp)
- On a cache miss malloc computes mem_required = active + cache + size and, if it is at least gc_limit_ or the resource count is at the limit, releases cached buffers worth mem_required - gc_limit_. — [source](https://raw.githubusercontent.com/ml-explore/mlx/main/mlx/backend/metal/allocator.cpp)
- Buffers smaller than 256 bytes are taken from a 1 MiB MTL heap (small_size_ = 256, heap_size_ = 1 << 20); the heap is skipped on the "Apple Paravirtual device" VM. — [source](https://raw.githubusercontent.com/ml-explore/mlx/main/mlx/backend/metal/allocator.h)
- malloc throws "[metal::malloc] Resource limit (N) exceeded." when num_resources_ >= resource_limit_ after reclaiming. — [source](https://raw.githubusercontent.com/ml-explore/mlx/main/mlx/backend/metal/allocator.cpp)
- malloc throws a dedicated error when a request exceeds device maxBufferLength. — [source](https://raw.githubusercontent.com/ml-explore/mlx/main/mlx/backend/metal/allocator.cpp)
- free() recycles the buffer into the cache only while cache memory is below max_pool_size_; otherwise it erases the buffer from the residency sets and releases it. — [source](https://raw.githubusercontent.com/ml-explore/mlx/main/mlx/backend/metal/allocator.cpp)
- Every non-heap buffer is inserted into the residency set on creation and erased on release, so cached buffers stay resident until released. — [source](https://raw.githubusercontent.com/ml-explore/mlx/main/mlx/backend/metal/allocator.cpp)
- make_buffer wraps existing memory with device newBuffer(ptr, size, options, nullptr), inserts it into the residency set and counts it in active memory; release() frees it without caching. — [source](https://raw.githubusercontent.com/ml-explore/mlx/main/mlx/backend/metal/allocator.cpp)
- All MLX Metal buffers use shared storage mode with untracked hazard tracking. — [source](https://raw.githubusercontent.com/ml-explore/mlx/main/mlx/backend/metal/allocator.cpp)
- set_wired_limit throws std::invalid_argument when the limit exceeds max_recommended_working_set_size. — [source](https://raw.githubusercontent.com/ml-explore/mlx/main/mlx/backend/metal/allocator.cpp)
- The global allocator is heap-allocated and never destroyed, so cached buffers are leaked at exit on purpose to save exit time. — [source](https://raw.githubusercontent.com/ml-explore/mlx/main/mlx/backend/metal/allocator.cpp)
- mlx-lm issue 1015 (opened 2026-03-17) reports generate() terminating the process with "[METAL] Command buffer execution failed: Insufficient Memory (kIOGPUCommandBufferCallbackErrorOutOfMemory)" and no recovery path. — [source](https://github.com/ml-explore/mlx-lm/issues/1015)
- The 1015 reporter attributes a SIGABRT after about 14 hours at 50 requests/hour on Llama-3.1-8B-Instruct-4bit to cache fragmentation, and works around it with mx.clear_cache(), gc.collect(), set_memory_limit/set_cache_limit, and process recycling every 24 hours. — [source](https://github.com/ml-explore/mlx-lm/issues/1015)
- Because buffers are page-multiple sized and reused only by size, a long-running server with varied prompt lengths can hold many unreusable cached sizes until the cache limit or gc_limit_ forces release. — source: `asserted`

## Corrections and disagreements

- CONTRADICTS (docs vs source, no existing file): the set_cache_limit doc says the cache limit defaults to the memory limit; in source this holds only until set_memory_limit is called. — [source](https://ml-explore.github.io/mlx/build/html/python/_autosummary/mlx.core.set_cache_limit.html)
