MLX buffer pool reuse window and cache sizing (mlx issue 3886)
Parent: Mac local LLMs: Memory and wired limits · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
The reuse test is `it->first >= min(2*size, size + 2*page_size_)` returns nullptr, so for buffers over about 32 KB the window is [size, size + 2 pages).
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- The reuse test is `it->first >= min(2*size, size + 2*page_size_)` returns nullptr, so for buffers over about 32 KB the window is [size, size + 2 pages). [source]
- Appending one position to 10 K/V pairs of [4, N, 512] bf16 per step cost 1.19/1.28/1.74/3.12 ms at ctx 512/1024/2048/4096 with growing concatenate, 0.58/0.82/1.13/1.96 ms with constant-size concatenate, and 0.38/0.36/0.35/0.35 ms with preallocated slice_update. [source]
- The reporter extrapolates growing concatenate on a 60-layer Gemma-4-31B to about 50 ms per token at 4096 context. [source]
- mlx-lm's Python KVCache avoids the trap by preallocating in 256-step chunks and using slice_update. [source]
- After a growing-cache phase, unrelated workloads ran 10 to 35 percent slower until process exit because the pool held never-reusable sizes. [source]
- The reporter's engine fix was chunk-preallocated buffers plus slice_update, where donation makes the append in place. [source]
- zcbenz closed issue 3886 on 2026-08-08 as won't fix, citing the preallocated-slice requirement of cuDNN SDPA. [source]
- Issue 3886 is a different issue from 3896 and the existing dossier's pointer to it is a second-hand citation. [source]
Children
- No children recorded.