<!-- llms-explorer concept facts · https://llms-explorer.com/tree/gpu-written-no-copy-buffer-footprint-charging/ · pack 2026-10-05 · ~811 tokens -->

# GPU-written no-copy buffer footprint charging

> Buffers MLX allocates (not no-copy wraps) and fills by GPU eval are charged to phys_footprint in full. A 60-iteration loop of ~500 MB alloc, eval, drop left 60.19 GB footprint, which matched active + cache within 0.14 GB.

Parent: [Mac local LLMs: Memory and wired limits](https://llms-explorer.com/tree/mac-local-llms-memory-and-wired-limits/) · 2 facets · 13 facts · page: https://llms-explorer.com/tree/gpu-written-no-copy-buffer-footprint-charging/

## Facts

- Buffers MLX allocates (not no-copy wraps) and fills by GPU eval are charged to phys_footprint in full. A 60-iteration loop of ~500 MB alloc, eval, drop left 60.19 GB footprint, which matched active + cache within 0.14 GB. — [source](https://github.com/ml-explore/mlx/issues/3896)
- Freed MLX buffers stay in the allocator pool, remain live Metal allocations and stay in phys_footprint; get_peak_memory tracks only active memory. — [source](https://github.com/ml-explore/mlx/issues/3896)
- clear_cache returns memory but phys_footprint trails the call by about 4 seconds (15.14 GB immediately, 0.02 GB at the next sample). — [source](https://github.com/ml-explore/mlx/issues/3896)
- A first reading right after clear_cache looked like a leak and was retracted by its author as a sampling artifact. — [source](https://github.com/ml-explore/mlx/issues/3896)
- Real incidents: a hard reboot and a Metal command-buffer GPU timeout on M5 Max 128 GB when the gate used get_peak_memory (46 GB reported, 110 GB real). — [source](https://github.com/ml-explore/mlx/issues/3896)
- A commenter reports a kernel panic (watchdogd 90 s timeout, 123.2 GiB resident, 50.1 GiB compressor) running Ollama 0.33.x MLX on an M5 Max 128 GB. — [source](https://github.com/ml-explore/mlx/issues/3896)
- GPU-written MLX allocator buffers are charged to phys_footprint as IOAccelerator (graphics) dirty memory. — [source](https://github.com/ml-explore/mlx/issues/3896)
- A 60 x 500 MB churn showed get_peak_memory 1.00 GB, active 0, cache 60.06 GB, footprint 60.19 GB. — [source](https://github.com/ml-explore/mlx/issues/3896)
- With cache_limit=0 the same churn ended at 1.14 GB footprint. — [source](https://github.com/ml-explore/mlx/issues/3896)
- The pool reuse window is [size, size + 2 pages), so varying request sizes grow the pool (cited as mlx issue 3886). — [source](https://github.com/ml-explore/mlx/issues/3896)
- The mlx maintainer's position: get_peak_memory is for single-model runs where cache hit rate is near 100 percent; for long-running or parallel serving use active + cache. — [source](https://github.com/ml-explore/mlx/issues/3896)
- After clear_cache, phys_footprint falls over a few seconds, not instantly. — [source](https://github.com/ml-explore/mlx/issues/3896)

## Corrections and disagreements

- CONTRADICTS (partly): iokit-mapped-metal-buffer-accounting-in-phys-foo.md measured a GPU-blit-filled untouched no-copy region at 0.12 GB footprint. Here GPU-written MLX-allocated buffers (not no-copy wraps) are charged in full as IOAccelerator dirty memory. The difference is plausibly allocation type (device allocation vs wrapped anonymous pages), but no source tests a GPU-written no-copy buffer directly. — source: `asserted`
