llama.cpp PR 25309 sparse mmap tensor range coalescing
Parent: Mac local LLMs: Memory and wired limits · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Diff is 3 files (`src/llama-model-loader.cpp` +3/-6, `src/llama-model-loader.h` +31/-1, `src/llama-model.cpp` +43/-10), 1 commit (b970f06), +77/-17 total.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Diff is 3 files (`src/llama-model-loader.cpp` +3/-6, `src/llama-model-loader.h` +31/-1, `src/llama-model.cpp` +43/-10), 1 commit (b970f06), +77/-17 total. [source]
- `llama_buf_map` changes from `unordered_map<idx, buf>` to `vector<llama_buf_range{idx, first, last, buf}>`. `llama_buf_map_find(bufs, idx, offs, size)` returns the buffer whose range fully contains `[offs, offs+size)` and guards against `offs+size` overflow. Non-mmap allocation registers one catch-all range (0 to SIZE_MAX) per file. Async-upload code uses `llama_buf_map_first(bufs)` instead of `bufs.at(0)`. [source]
- The gap is measured between tensors of this backend context only (spans come from `ggml_get_first_tensor(ctx)` walk filtered by `weight->idx`), so CPU-resident tensors never enter a range unless they sit within 64 MiB of a GPU tensor. A 27 GiB CPU tensor between GPU tensors therefore splits the span. [source]
- New log lines: `mapping N sparse mmap ranges for <buft> file <idx> instead of one X GiB span` and, per range, `mapped <buft> mmap range file <idx> offset X GiB size Y MiB`. These are the way to verify the patch took effect (expect several `MTL0_Mapped` buffers whose sum drops by the gap size). [source]
- `mmap_coalesce_gap` is a hard-coded `constexpr` (64 MiB), not a parameter or env var. [source]
- Opened 2026-07-04 by Takinggg as a draft; last updated 2026-08-31; one commit; never rebased in the data seen (mergeable_state clean). [source]
- Author validation is only a local build (`cmake --build build --target llama-cli -j 8`); no before/after buffer sizes, no test, no benchmark are in the PR text. All measured numbers on the span effect come from issue 29465 (file reorder), not from this patch. [source]
- Xjs (M5 Max 128 GB, 2026-08-31) is the only runtime report: patch on unsloth's llama.cpp fit Qwen3.8-Flash-Next Q4_K_XL at full context. Anecdotal, no sizes given. [source]
- A range still has to obey the Metal max-buffer view splitting described in the existing dossier; coalescing reduces the bytes mapped, not the per-buffer limit. [source]
- Tensors within 64 MiB of GPU tensors still get wired; the saving is bounded to gaps above 64 MiB. [source]
- Existing dossier says the PR is "awaiting review". Corrected: GitHub shows `draft: true`, so code owners CISC and ggerganov are listed as "will be requested when the pull request is marked ready for review", and 2 approvals are required. No maintainer has commented. It is stalled by draft status, not by review. [source]
- Alternative fixes (not averaged): loader-side range splitting (25309, and the 29465 suggestion to skip CPU/lazy tensors), file-side reorder, `--no-mmap`. [source]
- Whether the author marks it ready; whether the hard-coded 64 MiB survives review. [source]
- Whether the Reddit repack script for Qwen3.8-Flash-Next (linked in 29465 by nazeshinjite, r/LocalLLM 1vz927j) is retrievable; the fetch helper refuses Reddit, so its text and any published hash were not read. [source]
- PR 25309 "loader: map sparse mmap tensor ranges" is open and a draft (merged=false, draft=true), created 2026-07-04, updated 2026-08-31, 1 commit, +77/-17 across 3 files, 1 comment, 0 review comments, mergeable_state clean. [source]
- PR 25309 is not merged, so no llama.cpp release tag contains it as of 2026-10-04 (latest tag seen b11344/b11382). [source]
- CISC and ggerganov (code owners) review will be requested only when PR 25309 is marked ready; two approving reviews are required. [source]
- PR 25309 sorts per-file backend tensor spans by offset and merges when the next span starts no more than 64 MiB (`mmap_coalesce_gap`, hard-coded) after the previous range end. [source]
- Gaps are computed only over tensors of the current backend context, so CPU-resident interior tensors larger than 64 MiB are excluded from the Metal ranges. [source]
- `llama_buf_map` becomes a vector of `{idx, first, last, buf}` and tensor loads resolve their buffer via `llama_buf_map_find(bufs, idx, offs, size)` requiring full containment. [source]
- The PR logs `mapping N sparse mmap ranges ... instead of one X GiB span` and one `mapped ... mmap range file I offset X GiB size Y MiB` line per range. [source]
- The PR author's only stated validation is a local macOS/arm64 Metal build of `llama-cli`; no measured memory numbers are given. [source]
- No measurement of the PR's effect on Metal span size exists in sources read; the quantitative span effect (27,466 MiB, 106,166 to 78,701 MiB MTL0_Mapped for unsloth Qwen3.8-Flash-Next UD-Q4_K_XL; UD-IQ4_XS 89,332 MiB as shipped) comes from moving the tensor in the file, reported in issue 29465. [source]
- Issue 29465 (open, 2026-09-26, no labels, no linked PR) notes the 27,466 MiB difference equals the `per_layer_token_embd.weight` size exactly, and that `--lazy-mode off` does not help because the table becomes resident and stays in the Metal range. [source]
- The 29465 reporter verified the reordered GGUF was bit-identical per tensor by sha256 and produced it with a plain gguf-py rewrite (no requantisation); the script text was not posted in the issue, a commenter pointed to a Reddit repack script (r/LocalLLM 1vz927j) for that exact quant. [source]
- No gguf-py reorder script text or published hash list was retrievable in this run (Reddit page refused by the fetch helper). [source]
- A hash-verified reorder procedure that generalises: write all tensors via `GGUFWriter` with the large CPU-only tensor last, then compare per-tensor sha256 of data between source and output, and compare KV metadata. [source]
- PR 22941 (MTP head Metal buffer duplication) was closed unmerged on 2026-05-11; the duplication pattern (two identical 18760.13 MiB MTL0_Mapped lines) is the same span effect, and 29465 shows the MTP head loaded separately fails with `failed to load draft model` under memory pressure. [source]
- PR 26082 "metal: fix memory leak if model is freed without any GPU operations" merged 2026-07-30 08:11 UTC as commit d0bfb1981266c271cd0536a8aa7c5e863e7cdf61 by ggerganov, 5 commits, +94/-3 across 4 files, labelled merge ready, 29 of 34 checks passed. [source]
- PR 26082 has two halves: a dummy GPU job (1-byte private buffer fill) at `ggml_metal_rsets_init`, later gated to `#if defined(GGML_METAL_HAS_RESIDENCY_SETS)`, and an added `[rset commit]` after `removeAllAllocations` in `ggml_metal_buffer_rset_free`. [source]
- PR 26082's regression test `tests/test-memory-release.cpp` first measured `phys_footprint`, which drops after free even without the fix, so it was reworked to system-wide wired memory with non-mmap load and disabled by default (needs a multi-GB model); on M5 Max it passed always, on M2 Ultra about 50% (macOS 26.5.2). [source]
- Per-process counters (`ri_wired_size`, `ri_resident_size`, `currentAllocatedSize`, `vmmap`, `footprint`) all drop without the fix, so the leak is observable only through system-wide wired memory (`top -l 1 | grep PhysMem`). [source]
- Commit d0bfb19 is contained in master and tags from b10188 upward (earliest tag in the branch_commits list), so any build at or after b10188 includes the 26082 fix. [source]
Children
- No children recorded.