<!-- llms-explorer concept facts · https://llms-explorer.com/tree/llama-cpp-lazy-mode-lazy-tensor-reading-of-overs/ · pack 2026-10-05 · ~918 tokens -->

# llama.cpp lazy-mode lazy tensor reading of oversized embedding tables

> Small models lose throughput; large models barely notice.

Parent: [Mac local LLMs: Memory and wired limits](https://llms-explorer.com/tree/mac-local-llms-memory-and-wired-limits/) · 1 facets · 18 facts · page: https://llms-explorer.com/tree/llama-cpp-lazy-mode-lazy-tensor-reading-of-overs/

## Facts

- Small models lose throughput; large models barely notice. — source: `asserted`
- Streaming needs mmap; with mmap off there is nothing to read lazily. — source: `asserted`
- A tensor under 4 GiB stays resident in auto mode. — source: `asserted`
- Issue 29465 (existing dossier) shows the lazy table still sits inside the Metal-mapped span, so lazy reading does not shrink MTL0_Mapped. — source: `asserted`
- Flag name: the PR title and body say `--tensor-read-lazy on|off|auto`; existing dossiers cite `--lazy-mode on|auto|off` from the server README and modelfit. The PR page does not show a rename. Unresolved; check `llama-server --help` on the installed build. — source: `asserted`
- Whether `--lazy-mode` is a rename or alias of `--tensor-read-lazy`, and in which PR. — source: `asserted`
- Measured effect on a Mac with an unquantized 97.7 GiB n-gram table (only gemma-4 E4B numbers exist). — source: `asserted`
- PR 27794 states models with PLE and engram embeddings do not need the whole embedding table in RAM and can read it lazily via mmap. — [source](https://github.com/ggml-org/llama.cpp/pull/27794)
- The PR defines `--tensor-read-lazy on|off|auto`, where auto means lazy when the tensor is larger than 4 GiB. — [source](https://github.com/ggml-org/llama.cpp/pull/27794)
- The auto threshold is a constant auto_lazy_min_size = 4 GiB in the loader, and the condition is lazy ON or ggml_nbytes(cur) > threshold. — [source](https://github.com/ggml-org/llama.cpp/pull/27794)
- The PR author's test model was unsloth/gemma-4-E4B-it-GGUF:Q4_K_M, whose lazy tensor per_layer_token_embd.weight is Q5_K [10752, 262144], 1.94 GB, 39 percent of the 4.96 GB file. — [source](https://github.com/ggml-org/llama.cpp/pull/27794)
- Measured on that model: eager prefetch baseline 4.6 G of 4.6 G resident, peak RSS 7.37 GB, prompt processing 571-583 t/s, generation 105.8 t/s. — [source](https://github.com/ggml-org/llama.cpp/pull/27794)
- Lazy with prefetch skipped plus MADV_RANDOM: 2.8 G resident, peak RSS 6.16 GB, prompt processing 514-547 t/s, generation 94.2-94.7 t/s (minus 10.7 percent). — [source](https://github.com/ggml-org/llama.cpp/pull/27794)
- Lazy with prefetch skipped only: 2.8 G resident, prompt processing 561-563 t/s, generation 97.0-97.3 t/s (minus 8 percent). — [source](https://github.com/ggml-org/llama.cpp/pull/27794)
- The author says small models like gemma 4 take a significant performance hit because read delay is large relative to token generation, while larger models like qwen4 see a minor effect; this is why auto has a 4 GiB floor. — [source](https://github.com/ggml-org/llama.cpp/pull/27794)
- PR 27794 merged on 2026-08-27 as commit fac889fb38fd0e267636bd95bf096555e45b2270 with 23 of 26 checks passed. — [source](https://github.com/ggml-org/llama.cpp/pull/27794)
- A reviewer noted that qwen4 feeds n-gram embeddings at the second layer, not the first, which can hide the read latency of the lazy lookup. — [source](https://github.com/ggml-org/llama.cpp/pull/27794)
- A downstream fork's merge note says upstream's lazy path drops post-load advise passes and issues WILLNEED over the non-lazy complement, and that upstream's own commit measured MADV_RANDOM without batched row prefetch at 94.4 s. — [source](https://github.com/ggml-org/llama.cpp/pull/27794)
