Ollama runner no-mmap default on Metal and use_mmap option
Parent: Mac local LLMs: Ollama internals · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
2024-06: maintainer advised {"options": {"use_mmap": false}} to cure thrashing (issue 3940).
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- 2024-06: maintainer advised {"options": {"use_mmap": false}} to cure thrashing (issue 3940). [source]
- 2024-09: PR 6854 proposed OLLAMA_NO_MMAP to force --no-mmap globally; no maintainer merge is visible on the page. [source]
- 2025-01 and 2025-04 logs show --no-mmap on Metal and Linux partial offload (issues 8571, 10104, 7081). [source]
- By v0.12.0 the loader is split into an Ollama engine (no mmap) and a llama engine (mmap rules above). [source]
- Layer count: the Metal rule triggers on NumGPU between 1 and block_count, so a model Ollama decides to split (for example 9 of 62 layers) always loads with no mmap, which cannot load a model larger than RAM. [source]
- Forcing mmap on is only possible per request or Modelfile (use_mmap true), not with an environment variable. [source]
- Disabling mmap crashed some models in a 2024 report even with fewer layers (issue 3940). [source]
- Reporters in issue 4895 say --no-mmap is slightly faster and avoids RAM spikes on 8 GB machines; the Ollama staff reply in PR 6854 asks whether mmap should instead be disabled by default. Not resolved on the page. [source]
- Whether PR 6854 was merged (page shows no merged marker; "Closed" appears twice, which I read as closed unmerged but did not verify). [source]
- Behavior in releases after v0.12.0; the main-branch llm/server.go fetched here is a 289-line stub with no mmap logic. [source]
- Ollama's code comment reads "mmap has issues with partial offloading on metal" and sets UseMMap to false when a metal GPU has 0 < NumGPU < BlockCount + 1. [source]
- In v0.6.2, --no-mmap is appended when (Windows, CUDA, UseMMap nil) or (Linux, free memory < estimated total, UseMMap nil) or (CPU library, UseMMap nil) or UseMMap explicitly false. [source]
- The same rules appear in v0.12.0 as loadRequest.UseMmap = false, and v0.12.0 notes "Mmap is only supported on the llama engine"; the Ollama engine path leaves the legacy UseMmap field unused. [source]
- Ollama's Options struct has UseMMap *bool with JSON name use_mmap, and the default options set UseMMap to nil. [source]
- A Metal log (Ollama.app, 2025-01-24) offloading 9 of 62 layers of a 413.6 GiB model on a 48 GiB machine shows --n-gpu-layers 9 --no-mmap. [source]
- A Linux log (Ollama 0.6.2, 2025-04-03) shows --n-gpu-layers 21 --no-mmap and "load_tensors ... (mmap = false)" for a 250 GB model. [source]
- A maintainer's workaround for forcing mmap is the client option {"use_mmap": true}; ollama serve --help and ollama run --help expose no mmap flag. [source]
- A 2024-06 maintainer reply tells a user whose model thrashed to send options {"use_mmap": false}, and the user reported it "works nicely"; a second user saw crashes with mmap disabled unless num_gpu was also reduced. [source]
- PR 6854 proposed OLLAMA_NO_MMAP=1 to always add --no-mmap to the llama runner; a reviewer asked what problem mmap caused and whether it should be off by default, and commenters later asked for updates and for a way to force mmap on. [source]
- Issue 4895 requested a global use_mmap environment variable because many front ends cannot set it; the reporter claims --no-mmap costs 5-10 s more load time on an 8 GB RAM, 6 GB VRAM PC but avoids RAM at 99 percent. [source]
- Issue 10539 (2025-05-02, Ollama 0.6.7, closed) asked for a global use_mmap variable after reporting that with use_mmap true, model RAM was not fully released after unload and load times were worse. [source]
- On a Mac, a model that fits fully on the GPU loads with mmap by default under the llama engine, so its weights are reclaimable file cache until Metal touches them; a split model loads into anonymous memory. [source]
Children
- No children recorded.