Mac local LLMs: Ollama internals
Parent: Running LLM models locally on a Mac · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
0.40.0-rc0 (pre-release, 2026-09-25): models on MLX-supported architectures run on MLX automatically on Apple silicon; GGUF tags may switch runner; no opt-out found; final 0.40.0 unknown (release URL 404 on 2026-10-04).
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Runner choice and diagnosis
- 0.40.0-rc0 (pre-release, 2026-09-25): models on MLX-supported architectures run on MLX automatically on Apple silicon; GGUF tags may switch runner; no opt-out found; final 0.40.0 unknown (release URL 404 on 2026-10-04). [source]
- Which runner is live: grep ~/.ollama/logs/server.log for `--mlx-engine`; with OLLAMA_DEBUG=1 look for `Parallel:`. MLX tags carry `-mlx` (qwen3.6:27b-mlx 19 GB, 35b-mlx 24 GB, 256K ctx); run `ollama run gemma4:12b-mlx`. [source]
- GGUF models go through an upstream llama-server subprocess on main (CGO engines removed 2026-05-29); 2025 "Ollama engine" claims describe old builds. [source]
MLX memory and concurrency (act on these)
- OLLAMA_FLASH_ATTENTION and OLLAMA_KV_CACHE_TYPE are read only in the non-MLX branch of server/sched.go: no effect on MLX models. [source]
- OLLAMA_NUM_PARALLEL has no effect on MLX (single slot; 5 concurrent requests spread 4.0-75.0 s). PR 17317 (continuous batching) open 2026-08-14. qwen35/qwen35moe log "model architecture does not currently support parallel requests" and force Parallel:1. [source]
- Issue 16698 (0.30.8, M4 Max 64 GB, qwen3.6:35b-mlx): RSS 24 to 75.2 GB, swap 31.6 GB, 131k request 58.7 to 13.35 tok/s; 0.30.7 and earlier unaffected. Cause: recurrent conv-state slices pinned the forward buffer. Fixed by PR 17077 (commit 4e96f4d), shipped by 0.32.11: upgrade. [source]
- Workaround before the fix: OLLAMA_KEEP_ALIVE=0 (46.61 tok/s standalone 131k) kills the runner after each request, dropping the prefix cache. [source]
- 32 GB hosts: 18131 (M1 Max, 0.32.14, qwen3.8:27b-mlx, num_ctx 73728) process 20.41 to 28.16 GB, swap up to 10.10 GB; MLX 16.58 vs GGUF+MTP 11.08 tok/s. Attributed to the hard-coded 8 GiB prefix cache; proposed OLLAMA_MLX_PREFIX_CACHE_BYTES not found in releases. 48 GB plateaus (17924: 28.5 GiB). [source]
Create, quantize, import
- `ollama create --experimental --quantize int4|int8|mxfp4|mxfp8|nvfp4` panicked "mlx: There is no Stream(gpu, 1) in current thread" from 0.23.1 (0.23.0 worked); issue 16070 closed 2026-07-07; `q4_K_M` rejected there. [source]
- 0.34.1: safetensors create no longer experimental; GGUF conversion/quantization needs llama.cpp tooling. Import: `FROM /path/to/safetensors/directory` then `ollama create`; Ollama does not quantize GGUF on import. [source]
Context length
mmap
- Metal rule ("mmap has issues with partial offloading on metal"): use_mmap off when 0 < NumGPU < block_count+1, so a split model cannot exceed RAM; no env var forces mmap. [source]
- Per request/Modelfile: `{"options":{"use_mmap":false}}` cured thrashing (3940; some crashes unless num_gpu lowered); `use_mmap:true` forces on. On main, false gives `--load-mode none`. OLLAMA_NO_MMAP (PR 6854) was not merged as far as seen. [source]
- Compat-translated GGUFs load fully into RAM; `OLLAMA_LLAMA_CPP_COMPAT=0` disables hooks. [source]
llama-server launcher and bundled llama.cpp
- Flags always set: `--model --port --host 127.0.0.1 --no-webui --offline -c NumCtx*numParallel -np`, `--log-verbosity 4`. `-ngl` only if num_gpu set; `-b/-ub` from NumBatch; `--flash-attn on|off|auto`; `--cache-type-k/v` from OLLAMA_KV_CACHE_TYPE (default f16); `--context-shift --keep N`; `-t`; `--split-mode none --main-gpu`. Other llama-server flags are unreachable. [source]
- Metal startup failure ("failed to initialize ggml backend device: metal", "failed to allocate context") retries with GGML_METAL_TENSOR_DISABLE=1. [source]
- OLLAMA_KV_ROTATE does not exist (0 issue/PR hits, absent from envconfig/config.go); do not recommend it. KV rotation, if any, is inherited from llama-server; LLAMA_ATTN_ROT_DISABLE opts out only if it reaches the subprocess. [source]
Renderer, parser, Modelfile (undocumented)
- PR 18786: GGUF import auto-sets renderer `qwen3.8` + parser `qwen3.5` for arch qwen35/qwen35moe only if the embedded chat template contains both `resolved_reasoning_effort` and `preserve_thinking`. Detection runs only when Renderer or Parser is empty; values set in the request or Modelfile win. [source]
- Modelfile accepts `RENDERER` and `PARSER`, absent from docs/modelfile.mdx. Setting either (or harmony, or a Go template) makes Ollama render the prompt and sets DisableJinja, so PARSER alone also bypasses the GGUF Jinja template. [source]
- Names do not pair 1:1: renderer `qwen3.8` has no parser (use `qwen3.5`); `deepseek3.1` renders, parser is `deepseek3`; `harmony`, `ministral`, `passthrough` are parser-only. Unknown renderer errors `unknown renderer %q`; unknown parser silently returns nil. [source]
Template selection (Go TEMPLATE vs GGUF Jinja)
- OLLAMA_GO_TEMPLATE ("Enable Modelfile TEMPLATE based rendering when available"; an unparseable value counts as true): 1 forces Go TEMPLATE rendering, 0 the native llama-server GGUF chat template. [source]
- Unset: Ollama prefers the GGUF template only if it has strictly more capabilities and the Go template lacks a tool round trip; never when a Renderer, Parser or harmony (gptoss with `<|start|>`/`<|end|>`) applies. Log "model is missing tokenizer.chat_template and Go TEMPLATE support is unavailable; chat responses may be poorly formatted": set OLLAMA_GO_TEMPLATE=1. Release that added PreferChatTemplate unknown. [source]
- Native path posts to 127.0.0.1:<port>/v1/chat/completions and maps think to enable_thinking/reasoning_effort; Go-rendered prompts add `--no-jinja --chat-template chatml`. [source]
Speculative decoding (MLX and GGUF)
- MLX picks DFlash if the draft implements BlockDraft (draft limit blockSize-1), else MTP (unbounded); depth is adaptive and the old OLLAMA_MLX_MTP_* variables were removed. gemma4:26b-mlx M5 Max 148-157 tok/s. [source]
Open questions
- Final 0.40.0 architecture list, MLX opt-out, and whether PR 17317 or a prefix-cache setting shipped; dense-model growth past 8 GiB unresolved; minimum chip for MLX unknown. [source]
Corrections and disagreements
- CONTRADICTS claude-code-context-window-mismatch-against-local-servers.md on thresholds: docs.ollama.com/context-length lists under 24 GiB as 4k, 24-48 GiB as 32k and 48 GiB or more as 256k, while that dossier cites 23 GiB and 47 GiB from commit 0334ffa6. [source]
Concepts in this cluster
- Ollama MLX backend and NVFP4 on Apple silicon [source]
- Ollama runner no-mmap default on Metal and use_mmap option [source]
- ollama create from safetensors with MLX quantization [source]
- Ollama default 262144 context and num_ctx memory cost on Mac [source]
- Ollama VRAM-tiered default num_ctx and v1 num_ctx dropping [source]
- Ollama KV rotation port (OLLAMA_KV_ROTATE) [source]
- Ollama engine vs llama engine loader split and mmap support [source]
- Ollama launcher flag set passed to llama-server [source]
- Ollama qwen3_5 MTP loader and SelfDraft path [source]
- llama.cpp bundled revision inside Ollama releases [source]
- Ollama DFlash block-diffusion draft model in mlxrunner [source]
- Ollama llama/compat patch layer for llama.cpp [source]
- Ollama mlxrunner speculative depth controller (committed tokens per second) [source]
- Ollama multimodal projector offload disable policy [source]
- Ollama GGUF import renderer detection (qwen3.8 renderer selection) [source]
- Ollama Modelfile RENDERER and PARSER directives and the renderer/parser registri [source]
- Ollama usesOllamaRenderedChat versus native llama-server Jinja chat path [source]
Children
- ollama create from safetensors with MLX quantization
- Ollama default 262144 context and num_ctx memory cost on Mac
- Ollama DFlash block-diffusion draft model in mlxrunner
- Ollama engine vs llama engine loader split and mmap support
- Ollama GGUF import renderer detection (qwen3.8 renderer selection)
- Ollama KV rotation port (OLLAMA_KV_ROTATE)
- Ollama launcher flag set passed to llama-server
- Ollama llama/compat patch layer for llama.cpp
- Ollama MLX backend and NVFP4 on Apple silicon
- Ollama mlxrunner speculative depth controller (committed tokens per second)
- Ollama Modelfile RENDERER and PARSER directives and the renderer/parser registri
- Ollama multimodal projector offload disable policy
- Ollama qwen3_5 MTP loader and SelfDraft path
- Ollama runner no-mmap default on Metal and use_mmap option
- Ollama usesOllamaRenderedChat versus native llama-server Jinja chat path
- Ollama VRAM-tiered default num_ctx and v1 num_ctx dropping
- llama.cpp bundled revision inside Ollama releases