Ollama engine vs llama engine loader split and mmap support
Parent: Mac local LLMs: Ollama internals · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
The `NewLlamaServer` doc comment on main reads "All GGUF models are served via the upstream llama-server subprocess" and the function logs "using llama-server for model".
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- The `NewLlamaServer` doc comment on main reads "All GGUF models are served via the upstream llama-server subprocess" and the function logs "using llama-server for model". [source]
- `LlamaServerConfig` carries `DisableJinja`, `ContextShift`, `EnableMTP`, `ManifestDigest`, `DraftModelPath` and `DraftModelShardPaths`. [source]
- The `llm/` directory on main lists `llama_server.go`, `llama_binary.go`, `llama_server_score.go`, `metal_retry.go`, `model_split.go`, `gbnf.go` and `media.go`. [source]
- `runner/ollamarunner/runner.go`, `ml/backend/ggml/ggml.go` and `fs/ggml/ggml.go` return HTTP 404 on the main branch on 2026-10-04. [source]
- Ollama starts llama-server with `--model`, `--port`, `--host 127.0.0.1`, `--no-webui`, `--offline`, `-c NumCtx*numParallel` and `-np numParallel`, with the comment "minimal set, let llama-server auto-detect the rest". [source]
- `-ngl` is passed only when `NumGPU` is above 0 (value) or exactly 0 (`-ngl 0`); the default -1 passes nothing so llama-server auto-detects offload. [source]
- `appendLoadModeArgs` adds `--load-mode dio` for Linux integrated CUDA or ROCm GPUs and `--load-mode none` when `use_mmap` is explicitly false, and nothing otherwise. [source]
- llama-server's `-lm, --load-mode` default is `auto`: mmap unless a device does not support it; the other modes include `none`, `mmap`, `mlock` and `mmap+mlock`. [source]
- llama-server's `-lzm, --lazy-mode` reads certain large tensors on demand and requires mmap; `auto` applies it only to tensors above 4 GiB. [source]
- Ollama's memory accounting parses llama-server logs for `_Mapped` buffers and tracks `memModelFileBacked` and `memCPUMappedModel`, trimming overlap from the mmap-backed (reclaimable page cache) portion. [source]
- Flash attention is passed as `--flash-attn on` or `off` when `OLLAMA_FLASH_ATTENTION` is set and as `auto` when unset and the devices support it, and K/V cache type is passed as `--cache-type-k` and `--cache-type-v` from `OLLAMA_KV_CACHE_TYPE`. [source]
- `ShouldRetryWithMetalTensorDisabled` retries on darwin when startup errors match strings such as "failed to initialize ggml backend device: metal", "failed to allocate context" or "input types must match cooperative tensor types", using `GGML_METAL_TENSOR_DISABLE=1`. [source]
- A `DisableJinja` config launches llama-server with `--no-jinja --chat-template chatml` as a startup-only placeholder, because Go-rendered prompts go through completion endpoints. [source]
- LoRA adapters are logged as deprecated ("will be removed in a future release") and passed with `--lora`. [source]
- The release page URL `releases/tag/v0.40.0` returns a GitHub 404 on 2026-10-04; the pre-release tag is `v0.40.0-rc0`, whose compare range starts at v0.34.4 per the existing MLX-rollout dossier. [source]
- Inferred: on main a Mac user who needs mmap off sets `use_mmap: false` per request or Modelfile and gets `--load-mode none`; there is still no environment variable for it. [source]
- Inferred: the older "Ollama engine" claims in blog posts and issues from 2025 describe pre-removal builds and should not be applied to main. [source]
Children
- No children recorded.