<!-- llms-explorer concept facts · https://llms-explorer.com/tree/ollama-engine-vs-llama-engine-loader-split-and-m/ · pack 2026-10-05 · ~1175 tokens -->

# Ollama engine vs llama engine loader split and mmap support

> The `NewLlamaServer` doc comment on main reads "All GGUF models are served via the upstream llama-server subprocess" and the function logs "using llama-server for model".

Parent: [Mac local LLMs: Ollama internals](https://llms-explorer.com/tree/mac-local-llms-ollama-internals/) · 1 facets · 17 facts · page: https://llms-explorer.com/tree/ollama-engine-vs-llama-engine-loader-split-and-m/

## Facts

- The `NewLlamaServer` doc comment on main reads "All GGUF models are served via the upstream llama-server subprocess" and the function logs "using llama-server for model". — [source](https://raw.githubusercontent.com/ollama/ollama/main/llm/server.go)
- `LlamaServerConfig` carries `DisableJinja`, `ContextShift`, `EnableMTP`, `ManifestDigest`, `DraftModelPath` and `DraftModelShardPaths`. — [source](https://raw.githubusercontent.com/ollama/ollama/main/llm/server.go)
- The `llm/` directory on main lists `llama_server.go`, `llama_binary.go`, `llama_server_score.go`, `metal_retry.go`, `model_split.go`, `gbnf.go` and `media.go`. — [source](https://github.com/ollama/ollama/tree/main/llm)
- `runner/ollamarunner/runner.go`, `ml/backend/ggml/ggml.go` and `fs/ggml/ggml.go` return HTTP 404 on the main branch on 2026-10-04. — [source](https://raw.githubusercontent.com/ollama/ollama/main/ml/backend/ggml/ggml.go)
- Ollama starts llama-server with `--model`, `--port`, `--host 127.0.0.1`, `--no-webui`, `--offline`, `-c NumCtx*numParallel` and `-np numParallel`, with the comment "minimal set, let llama-server auto-detect the rest". — [source](https://raw.githubusercontent.com/ollama/ollama/main/llm/llama_server.go)
- `-ngl` is passed only when `NumGPU` is above 0 (value) or exactly 0 (`-ngl 0`); the default -1 passes nothing so llama-server auto-detects offload. — [source](https://raw.githubusercontent.com/ollama/ollama/main/llm/llama_server.go)
- `appendLoadModeArgs` adds `--load-mode dio` for Linux integrated CUDA or ROCm GPUs and `--load-mode none` when `use_mmap` is explicitly false, and nothing otherwise. — [source](https://raw.githubusercontent.com/ollama/ollama/main/llm/llama_server.go)
- llama-server's `-lm, --load-mode` default is `auto`: mmap unless a device does not support it; the other modes include `none`, `mmap`, `mlock` and `mmap+mlock`. — [source](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md)
- llama-server's `-lzm, --lazy-mode` reads certain large tensors on demand and requires mmap; `auto` applies it only to tensors above 4 GiB. — [source](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md)
- Ollama's memory accounting parses llama-server logs for `_Mapped` buffers and tracks `memModelFileBacked` and `memCPUMappedModel`, trimming overlap from the mmap-backed (reclaimable page cache) portion. — [source](https://raw.githubusercontent.com/ollama/ollama/main/llm/llama_server.go)
- Flash attention is passed as `--flash-attn on` or `off` when `OLLAMA_FLASH_ATTENTION` is set and as `auto` when unset and the devices support it, and K/V cache type is passed as `--cache-type-k` and `--cache-type-v` from `OLLAMA_KV_CACHE_TYPE`. — [source](https://raw.githubusercontent.com/ollama/ollama/main/llm/llama_server.go)
- `ShouldRetryWithMetalTensorDisabled` retries on darwin when startup errors match strings such as "failed to initialize ggml backend device: metal", "failed to allocate context" or "input types must match cooperative tensor types", using `GGML_METAL_TENSOR_DISABLE=1`. — [source](https://raw.githubusercontent.com/ollama/ollama/main/llm/metal_retry.go)
- A `DisableJinja` config launches llama-server with `--no-jinja --chat-template chatml` as a startup-only placeholder, because Go-rendered prompts go through completion endpoints. — [source](https://raw.githubusercontent.com/ollama/ollama/main/llm/llama_server.go)
- LoRA adapters are logged as deprecated ("will be removed in a future release") and passed with `--lora`. — [source](https://raw.githubusercontent.com/ollama/ollama/main/llm/llama_server.go)
- The release page URL `releases/tag/v0.40.0` returns a GitHub 404 on 2026-10-04; the pre-release tag is `v0.40.0-rc0`, whose compare range starts at v0.34.4 per the existing MLX-rollout dossier. — [source](https://github.com/ollama/ollama/releases/tag/v0.40.0)
- Inferred: on main a Mac user who needs mmap off sets `use_mmap: false` per request or Modelfile and gets `--load-mode none`; there is still no environment variable for it. — source: `asserted`
- Inferred: the older "Ollama engine" claims in blog posts and issues from 2025 describe pre-removal builds and should not be applied to main. — source: `asserted`
