llama.cpp bundled revision inside Ollama releases
Parent: Mac local LLMs: Ollama internals · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
File `LLAMA_CPP_VERSION` at the Ollama repository root holds a single llama.cpp build tag such as `b11232`.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- File `LLAMA_CPP_VERSION` at the Ollama repository root holds a single llama.cpp build tag such as `b11232`. [source]
- `llama/server/CMakeLists.txt` reads that file, fetches the tagged llama.cpp source with FetchContent, and applies the patches in `llama/compat/` (model shims, clip handlers, hooks). The Dockerfile reads the same file. [source]
- A developer can point the build at a local llama.cpp tree with `OLLAMA_LLAMA_CPP_SOURCE`, and skip the patch with `OLLAMA_LLAMA_CPP_SKIP_COMPAT_PATCH`. [source]
- Ollama's `llama/README.md` lists what a bump must review: build options, backend discovery symbols, llama-server launch args and log lines Ollama parses, streaming frames, and "speculative/MTP paths". [source]
- A user can read the bundled build from any install with `<runner>/llama-server --version`; the Homebrew 0.34.4 runner prints `version: 0.4.1-dev (build 11081, commit 161755f29)`. [source]
- 2026-05-29: Ollama removes its CGO engines and uses llama-server exclusively for GGML models. [source]
- Bumps visible in the file history (all by the same maintainer in the commits read): b9493 (06-03), b9509 (06-04), b9637 (06-15), b9672 (06-17), b9888 (07-06), 08-04, 08-10, 08-12, 08-15, 08-18, b10630 (08-26), b10729 (09-01), b10760 (09-02), b10864 (09-10), b10969 (09-15), an unnumbered "version update" (09-22), b11232 (09-29). That is 17 bumps in about four months, roughly weekly since August. [source]
- Patch churn rides along: b10729 regenerated the compat hooks patch because upstream removed a whole-tensor `load_data_for` read; b10969 moved the compat patch into libllama because libllama and libmtmd had duplicate symbols; the 09-22 bump removed the Laguna Metal patch as fixed upstream. [source]
- A feature merged into llama.cpp master is not in a given Ollama release until a bump after the merge ships. PR 25592 (hybrid checkpoint restore) is still unmerged upstream, so no Ollama release has it. [source]
- The Ollama server's own prompt cache and the MLX engine do not depend on this pin. [source]
- The pinned llama-server is built without router support (see the router dossier), so llama.cpp's `--models-max` features are unavailable through Ollama. [source]
- Whether the 09-22 "version update" is exactly b11081; the tags v0.34.4 and v0.35.0 both pin b11081, so it very likely is. [source]
- Which llama.cpp build the 0.40.0 release will pin; the pre-release tag v0.40.0-rc0 pins b11081. [source]
- The Ollama repository root has a file `LLAMA_CPP_VERSION` containing `b11232` on main on 2026-10-04. [source]
- `llama/server/CMakeLists.txt` reads the pinned commit from `../../LLAMA_CPP_VERSION` ("shared with Dockerfile") and fetches llama.cpp with FetchContent. [source]
- The same CMake file honors the environment variable `OLLAMA_LLAMA_CPP_SOURCE` for a local source tree and the option `OLLAMA_LLAMA_CPP_SKIP_COMPAT_PATCH`. [source]
- The root CMake file says GGML inference is provided by llama-server, built separately from "the pinned llama.cpp source". [source]
- Release tag v0.33.3 pins b10760. [source]
- Release tag v0.34.0 pins b10760. [source]
- Release tag v0.34.4 pins b11081. [source]
- Release tag v0.35.0 pins b11081. [source]
- Release tag v0.35.1 pins b11232. [source]
- Pre-release tag v0.40.0-rc0 pins b11081. [source]
- The v0.35.1 release notes list "Updated llama.cpp and the MLX engine". [source]
- The commit "llama.cpp: version bump b11232" (PR 18652) landed on 2026-09-29. [source]
- Earlier bumps in the file history: b10969 (2026-09-15, PR 18446), b10864 (09-10, PR 18317), b10760 (09-02, PR 18199), b10729 (09-01, PR 18160), b10630 (08-26, PR 18003). [source]
- The 2026-09-22 bump (PR 18577) refined memory-allocation failure log substrings and removed the Laguna Metal patch because upstream fixed it. [source]
- The b10729 bump regenerated the compat hooks patch because upstream removed the whole-tensor `load_data_for` read, whose last consumer was llama-quantize. [source]
- The b10969 bump moved the compat patch into libllama with exported symbols because upstream build changes created duplicate symbols between libllama and libmtmd. [source]
- The 2026-06-04 bump to b9509 brought "upstream Gemma 4 12B multimodal projector fixes for the n_head=0 divide-by-zero crash" on x86, CUDA, Linux and Windows. [source]
- The 2026-05-29 commit "runner: Remove CGO engines, use llama-server exclusively for GGML models" is the oldest entry read in the history page. [source]
- `llama/README.md` says `LLAMA_CPP_VERSION` pins Ollama's llama.cpp source and that an update can affect model loading, GPU discovery, scheduler inputs, runtime logs, streaming and compatibility patches. [source]
- `llama/README.md` tells reviewers of a bump to check "speculative/MTP paths", launch args consumed by `llm/llama_server.go`, and logs for device discovery, offload and memory accounting. [source]
- The `llama/` directory on main contains `clef`, `compat` and `server` subdirectories and a README. [source]
- Measured on this Mac: the Homebrew Ollama 0.34.4 runner `llama-server --version` prints `version: 0.4.1-dev (build 11081, commit 161755f29)` built with AppleClang 21.0.0.21000334 for Darwin arm64, matching the tag's pin. [source]
- The Homebrew client warns "client version is 0.35.1" while the server is 0.34.4, so the tag pin for a running server is the server's, not the CLI's. [source]
- Inference: lazy-mode (llama.cpp PR 27794, merged 2026-08-27) is in b10729 and later pins, so every Ollama release from v0.34.0 includes it, while v0.33.3 shares the b10760 pin and therefore does too. [source]
- Inference: llama.cpp PR 28335 (Gemma 4 vision attention, merged 2026-09-04) is absent from b10760 and present from b11081. [source]
- Inference: PR 28751 (merged 2026-09-28) is absent from b11081 and, depending on merge time, may be in b11232. [source]
Children
- No children recorded.