<!-- llms-explorer concept facts · https://llms-explorer.com/tree/llama-cpp-bundled-revision-inside-ollama-release/ · pack 2026-10-05 · ~1982 tokens -->

# llama.cpp bundled revision inside Ollama releases

> File `LLAMA_CPP_VERSION` at the Ollama repository root holds a single llama.cpp build tag such as `b11232`.

Parent: [Mac local LLMs: Ollama internals](https://llms-explorer.com/tree/mac-local-llms-ollama-internals/) · 1 facets · 39 facts · page: https://llms-explorer.com/tree/llama-cpp-bundled-revision-inside-ollama-release/

## Facts

- File `LLAMA_CPP_VERSION` at the Ollama repository root holds a single llama.cpp build tag such as `b11232`. — source: `asserted`
- `llama/server/CMakeLists.txt` reads that file, fetches the tagged llama.cpp source with FetchContent, and applies the patches in `llama/compat/` (model shims, clip handlers, hooks). The Dockerfile reads the same file. — source: `asserted`
- A developer can point the build at a local llama.cpp tree with `OLLAMA_LLAMA_CPP_SOURCE`, and skip the patch with `OLLAMA_LLAMA_CPP_SKIP_COMPAT_PATCH`. — source: `asserted`
- Ollama's `llama/README.md` lists what a bump must review: build options, backend discovery symbols, llama-server launch args and log lines Ollama parses, streaming frames, and "speculative/MTP paths". — source: `asserted`
- A user can read the bundled build from any install with `<runner>/llama-server --version`; the Homebrew 0.34.4 runner prints `version: 0.4.1-dev (build 11081, commit 161755f29)`. — source: `asserted`
- 2026-05-29: Ollama removes its CGO engines and uses llama-server exclusively for GGML models. — source: `asserted`
- Bumps visible in the file history (all by the same maintainer in the commits read): b9493 (06-03), b9509 (06-04), b9637 (06-15), b9672 (06-17), b9888 (07-06), 08-04, 08-10, 08-12, 08-15, 08-18, b10630 (08-26), b10729 (09-01), b10760 (09-02), b10864 (09-10), b10969 (09-15), an unnumbered "version update" (09-22), b11232 (09-29). That is 17 bumps in about four months, roughly weekly since August. — source: `asserted`
- Patch churn rides along: b10729 regenerated the compat hooks patch because upstream removed a whole-tensor `load_data_for` read; b10969 moved the compat patch into libllama because libllama and libmtmd had duplicate symbols; the 09-22 bump removed the Laguna Metal patch as fixed upstream. — source: `asserted`
- A feature merged into llama.cpp master is not in a given Ollama release until a bump after the merge ships. PR 25592 (hybrid checkpoint restore) is still unmerged upstream, so no Ollama release has it. — source: `asserted`
- The Ollama server's own prompt cache and the MLX engine do not depend on this pin. — source: `asserted`
- The pinned llama-server is built without router support (see the router dossier), so llama.cpp's `--models-max` features are unavailable through Ollama. — source: `asserted`
- Whether the 09-22 "version update" is exactly b11081; the tags v0.34.4 and v0.35.0 both pin b11081, so it very likely is. — source: `asserted`
- Which llama.cpp build the 0.40.0 release will pin; the pre-release tag v0.40.0-rc0 pins b11081. — source: `asserted`
- The Ollama repository root has a file `LLAMA_CPP_VERSION` containing `b11232` on main on 2026-10-04. — [source](https://raw.githubusercontent.com/ollama/ollama/main/LLAMA_CPP_VERSION)
- `llama/server/CMakeLists.txt` reads the pinned commit from `../../LLAMA_CPP_VERSION` ("shared with Dockerfile") and fetches llama.cpp with FetchContent. — [source](https://raw.githubusercontent.com/ollama/ollama/main/llama/server/CMakeLists.txt)
- The same CMake file honors the environment variable `OLLAMA_LLAMA_CPP_SOURCE` for a local source tree and the option `OLLAMA_LLAMA_CPP_SKIP_COMPAT_PATCH`. — [source](https://raw.githubusercontent.com/ollama/ollama/main/llama/server/CMakeLists.txt)
- The root CMake file says GGML inference is provided by llama-server, built separately from "the pinned llama.cpp source". — [source](https://raw.githubusercontent.com/ollama/ollama/main/CMakeLists.txt)
- Release tag v0.33.3 pins b10760. — [source](https://raw.githubusercontent.com/ollama/ollama/v0.33.3/LLAMA_CPP_VERSION)
- Release tag v0.34.0 pins b10760. — [source](https://raw.githubusercontent.com/ollama/ollama/v0.34.0/LLAMA_CPP_VERSION)
- Release tag v0.34.4 pins b11081. — [source](https://raw.githubusercontent.com/ollama/ollama/v0.34.4/LLAMA_CPP_VERSION)
- Release tag v0.35.0 pins b11081. — [source](https://raw.githubusercontent.com/ollama/ollama/v0.35.0/LLAMA_CPP_VERSION)
- Release tag v0.35.1 pins b11232. — [source](https://raw.githubusercontent.com/ollama/ollama/v0.35.1/LLAMA_CPP_VERSION)
- Pre-release tag v0.40.0-rc0 pins b11081. — [source](https://raw.githubusercontent.com/ollama/ollama/v0.40.0-rc0/LLAMA_CPP_VERSION)
- The v0.35.1 release notes list "Updated llama.cpp and the MLX engine". — [source](https://github.com/ollama/ollama/releases)
- The commit "llama.cpp: version bump b11232" (PR 18652) landed on 2026-09-29. — [source](https://github.com/ollama/ollama/commits/main/LLAMA_CPP_VERSION)
- Earlier bumps in the file history: b10969 (2026-09-15, PR 18446), b10864 (09-10, PR 18317), b10760 (09-02, PR 18199), b10729 (09-01, PR 18160), b10630 (08-26, PR 18003). — [source](https://github.com/ollama/ollama/commits/main/LLAMA_CPP_VERSION)
- The 2026-09-22 bump (PR 18577) refined memory-allocation failure log substrings and removed the Laguna Metal patch because upstream fixed it. — [source](https://github.com/ollama/ollama/commits/main/LLAMA_CPP_VERSION)
- The b10729 bump regenerated the compat hooks patch because upstream removed the whole-tensor `load_data_for` read, whose last consumer was llama-quantize. — [source](https://github.com/ollama/ollama/commits/main/LLAMA_CPP_VERSION)
- The b10969 bump moved the compat patch into libllama with exported symbols because upstream build changes created duplicate symbols between libllama and libmtmd. — [source](https://github.com/ollama/ollama/commits/main/LLAMA_CPP_VERSION)
- The 2026-06-04 bump to b9509 brought "upstream Gemma 4 12B multimodal projector fixes for the n_head=0 divide-by-zero crash" on x86, CUDA, Linux and Windows. — [source](https://github.com/ollama/ollama/commits/main/LLAMA_CPP_VERSION)
- The 2026-05-29 commit "runner: Remove CGO engines, use llama-server exclusively for GGML models" is the oldest entry read in the history page. — [source](https://github.com/ollama/ollama/commits/main/LLAMA_CPP_VERSION)
- `llama/README.md` says `LLAMA_CPP_VERSION` pins Ollama's llama.cpp source and that an update can affect model loading, GPU discovery, scheduler inputs, runtime logs, streaming and compatibility patches. — [source](https://raw.githubusercontent.com/ollama/ollama/main/llama/README.md)
- `llama/README.md` tells reviewers of a bump to check "speculative/MTP paths", launch args consumed by `llm/llama_server.go`, and logs for device discovery, offload and memory accounting. — [source](https://raw.githubusercontent.com/ollama/ollama/main/llama/README.md)
- The `llama/` directory on main contains `clef`, `compat` and `server` subdirectories and a README. — [source](https://github.com/ollama/ollama/tree/main/llama)
- Measured on this Mac: the Homebrew Ollama 0.34.4 runner `llama-server --version` prints `version: 0.4.1-dev (build 11081, commit 161755f29)` built with AppleClang 21.0.0.21000334 for Darwin arm64, matching the tag's pin. — source: `asserted`
- The Homebrew client warns "client version is 0.35.1" while the server is 0.34.4, so the tag pin for a running server is the server's, not the CLI's. — source: `asserted`
- Inference: lazy-mode (llama.cpp PR 27794, merged 2026-08-27) is in b10729 and later pins, so every Ollama release from v0.34.0 includes it, while v0.33.3 shares the b10760 pin and therefore does too. — source: `asserted`
- Inference: llama.cpp PR 28335 (Gemma 4 vision attention, merged 2026-09-04) is absent from b10760 and present from b11081. — source: `asserted`
- Inference: PR 28751 (merged 2026-09-28) is absent from b11081 and, depending on merge time, may be in b11232. — source: `asserted`
