<!-- llms-explorer concept facts · https://llms-explorer.com/tree/llama-cpp-build-number-drift-in-published-mac-be/ · pack 2026-10-05 · ~1689 tokens -->

# llama.cpp build-number drift in published Mac benchmarks

> Discussion 4167 opened 2023-11-22 with one pinned commit; by 2026-08-25 its instructions pin c1d0e7a00 for M5+ chips, and by 2026-09-30 it holds M5 Ultra and M6 rows on that commit.

Parent: [Mac local LLMs: Benchmarking and comparisons](https://llms-explorer.com/tree/mac-local-llms-benchmarking-and-comparisons/) · 1 facets · 25 facts · page: https://llms-explorer.com/tree/llama-cpp-build-number-drift-in-published-mac-be/

## Facts

- Discussion 4167 opened 2023-11-22 with one pinned commit; by 2026-08-25 its instructions pin c1d0e7a00 for M5+ chips, and by 2026-09-30 it holds M5 Ultra and M6 rows on that commit. — source: `asserted`
- The thread now holds rows from builds 1493 to 10621, a span of roughly 7x in build number, which is the largest drift inside one community table in the sources read. — source: `asserted`
- The same table is not same-artifact across eras: 2023 rows label the Q4_0 model "mostly Q4_0 3.56 GiB" on backend "Metal" with "ngl 99", and 2026 rows label it "Q4_0 3.53 GiB" on backend "MTL,BLAS" with a "threads" column. — source: `asserted`
- The pin is not uniform even in the new era: the M5 Pro 16-core row cites commit 60081bb while the M5 Pro 20-core and M5 Max rows cite c1d0e7a. — source: `asserted`
- Run-to-run noise on one machine and one commit can reach 10% when cooling is limited: a 2023 main-post Q4_0 pp512 of 690.99 +- 33.76 was reproduced at 759.70 by one user and 760.65 and 787.24 by the maintainer after a laptop-cooling suspicion. — source: `asserted`
- A beta OS or toolchain can drift under a fixed binary: macOS 27 beta 26A5378j broke every `coreai-build` AOT compile until the IR generator moved from coreai-torch 0.4.0 to 0.4.1, and `uv run` silently resynced the virtual environment back to the 0.4.0 pin, invalidating the first fix probe. — source: `asserted`
- A published tok/s that does not record the OS build cannot be reproduced: June and July iPhone MLX rows differed 1.3-1.4x with an identical app binary and the OS build was unrecorded. — source: `asserted`
- Compare across builds vs re-run on one build. A thread commenter (2026-08-30) says the M5 rows and old rows are "incomparable" and the options are to benchmark M5 on the old build or re-benchmark old machines on the new one. The maintainer's answer to a duplicate M-series submission is that the missing data are for the M5 family, i.e. the table is extended forward and not re-run. Both positions stand in the thread. — source: `asserted`
- How much of the M4 Max to M5 Max Q4_0 tg gain (83.06 to 119.92) is the build and how much is the chip: no row measures an M4 Max on c1d0e7a in the pages read. — source: `asserted`
- Whether the 3.56 to 3.53 GiB difference is a different GGUF file or a changed size calculation. — source: `asserted`
- Discussion 4167 prints the llama.cpp build number beside each commit, and 8e672ef is build 1550 while c1d0e7a is build 10621. — [source](https://github.com/ggml-org/llama.cpp/discussions/4167)
- Other builds printed in discussion 4167 include 795cd5a (1493), d103d93 (1553), 55978ce (1555), e9c13ff (1560), 22da055 (1566), c3a1128 (8509), 073bb2c (8762), 2e97c5f (9100) and 52b3df0 (9754). — [source](https://github.com/ggml-org/llama.cpp/discussions/4167)
- The M5 Pro 16-core row cites commit 60081bb whereas the M5 Pro 20-core and M5 Max rows cite c1d0e7a. — [source](https://github.com/ggml-org/llama.cpp/discussions/4167)
- 2023 rows describe the Q4_0 model as "mostly Q4_0" at 3.56 GiB on backend Metal with ngl 99, while 2026 rows describe it as Q4_0 at 3.53 GiB on backend MTL,BLAS with a threads column. — [source](https://github.com/ggml-org/llama.cpp/discussions/4167)
- For the M2 Ultra, F16 pp512 rose from 1401.85 to 1627.82 and F16 tg128 from 41.02 to 49.55 between the 2023 row and the 2026-08-25 c1d0e7a row with flash attention. — [source](https://github.com/ggml-org/llama.cpp/discussions/4167)
- For the M2 Ultra, Q8_0 pp512 rose from 1248.59 to 1486.51 and Q8_0 tg128 from 66.64 to 82.83 between the same two rows. — [source](https://github.com/ggml-org/llama.cpp/discussions/4167)
- In 2023 a main-post Q4_0 pp512 of 690.99 +- 33.76 was reproduced by another user at 759.70 and by the maintainer at 760.65 and 787.24 after the maintainer suspected cooling limits of a smaller laptop. — [source](https://github.com/ggml-org/llama.cpp/discussions/4167)
- A commenter on 2026-08-30 called results on different llama.cpp versions "incomparable" and suggested either benchmarking M5 on the old build or re-benchmarking old machines on the new one. — [source](https://github.com/ggml-org/llama.cpp/discussions/4167)
- llm-benchpacks' 2026-05-05 Qwen3.6 M4/M5 sweep recorded llama.cpp llama-server build 9020, Ollama 0.20.5 and mlx_lm.server 0.31.3 in its summary and run metadata. — [source](https://raw.githubusercontent.com/ephes/llm-benchpacks/main/docs/qwen36-m4-m5-benchmark-summary.md)
- llm-benchpacks states that the MLX and GGUF artifacts in its sweep are not bit-identical, so its numbers are runtime-and-format comparisons and not runtime-only comparisons. — [source](https://raw.githubusercontent.com/ephes/llm-benchpacks/main/docs/qwen36-m4-m5-benchmark-summary.md)
- macOS 27 beta 26A5378j made every `xcrun coreai-build compile` abort with "LLVM ERROR: cannot unwrap empty `odiec_module_t`" for both iOS and macOS platforms and both tried toolchains. — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/methodology/coreai-build-regression-2026-07.md)
- The cause was that the updated beta's odiec accepts only IR with "AICode versioned locations", which coreai-torch 0.4.1 emits and 0.4.0 does not, and every 0.4.0-era IR became uncompilable on that OS. — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/methodology/coreai-build-regression-2026-07.md)
- `uv run` silently resyncs the virtual environment to the 0.4.0 pin and invalidated the first probe of the 0.4.1 export. — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/methodology/coreai-build-regression-2026-07.md)
- The bench requires re-measured Core AI numbers to be labelled with their artifact lineage because 0.4.1 IRs are a new artifact generation whose lowering may differ from the June bundles. — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/methodology/coreai-build-regression-2026-07.md)
- A binary built before a beta OS update was ABI-broken against the updated FoundationModels.framework and had to be rebuilt with the beta Xcode toolchain. — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/methodology/coreai-build-regression-2026-07.md)
