Benchmark methodology pitfalls for local LLM runtimes
Parent: Mac local LLMs: Benchmarking and comparisons · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
llama.cpp's canonical Apple-Silicon table (discussion 4167) kept one 2023 build (8e672ef) for years to hold "all performance factors even"; as of the 2026-08-25 edit it pins a new commit (c1d0e7a00) for M5+ chips, so old and new rows in that table are not comparable.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- llama.cpp's canonical Apple-Silicon table (discussion 4167) kept one 2023 build (8e672ef) for years to hold "all performance factors even"; as of the 2026-08-25 edit it pins a new commit (c1d0e7a00) for M5+ chips, so old and new rows in that table are not comparable. [source]
- john-rocky/apple-silicon-llm-bench evolved from a block-ordered, cold-launch table into 11 fairness rules after audits (2026-07-17 quant audit, 2026-08-15 interleave finding). [source]
- Block-ordered arms: a hot GPU from arm A depresses arm B uniformly, so the per-trial spread check cannot catch it. Interleave arms. [source]
- First run after another arm pulled a multi-GB model through the page cache reads far below true (16.2 vs 27.4 tok/s). [source]
- GPU contention (another app, a second server) roughly halves decode with the only tell being trial spread (72.7/117.9/72.4 contended vs 171.3/171.6/171.0 idle). [source]
- Random-token prompts with EOS disabled (mlx_lm.benchmark) measure raw throughput but not realistic text, MTP/speculative acceptance, or MoE routing on natural text; treat as a ceiling for speculative setups. [source]
- exo_bench defaults to --repeat 1 and --warmup 0: a default run is one cold sample. [source]
- Budget semantics differ: max_tokens is generation budget in one runtime and total context in another; undersizing corrupts rather than truncates. [source]
- Thinking mode defaults differ by chat template (HF renders ON, swift-transformers OFF); output token counts then differ. [source]
- Quant label is not a spec (wNa8o8 schema, 4-bit affine vs Q4_K_M vs MXFP4), and a QAT vs PTQ checkpoint swing was worth ~9 GSM8K points on MLX. [source]
- Cold vs warm as headline: apple-silicon-llm-bench says warm (median of runs 2-4) is the cross-vendor headline and cold is a first-use latency metric; app-launch and agent-restart use cases need cold. Report both. [source]
- Warm-up sensitivity is engine-dependent (Core AI ~2.5x, LiteRT ~1.1x, MLX ~flat), so cold rankings do not predict warm rankings. [source]
- exo_bench's `/bench/chat/completions` timing method vs client-side streaming timing: only energy parity was published (server 1940 J vs client 1931 J, +0.5%), not tok/s parity. [source]
- stared/benching-local-llms returned a 404 on 2026-10-04; could not be assessed (may be renamed or private). [source]
- Whether llama-bench and llama-server agree on pp/tg for the same build and flags is not measured in sources read. [source]
- 1 Same weights lineage: same BF16 parent and tokenizer/chat template; compare prompt token counts across engines before trusting a run (probe one item through every arm). [source]
- 2 Same quant and bpw: record effective bpw and file size per arm; state recipe in the table. [source]
- 3 Same build: pin llama.cpp commit and mlx/mlx-lm version; Debug vs Release differ; prefer the official runtime SDK and note the version. [source]
- 4 Same prompt, same token budget, same sampling (greedy, batch 1 for decode comparisons). [source]
- 5 Declare regime per number: first-ever (shader/pipeline caches built), cold (fresh process, caches on disk), warm (in-process run N>=2). Discard run 1 for warm; report median of runs 2-4. [source]
- 6 Cooldown: at least 100 s between cells in warm campaigns (tens of seconds minimum per run for a 30B-class model); confirm thermal state nominal after the fact and re-run flagged cells. [source]
- 7 Interleave arms per prompt (A/B/A/B); never run one arm's block then the next; publish every round. [source]
- 8 Disclose hardware state: charging state, Low Power Mode, thermal state at start; any can move throughput 30%+. [source]
- 9 Trial spread rule: decode trials must agree within a few percent; quote the spread; discard a wide-spread number. [source]
- 10 Median of N>=3, never best-of; failed/OOM/unsupported rows stay in the table with a reason. [source]
- 11 Do not average across device classes; show them in separate rows. [source]
- 12 Disclose integration difficulty (glue code, no streaming, no cancel) separately from speed. [source]
- 13 A number without a stored raw report (per-run JSONL with repo id and quantization) is not a measurement. [source]
- 14 Never mix modes: thinking on/off, max_tokens semantics, MTP/speculative on in one arm and off in another. [source]
- 15 Test at depth, not only at 512: use llama-bench -d (prefilled KV) and mlx_lm.benchmark -p with long prompts; decode at depth 512 already differs from depth 0. [source]
- 16 Concurrency is a separate axis: single-stream decode and N-stream aggregate decode rank engines differently; label which. [asserted, consistent with held Rapid-MLX 8-stream data] [source]
- 17 Dtype: fix KV cache type (llama-bench -ctk/-ctv default f16) and activation dtype; do not compare f16 KV against quantized KV. [source]
- 18 Prefix-cache hygiene: for cold runs, use a unique random prefix per request or clear the runtime cache dir; for warm-cache tests state the hit fraction. [asserted, builds on held oMLX/Ollama cache claims] [source]
- 19 Page-cache hygiene for load-time and first-run numbers: a model recently read is resident in the page cache; separate load time (cold disk vs warm cache) from throughput. [source]
- Options and defaults: -p 512, -n 128, -pg <pp,tg> (combined), -d 0 (context depth, prefills KV), -b 2048, -ub 512, -r 5 repetitions, -ctk/-ctv f16, -ngl -1, -fa auto, -ncmoe 0, -lm auto (load mode), --delay 0 s, --no-warmup flag (warmup runs happen by default), -o md|csv|json|jsonl|sql. [source]
- Results are mean tok/s with standard deviation over -r repeats; json also keeps per-repetition values. [source]
- Measurements exclude tokenization and sampling time, so llama-bench tok/s is an upper bound on what a server or CLI delivers. [source]
- Every list-valued option is swept as a cross-product; ranges use first-last, first-last+step, first-last*mult. [source]
- Recommended Mac command: `llama-bench -m MODEL.gguf -p 512,4096 -n 128 -d 0,8192,32768 -fa 1 -ngl 99 -r 5 --delay 30 -o json > out.json`; add `-ncmoe N` for MoE offload. [asserted, composed from the documented flags] [source]
- Canonical community command (ggerganov table, updated 2026-08-25, pinned commit c1d0e7a00): `llama-bench -hf ggml-org/Llama-2-7B-GGUF:F16 -hf ...:Q8_0 -hf ...:Q4_0 -p 512 -n 128 -ngl 99 --delay 30 2> /dev/null`; PP uses batch 512, TG batch 1. [source]
- Flags: --model, -p/--prompt-tokens (512), -g/--generation-tokens (1024), -b/--batch-size (1), -n/--num-trials (5), --prefill-step-size (2048), --delay (0 s), --pipeline, -qa/--quantize-activations, --trust-remote-code. [source]
- It seeds mx.random with 0, builds the prompt from random token ids (not text), empties the tokenizer EOS set so generation never stops early, runs one unmeasured warmup, then N timed trials and prints prompt_tps, generation_tps, peak_memory, total_time per trial plus the plain mean (no median, no stdev). [source]
- Because the mean has no spread output, read the per-trial lines yourself and apply the spread rule; use --delay 30 on fanless machines. [source]
- Recommended: `mlx_lm.benchmark --model mlx-community/MODEL-4bit -p 512 -g 256 -n 5 --delay 30`, repeated with -p 4096 and 16384, and -b N for batch. [asserted, composed from flags] [source]
- llama-bench -p 512 -n 128 and mlx_lm.benchmark defaults (512/1024) differ in generation length (128 vs 1024); at long generation, decode drifts down with KV growth and heat, so set -g/-n equal before comparing. [asserted, from defaults in the two sources] [source]
- Apple's own M5 post uses mlx_lm.generate with prompt size 4096 and 128 generated tokens, reports TTFT (s) and tok/s, and notes TTFT is compute-bound and decode is bandwidth-bound. [source]
- `uv run bench/exo_bench.py --model M --pp 128,512 --tg 128 --max-nodes 2 --sharding tensor --repeat 3 --warmup 1 --json-out out.json`; needs nodes running via `uv run exo`; uses the /bench/chat/completions endpoint; options --instance-meta ring|jaccl|both, --sharding pipeline|tensor|both; defaults --repeat 1, --warmup 0, --max-nodes 4. [source]
- Output per placement: prompt_tps, generation_tps, peak memory; power sampling splits energy into prefill and generation (PowerSampler, boundary at the first non-prefill chunk), validated within 0.5% of client-side energy on M3 Ultra and M4 Pro. [source]
- apple-silicon-llm-bench (john-rocky): engines LiteRT-LM, llama.cpp, Core AI, MLX, Cactus; harness "short-chat" with the same 128-token budget, warm median with run 1 dropped; records per-run JSONL under results/raw/; also supports Apple's llm-benchmark protocol (512 prompt, 1024 gen, 5 trials, greedy; MLX arm via mlx_lm benchmark with same args); trial spread 0.1-0.3% reported on idle M4 Max; entry point `./bench`. [source]
- Hybrid-model CLI comparisons there use llama-bench tg256 (pp512) vs `mlx_lm generate` 256 tokens, greedy, batch 1, mains power, median of 3 processes. [source]
- ephes/llm-benchpacks targets workload-level comparison via OpenAI-compatible endpoints (`benchpack run runtime-sweep --adapter openai-chat --endpoint http://localhost:11434/v1 --openai-stream-usage omit`), covering TTFT, prefill speed, prompt-cache reuse, long-context stability, tool-call formatting and repo-task pass rates; runs a local M5 plus SSH-to-M4 workflow. [source]
- stared/benching-local-llms: URL returned GitHub 404 on 2026-10-04; no claims made. [source]
- Hardware: chip, GPU cores, RAM, macOS version, power source, Low Power Mode, thermal state at start, chassis (fan or fanless). [source]
- Software: runtime name plus commit/version, flags (-fa, -ub, -ctk/-ctv, wired limit), model repo id and revision, quant recipe, file size, effective bpw. [source]
- Protocol: prompt tokens, generation tokens, depth, batch/concurrency, sampling, thinking mode, MTP on/off, regime (first-ever/cold/warm), N trials, cooldown, run order. [source]
- Metrics: pp tok/s and TTFT (cold and cache-hit separately), tg tok/s with spread (stdev or min-max), peak memory, total wall time; energy J/token if measured. [source]
- Raw per-run records and failed runs. [source]
Children
- No children recorded.