<!-- llms-explorer concept facts · https://llms-explorer.com/tree/apple-silicon-llm-bench-fairness-rules-and-inter/ · pack 2026-10-05 · ~2807 tokens -->

# apple-silicon-llm-bench fairness rules and interleaved-arm protocol

> Run order is a hardware-state variable. A large model heats the GPU within seconds and decode tok/s on a Mac swings about 30% with thermal state, so block ordering measures the order. Disclosing thermal state (rule 9) does not neutralise it.

Parent: [Mac local LLMs: Benchmarking and comparisons](https://llms-explorer.com/tree/mac-local-llms-benchmarking-and-comparisons/) · 1 facets · 41 facts · page: https://llms-explorer.com/tree/apple-silicon-llm-bench-fairness-rules-and-inter/

## Facts

- Run order is a hardware-state variable. A large model heats the GPU within seconds and decode tok/s on a Mac swings about 30% with thermal state, so block ordering measures the order. Disclosing thermal state (rule 9) does not neutralise it. — source: `asserted`
- The per-trial spread check catches only an arm whose own trials drift. An arm that is uniformly depressed by the previous arm's heat passes the spread check, so spread cannot replace interleaving. — source: `asserted`
- Comparisons are made only inside one session window, and the session is part of every table cell. Across-session numbers are compared through an anchor cell re-measured in each sitting. — source: `asserted`
- Rule citations in code moved from numbers to slugs after the two numbering schemes collided and scripts cited "rule 3" ambiguously (the file's section 3 is quantization, CLAUDE.md's rule 3 is budget and mode mixing). — source: `asserted`
- 2026-07-13: the first-ever, cold and warm regimes were split after conflating them produced a 2.6x prefill "discrepancy" on the E2B card reconciliation. — source: `asserted`
- 2026-07-17: audit of the Gemma-4-E2B table found the "winner" trophy measured checkpoint quality (PTQ vs QAT) rather than runtime speed; rule 1 and rule 2 were added. — source: `asserted`
- 2026-07-28: audit addendum re-labelled the n>=7 floor as the bench's own convention and superseded an agreed-protocol section claiming LiteRT cannot report prompt tokens. — source: `asserted`
- 2026-08-15: rule 11 (interleave) added from a failed block-ordered capture of Muse-Glimmer-30B on M4 Max. — source: `asserted`
- Block-ordered capture, Muse-Glimmer-30B on M4 Max: ExecuTorch decayed 23.5 to 17.4 tok/s inside its own block, and Core AI, run second on the GPU the first block had heated, read 15.6 to 20.7 against its own true value of about 27.4. The two arms were wrong in opposite directions. — source: `asserted`
- An arm may be structurally unable to run an equalized protocol. Core AI on iPhone dies at setup with signal 9 (jetsam) because KV cap 1024 conflicts with maxTokens 2048; the bench records this as a finding (`arm_can_energy`) rather than dropping the cell or granting a smaller budget (the earlier row existed only because that session gave Core AI maxTokens=192 while others ran 2048). — source: `asserted`
- An engine may not be the pinned release. The Core AI arm in the 3-way capture is a fork (john-rocky/coreai-models at 58aab35 plus an archived 2-file diff) because release 0.2.0 neither runs on that OS seed nor supports the model; the row is labelled "fork engine" and the self-made-artifact asymmetry is disclosed. — source: `asserted`
- Harness defects found in the Gemma-4-E2B re-capture: model revision not recorded per row (MLX checkpoint 2c3e507 vs 2387675 was ambiguous), mlx-swift-lm package drifted so neither checkpoint loaded on Mac, a Mac capture stamped as an iPhone device, and a published number traced to the wrong model. — source: `asserted`
- A cell with n below the floor is published with disclosure, not topped up in a later sitting (iPhone A2 mlx and optiq n=6 after a launch was lost when the device left WiFi range). — source: `asserted`
- Block order vs interleave: the bench's earlier cold-launch tables were block-ordered and were corrected after the fact; the claim that "disclosure suffices" (rule 9) and the claim that "order must be eliminated" (rule 11) were both held by the same repo at different dates. The later rule supersedes. — source: `asserted`
- Neutrality: the repo says "neutral" is load-bearing, and also records that Google's LiteRT team is evaluating extending it into continuous competitive benchmarking and that each arm runs at its own best available build. Whether best-build-per-arm is neutral depends on whether the recipe is visible in the table (rule 1). — source: `asserted`
- Whether an N-stream or concurrency axis is covered by any fairness rule: none of the pages read states one. — source: `asserted`
- Whether 45 s cooldowns (3-way capture) and 100 s cooldowns (warm campaigns) give different order residue: no page measures it. — source: `asserted`
- The bench cites its rules by slug, and the slug table maps CLAUDE.md working rules to quant-per-arm-rule, quant-label-rule, budget-mode-rule, spread-rule and stored-report-rule. — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/methodology/fairness-rules.md)
- The 11 rule slugs are same-budget, cold-warm-split, quant-explicit, failed-runs-stay, disclose-difficulty, same-device-class, same-build-config, no-cherry-pick, disclose-hw-state, official-sdk and interleave-arms. — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/methodology/fairness-rules.md)
- Scripts once cited "rule 3" ambiguously because the file's section 3 is quantization and CLAUDE.md's rule 3 is budget and mode mixing. — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/methodology/fairness-rules.md)
- In the block-ordered first attempt ExecuTorch decayed from 23.5 to 17.4 tok/s inside its own block. — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/methodology/fairness-rules.md)
- In the same attempt Core AI, running second on a GPU heated by the first block, read 15.6 to 20.7 tok/s against its own true value of about 27.4. — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/methodology/fairness-rules.md)
- The per-trial spread check catches only the arm with a wide spread and misses an arm that is uniformly depressed by the previous arm's heat. — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/methodology/fairness-rules.md)
- The interleaved Muse-Glimmer-30B 3-way on M4 Max measured Core AI 27.43, MLX 27.36 and ExecuTorch 24.00 tok/s with 45 s cooldowns and every round published. — [source](https://github.com/john-rocky/apple-silicon-llm-bench)
- The failed block-ordered capture is kept in the repo as headtohead-blockordered-256-failed.log, per the rule that failed runs stay on record. — [source](https://github.com/john-rocky/apple-silicon-llm-bench)
- The 3-way Core AI arm is the fork john-rocky/coreai-models at 58aab35 plus an archived 2-file diff, labelled "fork engine", because Core AI 0.2.0 neither runs on that OS seed nor supports muse_glimmer. — [source](https://github.com/john-rocky/apple-silicon-llm-bench)
- The 3-way capture is reproducible with `./reproduce mac muse-glimmer-30b-3way`, which runs scripts/bench_muse_glimmer_3way_mac.sh. — [source](https://github.com/john-rocky/apple-silicon-llm-bench)
- A "first-ever / cold / warm" split was added on 2026-07-13 after conflating the regimes produced a 2.6x prefill discrepancy that was purely protocol. — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/methodology/fairness-rules.md)
- The 2026-07-17 audit found the published Gemma-4-E2B table crowned LiteRT while MLX ran PTQ and llama.cpp ran a third-party PTQ, so the trophy measured checkpoint quality. — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/CLAUDE.md)
- The same wNa8o8 LiteRT weights score 85% on LiteRT and 48% on any fp16-activation runtime, and two independent implementations returned identical wrong answers on the fp16-activation side. — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/CLAUDE.md)
- Each arm runs at its own best available build, which the bench calls fair only if the recipe is visible in the table. — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/CLAUDE.md)
- The n>=7 sample floor is the bench's own convention, not a published standard, and a cell below it (iPhone A2 mlx and optiq at n=6) is published with disclosure rather than topped up across sittings. — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/results/raw/2026-07-29-gemma4-e2b-protocol/README.md)
- The README prints no cross-cell ratios on purpose: cells are compared only within one session and the session is part of every cell. — [source](https://github.com/john-rocky/apple-silicon-llm-bench)
- An anchor cell (LiteRT chat) read 61.1 tok/s in one sitting and 59.4 in another, a 2.8% drift, and cross-block ratios are routed through that anchor. — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/results/raw/2026-07-29-gemma4-e2b-protocol/README.md)
- MLX Qwen3-0.6B on iPhone read 125.8 to 133.1 tok/s in June and 166 to 179 tok/s in July under apparently identical conditions. — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/results/raw/2026-07-13-mlx-variance/README.md)
- Rebuilding the June app tree with the June package pins and running it on the same device gave 164.8 cold and 169.1 warm, matching July, so the binary was exonerated. — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/results/raw/2026-07-13-mlx-variance/README.md)
- The bench attributes the June-to-July shift to device state, with an iOS 27.0 beta build update as prime suspect, because the app logged only systemVersion "27.0" and the June build number was never recorded. — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/results/raw/2026-07-13-mlx-variance/README.md)
- Within-session repeatability was +-1.5% and across the July sessions +-6%, and the MLX inter-token latency p50 moved from 8.20 ms to 5.6-6.0 ms. — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/results/raw/2026-07-13-mlx-variance/README.md)
- The Gemma-4-E2B re-capture added modelRevision per row, pinned one mlx-swift-lm revision for both builds, and made import_native_benchmark require --device after a Mac capture was stamped iPhone18,1. — [source](https://github.com/john-rocky/apple-silicon-llm-bench)
- Core AI cannot run the equalized iOS protocol because KV 1024 conflicts with maxTokens 2048 and the app dies at setup with signal 9 (jetsam); the bench encodes this as `arm_can_energy` rather than dropping the arm silently. — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/results/raw/2026-07-29-gemma4-e2b-protocol/README.md)
- Running an agent session from the wrong working directory (bootstrap script run from the wrong cwd) cost hours on 2026-07-17, so the repo instructs sessions to launch inside its own directory. — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/CLAUDE.md)
