<!-- llms-explorer concept facts · https://llms-explorer.com/tree/energy-per-token-measurement-on-apple-silicon/ · pack 2026-10-05 · ~3401 tokens -->

# Energy per token measurement on Apple Silicon

> powermetrics prints per-sample CPU, GPU and ANE power in mW; energy is the sum of those samples times the interval, so the result depends on the sampler set, the interval and the window edges.

Parent: [Mac local LLMs: Speed, bandwidth and prefill](https://llms-explorer.com/tree/mac-local-llms-speed-bandwidth-and-prefill/) · 2 facets · 52 facts · page: https://llms-explorer.com/tree/energy-per-token-measurement-on-apple-silicon/

## Facts

- powermetrics prints per-sample CPU, GPU and ANE power in mW; energy is the sum of those samples times the interval, so the result depends on the sampler set, the interval and the window edges. — source: `asserted`
- IOReport's "Energy Model" channel group gives cumulative per-subsystem energy counters read in-process with no sudo, which makes per-request windows possible; the values are believed to be model-based estimates (utilization, frequency, voltage), not sensor readings. — source: `asserted`
- Both routes measure SoC rails. DRAM is included only where the IOReport channel exists; powermetrics package figures exclude DRAM and peripherals; wall meters include everything. — source: `asserted`
- Decode is memory-bandwidth-bound, so package power moves little with model size and J/token mostly tracks 1/(tok/s). — source: `asserted`
- Zeus (ml-energy) wrapped IOReport as zeus-apple-silicon: M5 support landed 2026-03-29 (release 1.1.0), multi-die Ultra support 2026-04-02. — source: `asserted`
- powermetrics on a macOS 27 beta lost its CPU power line; the bench moved its ranked Mac column to a GPU-only figure on 2026-07-28. — source: `asserted`
- GreenBench (arXiv 2608.28667, 2026-08-24) proposed energy per token as a standard benchmark metric alongside accuracy and throughput. — source: `asserted`
- On the macOS 27 beta the powermetrics CPU sampler read `CPU Power: 0 mW` on every sample of the raw log (288 of 288) and printed no cluster-level lines, while the GPU sampler read real values; any combined figure is then GPU+ANE only. — source: `asserted`
- The same sampler was flaky per cell: it fired in 2 of 8 cells (16.8 W and 23.5 W) and read 0 in the other six, so a recorded "combined" J/token mixed two bases within one arm (MLX read 0.116 in one round and 0.215 in the other). — source: `asserted`
- The bias is not symmetric across arms: llama.cpp does more host-side CPU work than the Metal-decode arms, so a GPU-only figure is a lower bound for it with a larger missing share. — source: `asserted`
- Rows from before the beta are not comparable with the beta rows, because it is not recorded whether their combined figure had a live CPU sampler. — source: `asserted`
- Battery-delta on iPhone quantises at 5% steps in practice, so a single cell carries about 10-20% error at a 5% drop and n=2 medians are not rankable; the earlier "misdiagnosis" that battery windows measured something else was corrected in the audit addendum (deltas are an outcome of power draw at gauge resolution). — source: `asserted`
- An energy cell started thermally "fair" instead of "nominal" publishes a wrong ranking; the app-side thermal gate deferred four cells and re-ran them after a 600 s cool. — source: `asserted`
- powermetrics needs a window of at least a few samples; the wrapper aborts when fewer than 4 power samples land inside the run. — source: `asserted`
- Running the wrapper as root makes the benchmarked tool re-download models into /var/root/.cache/huggingface, so only powermetrics is elevated. — source: `asserted`
- Package power magnitude. Side A (GreenBench, M4 Pro 48 GB, Ollama v0.21.2): CPU+GPU package power averages 413 mW and peaks at 474 mW during decode, giving 4.4-11.6 mJ per token at package level. Side B (apple-silicon-llm-bench, M4 Max): whole-system powermetrics 12.7-24.7 W and GPU-only 0.116-0.189 J/token. A 30x gap is not explained by M4 Pro vs M4 Max; GreenBench samples every 2 s and treats 0.47 W as the "peak", so its package figure is the less credible of the two. Neither side is averaged. — source: `asserted`
- System-level J/token. GreenBench's system J/token (0.09-0.25) is a constant assumed 10 W times per-token latency, not a measurement; the bench's J/token comes from measured samples. The two columns are not the same kind of number. — source: `asserted`
- Mac decode-window figure vs whole-run GPU-only figure for MLX. Existing: 0.090 J/token at 14.6 W. New GPU-only whole-run: 0.1164 and 0.1155 J/token. The sources do not say whether the windows differ; both are published. — source: `asserted`
- Whether IOReport energy (zeus) and powermetrics agree on the same run: no source measures both together. — source: `asserted`
- Whether the macOS 27 beta CPU-power zero is fixed in later seeds. — source: `asserted`
- Whether a GPU-only figure ranks engines the same as a wall-meter figure on a Mac when the engines differ in host-side CPU work. — source: `asserted`
- GreenBench (arXiv 2608.28667) ran five Q4_K_M models of 3B to 9B on an M4 Pro MacBook Pro (Mac16,8, 48 GB, macOS 15.4, Ollama v0.21.2). — [source](https://arxiv.org/html/2608.28667v1)
- GreenBench sampled powermetrics every 2 seconds and reported idle 267 mW, inference average 413 mW and inference peak 474 mW of CPU+GPU package power. — [source](https://arxiv.org/html/2608.28667v1)
- GreenBench's package energy per token is 0.47 W times per-token latency and its system energy per token is an assumed 10 W times per-token latency, with the 8-12 W range estimated from Apple's thermal specifications. — [source](https://arxiv.org/html/2608.28667v1)
- GreenBench Table IV gives 115.2 tok/s and 0.09 J/token for Llama 3.2 3B and 42.0 tok/s and 0.25 J/token for Gemma 2 9B at system level. — [source](https://arxiv.org/html/2608.28667v1)
- GreenBench states its package figure excludes DRAM and recommends zeus-apple-silicon or hardware wall meters for validated measurements. — [source](https://arxiv.org/html/2608.28667v1)
- GreenBench reports HumanEval code generation used 5 to 8 times more energy than its MMLU sample although the set is only 1.6 times larger, because output length scales linearly with time. — [source](https://arxiv.org/html/2608.28667v1)
- zeus-apple-silicon reads Apple's private IOReport framework "Energy Model" channel group with no sudo and reports energy in millijoules per window. — [source](https://github.com/ml-energy/zeus-apple-silicon)
- zeus-apple-silicon says IOReport values are believed to be model-based estimates, resolution is 1 mJ for most channels, and windows under about 10 ms may be dominated by quantisation noise. — [source](https://github.com/ml-energy/zeus-apple-silicon)
- zeus-apple-silicon's AppleEnergyMetrics exposes cpu_total, per-core and per-cluster CPU energy, dram, gpu, gpu_sram and ane fields, and a field is None when the chip does not expose it. — [source](https://github.com/ml-energy/zeus-apple-silicon)
- On an M5 Max the zeus cluster totals can be 1.5x to 2x the sum of per-core values because clusters include shared L2 and interconnect energy. — [source](https://github.com/ml-energy/zeus-apple-silicon)
- zeus-apple-silicon maps the M5 Pro/Max Performance cores (MCPU) to its efficiency_* fields and Super cores (PCPU) to its performance_* fields, only the M5 Max is tested, and GPU SRAM is not available on M5. — [source](https://github.com/ml-energy/zeus-apple-silicon)
- zeus-apple-silicon supports overlapping labelled windows, a restart flag for crashed notebook cells and get_cumulative_energy since an unspecified fixed point. — [source](https://github.com/ml-energy/zeus-apple-silicon)
- The bench's Mac energy wrapper co-runs `sudo -n powermetrics -b 1 -s cpu_power,gpu_power,ane_power -i 100` and computes energy as CPU+GPU+ANE power times elapsed time. — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/scripts/measure_energy.py)
- The wrapper waits until powermetrics has written more than 1024 bytes before starting the benchmark, so sub-second runs do not fall outside the window. — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/scripts/measure_energy.py)
- The wrapper opens the powermetrics log as the user and passes the file descriptor as stdout, because --output-file would create a root-owned mode 0600 file. — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/scripts/measure_energy.py)
- The wrapper avoids preexec_fn=os.setsid because a new session loses the controlling tty and sudo's credential cache is keyed on uid and tty, which would make `sudo -n` fail silently and leave a 0-byte log. — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/scripts/measure_energy.py)
- The wrapper's docstring says powermetrics samples may carry a few hundred mW of bias, so results compare runtimes on the same Mac and not devices. — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/scripts/measure_energy.py)
- The wrapper has a --cmd mode that wraps an arbitrary command for engines the harness does not drive on Mac, given a --tokens count to compute J/token. — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/scripts/measure_energy.py)
- On a macOS 27 beta powermetrics reported CPU Power 0 mW on all 288 of 288 samples with no cluster-level power lines, while the GPU sampler read real values. — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/results/raw/2026-07-28-gemma4-e2b-protocol-mac/README.md)
- The CPU sampler fired in only 2 of 8 Mac energy cells (MLX round 2 at 16.8 W and OptiQ round 1 at 23.5 W), so recorded combined J/token mixed two bases (MLX 0.116 against 0.215 across its two rounds). — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/results/raw/2026-07-28-gemma4-e2b-protocol-mac/README.md)
- The bench's ranked Mac column is energyJoulesPerTokenGPU derived by derive_gpu_energy.py, and its two rounds agree within 1.5% per arm. — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/results/raw/2026-07-28-gemma4-e2b-protocol-mac/README.md)
- Mac GPU-only J/token on Gemma-4-E2B were MLX 0.1164 and 0.1155, OptiQ 0.1369 and 0.1366, LiteRT 0.1557 and 0.1533, and llama.cpp 0.1887 and 0.1879 across the two rounds. — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/results/raw/2026-07-28-gemma4-e2b-protocol-mac/README.md)
- Mac ENERGY_*.jsonl files without "PM" are sustained-decode captures with null energy, because macOS has no battery to read, and only the ENERGY_PM files from the powermetrics wrapper carry J/token. — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/results/raw/2026-07-28-gemma4-e2b-protocol-mac/README.md)
- The bench recommends naming the Mac energy column "GPU energy per token" because llama.cpp's missing CPU share makes the cross-arm ranking biased. — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/results/raw/2026-07-28-gemma4-e2b-protocol-mac/README.md)
- The iPhone energy re-run in one session with nominal start gave average power MLX-PTQ 4.83 W, Cactus 4.89 W, llama.cpp 4.72 W, LiteRT 9.51 W and OptiQ 9.06 W, with sustained rates of 29.4, 29.8, 22.2, 41.0 and 21.8 tok/s. — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/results/raw/2026-07-29-gemma4-e2b-protocol/README.md)
- iPhone battery drops in that session were 5% for MLX, Cactus and llama.cpp and 10% for LiteRT and OptiQ, and gauge quantisation implies about 10-20% error per cell at a 5% drop. — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/results/raw/2026-07-29-gemma4-e2b-protocol/README.md)
- A second iPhone energy round exposed the 5%-step battery-gauge quantisation, since cross-round pairs disagreed by about 2x wherever the two deltas landed on different steps. — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/results/raw/2026-07-29-gemma4-e2b-protocol/README.md)
- The app-side thermal gate (YARDSTICK_THERMAL_DEFER, exit code 7 before model load, plus a driver retry) fired four times in the iPhone energy block, and the deferred cells re-ran at nominal after a 600 s cool. — [source](https://github.com/john-rocky/apple-silicon-llm-bench)
- The earlier iPhone energy ranking LiteRT 0.122 below MLX 0.151 came from cells that started at thermal "fair", and at enforced nominal start MLX read 0.164 against LiteRT 0.232. — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/results/raw/2026-07-29-gemma4-e2b-protocol/README.md)
- The iPhone energy block was interrupted for about 3.9 hours between the llama.cpp cell and the two Cactus cells, and every cell on both sides of the gap passed the nominal thermal gate. — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/results/raw/2026-07-29-gemma4-e2b-protocol/README.md)

## Corrections and disagreements

- iPhone llama.cpp power. The existing dossier holds both about 9.8 W with 0.483 J/token (one campaign) and 0.213 J/token (nominal re-run) without a wattage for the second. The nominal-gated session gives 4.72 W for 0.213 J/token, so CONTRADICTS: sustained-load-thermal-throttling-ane-versus-gpu.md on the 9.8 W figure (at 0.213 J/token and 22.2 tok/s the implied power is 4.7 W, not 9.8 W). Different campaigns and gating; the nominal-gated one is the bench's own correction. — source: `asserted`
