Wall-power meter calibration of SoC-rail energy estimates
Parent: Mac local LLMs: Speed, bandwidth and prefill · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
The protocol in the one calibrated study: a metering smart plug (Shelly Plug US Gen4) on the Mac, a sustained-decode loop per model at three durations (120, 360 and 720 seconds), three passes each, and a 15-minute idle baseline between runs.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- The protocol in the one calibrated study: a metering smart plug (Shelly Plug US Gen4) on the Mac, a sustained-decode loop per model at three durations (120, 360 and 720 seconds), three passes each, and a 15-minute idle baseline between runs. [source]
- The tool's built-in calibration fits the SoC counters to the wall figure with no hand tuning, and every published calibrated number carries the fit's error band (+/-2.6 to 4.5%). [source]
- Two tiers exist in the same tool: a wall-calibrated tier (sustained loops, with bands) and an "estimated" tier (uncalibrated, used for the real-traffic ledger). The author trusts the estimated tier for ordering, not for the third decimal. [source]
- The calibration is not constant across workloads. The author reports that the calibration "shifts slightly" for the largest, memory-bound models, so one fitted gain applied to every model would mis-state the biggest ones. [source]
- Holding a large model resident adds standing power even between tokens. The measured overhead rose with model size and flattened near 21 W once the model was big enough. A per-request marginal estimate with an idle baseline taken with no model loaded would credit that overhead to the requests. [source]
- Real traffic is more expensive per token than a saturated loop: running below saturation raised cost per token roughly 2 to 3 times over the best-case loop, and the real-traffic ledger prices per total token (prompt plus output) rather than per generated token. [source]
- 2026-07: the TokenWatt calibration write-up on an M3 Ultra Mac Studio (96 GB, macOS 26) followed an RTX 3090 cost study; this is the only calibrated Apple silicon energy-per-token dataset found. [source]
- Comparing devices with uncalibrated powermetrics or IOReport numbers is unsupported by Apple's own man page ("should not be used for any comparison between devices"); only a wall fit gives a cross-device basis. [source]
- The fit was made on one machine: a Mac Studio M3 Ultra. Nothing found validates the same gain on a MacBook, an M4 or M5 chip, or a non-Ultra Studio. [source]
- Cost figures are marginal electricity only. The same article prices the machine at $5,299 (M3 Ultra, 96 GB, 1 TB), so the electricity cost of a few thousand requests ("a tenth of a cent") is dwarfed by amortized hardware. [source]
- A heavy individual user generates tokens for about 1 to 2 hours of wall-clock time per day, not 8, so a saturated-loop figure overstates daily energy and an always-on idle draw can dominate. [source]
- System power scale. GreenBench assumes 10 W for system energy per token on an M4 Pro MacBook (existing). The wall-calibrated M3 Ultra study measures 94 W (gpt-oss-120b) to 138 W (dense 27B) during sustained decode. Different machines and loads, so the two are not averaged and neither refutes the other. [source]
- Per-token cost ordering. In the calibrated loop the dense 27B costs 5 to 9 times the MoE models per output token; in the uncalibrated real-traffic ledger it costs about ten times per total token. The shape matches; the magnitudes use different token bases. [source]
- Does a GPU-only powermetrics figure rank engines the same as a wall-calibrated figure? No source calibrates two engines side by side. [source]
- Does the wall calibration gain differ between M-series generations or between MacBook and desktop? No source. [source]
- What share of the wall figure is the 21 W resident-model overhead on a 16 or 32 GB Mac? The study's lowest-memory row was a 2.9 GB model. [source]
- The calibrated study used a Shelly Plug US Gen4 on an M3 Ultra Mac Studio with 96 GB of memory under macOS 26, priced at $0.31/kWh. [source]
- Each model ran a sustained generation loop at 120, 360 and 720 seconds, three passes each, with a 15-minute idle baseline between runs. [source]
- Calibrated figures carry bands of +/-2.6% to +/-4.5%, while the 30-day real-traffic ledger (about 6,300 requests) is the uncalibrated estimated tier and is "directional". [source]
- Real-traffic cost per million total tokens was about $0.78 for the dense Qwen3.6-27B, $0.071 for Qwen3.6-35B-A3B and $0.057 for Qwen3-Coder-Next. [source]
- Below saturation, real traffic raised per-token cost roughly two to three times over the best-case loop. [source]
- Resident-model standing power rose with model size and flattened at roughly 21 W, and the calibration shifted slightly for the largest memory-bound models. [source]
- The study's figures are marginal electricity only; the author prices the tested Mac Studio at $5,299. [source]
- The powermetrics man page says average power values are estimated, may be inaccurate, and should not be used to compare devices. [source]
- An article on local-versus-cloud cost says a heavy individual user's token generation is closer to 1 to 2 hours of wall-clock time per day than 8. [source]
- An energy benchmark on a Mac should publish both the calibrated SoC-rail figure with its band and the idle baseline policy (loaded versus unloaded), because the standing resident-model overhead changes marginal cost. [source]
Children
- No children recorded.