Resident-model standing power overhead on unified memory (about 21 W)
Parent: Mac local LLMs: Speed, bandwidth and prefill · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
The article attributes the overhead to "memory bandwidth, regulators, fans", separate from the arithmetic of generation. It does not measure which component contributes.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- The article attributes the overhead to "memory bandwidth, regulators, fans", separate from the arithmetic of generation. It does not measure which component contributes. [source]
- The article gives no per-model overhead values, no table and no procedure for isolating the overhead. The "21 W" figure appears once, in prose. [source]
- The TokenWatt tool reads whole-SoC rails through IOReport and subtracts a rolling idle baseline "so you're measuring the marginal cost of the request, not the machine sitting there breathing". The article does not say whether that baseline is taken with the model loaded. [source]
- The same paragraph says the calibration "also shifts slightly" for the largest memory-bound models, so the overhead may partly be a wall-versus-rail calibration effect and not an isolated physical load. [source]
- The largest resident model was gpt-oss-120b at MXFP4, about 59 GB, on a 96 GB machine; no larger model was tested. [source]
- 2026-07: the article is dated by its image upload path (2026/07) and is the only source for the figure. [source]
- Scale against the generation draw: 21 W is about 22% of the 94 W sustained draw the study reports for gpt-oss-120b and about 15% of the 138 W for the dense 27B model (derived from those two figures in wall-power-meter-calibration-of-soc-rail-energy.md). [source]
- Scale against a year: 21 W held continuously is 184 kWh per year, about $57 at the study's $0.31 per kWh (21 x 8,760 / 1000 x 0.31). [source]
- The figure comes from one M3 Ultra Mac Studio with 96 GB. No source gives the overhead for a MacBook, where fans and regulators differ, or for a 16-32 GB Mac. [source]
- The article's own advice is to keep only actively used models resident, since the cost accrues while idle. [source]
- None found. The one competing statement in existing dossiers is a MacBook battery drain of 1.5-2% per hour for a resident 14B model, which measures a different quantity on a different machine and is not comparable in watts. [source]
- Whether the overhead persists when the weights are mapped but untouched and only the page cache holds them. [source]
- Whether it is DRAM and fabric power, a calibration artifact, or both; a test would compare mains power with the same bytes allocated by a non-model process. [source]
- The study attributes the standing overhead to "memory bandwidth, regulators, fans" without measuring which. [source]
- The article publishes no per-model overhead table; the roughly 21 W plateau is stated once in prose. [source]
- The article does not state whether TokenWatt's rolling idle baseline is measured with a model loaded. [source]
- The largest resident model in the study was gpt-oss-120b MXFP4 at about 59 GB on a 96 GB M3 Ultra. [source]
- The article recommends keeping only actively used models resident because of the standing overhead. [source]
- A 21 W standing overhead is about 22% of the 94 W gpt-oss-120b sustained draw and about 15% of the 138 W dense 27B draw. [source]
- A 21 W overhead held for a year costs about $57 at $0.31 per kWh. [source]
- The 21 W figure is a single-machine, single-author, prose-only claim and may include a calibration effect. [source]
Children
- No children recorded.