<!-- llms-explorer concept facts · https://llms-explorer.com/tree/thermal-power-and-battery-behavior-during-local/ · pack 2026-10-05 · ~3899 tokens -->

# Thermal, power and battery behavior during local inference

> Decode is memory-bandwidth-bound: each token streams all active weights, so dense models draw more watts and give fewer tok/s than MoE models of larger parameter count (M3 Ultra measurements: 27B dense 138 W / 21.5 tok/s vs 120B MoE 94 W / 74 tok/s).

Parent: [Mac local LLMs: Speed, bandwidth and prefill](https://llms-explorer.com/tree/mac-local-llms-speed-bandwidth-and-prefill/) · 1 facets · 68 facts · page: https://llms-explorer.com/tree/thermal-power-and-battery-behavior-during-local/

## Facts

- Decode is memory-bandwidth-bound: each token streams all active weights, so dense models draw more watts and give fewer tok/s than MoE models of larger parameter count (M3 Ultra measurements: 27B dense 138 W / 21.5 tok/s vs 120B MoE 94 W / 74 tok/s). — source: `asserted`
- Energy per token = watts / tokens-per-second; electricity cost per 1M output tokens followed that ratio ($0.554 dense 27B vs $0.063-0.109 for the others at $0.31/kWh). — source: `asserted`
- Every chassis has a sustained thermal budget below its burst budget; when die temperature reaches target, clocks drop. A fanless Air sheds heat only through the case. — source: `asserted`
- Apple's Energy Modes: Automatic default; Low Power reduces power and fan noise; High Power lets fans run faster so the chip may sustain performance in very intensive workloads. Modes can be set separately for battery and adapter. — source: `asserted`
- Inference is bursty in interactive use (duty cycle roughly 5-10% by one author), so battery cost depends on workload shape far more than peak wattage. — source: `asserted`
- High Power Mode started on Max-chip Macs (2021-2023 MacBook Pro Max), reached M4 Pro MacBook Pro and Mac mini in Nov 2024, and per Apple's page now covers 2026 Mac Studio, Pro/Max MacBook Pro 2024+ and Pro-chip Mac mini. — source: `asserted`
- Ollama 0.19 (30 Mar 2026) replaced its llama.cpp Metal backend with MLX, so "MLX vs llama.cpp" now partly means "Ollama before/after". — source: `asserted`
- Throttling on the Air arrives after minutes of continuous GPU load; first-minute benchmarks are burst numbers. — source: `asserted`
- Slowness from the first token is a memory/placement problem, not thermal. — source: `asserted`
- On battery macOS may cap sustained performance independent of temperature (single-source, llmconfigurator). — source: `asserted`
- Clamshell or soft-surface use removes heat paths (single-source blogs). — source: `asserted`
- Adapter wattage and cable can cap power: one M5 Max user drained battery while on a 60 W adapter and measured a 34 W (36%) gap between two cables on the same adapter. — source: `asserted`
- A resident large model carries standing overhead (~21 W plateau on the M3 Ultra for big models, single-source) and, on battery, memory pressure drain (~1.5-2%/hr above baseline for a resident 14B, single-source). — source: `asserted`
- Prefill on long context can sit at low GPU power for minutes (about 10 W for 3.5 minutes on a 40k-token prompt), so tok/s headlines hide wall-clock cost. — source: `asserted`
- Low Power Mode cost. Side A (thinkdifferent.blog, Jakub Jirak): Low Power "will absolutely strangle inference", people report "half the expected speed". Side B (same author, Medium battery post): Low Power slows generation ~20-30% (8B from ~30 to ~21 tok/s) and cuts drain more than that. Side C (Newport, Medium): M2 Max 8.02 tok/s at low power vs 16 at high power (his own "2.5x" label does not match 16/8.02). Same author contradicts himself; no source is controlled. Treat "halves" as an upper-range anecdote. — source: `asserted`
- High Power Mode benefit. thinkdifferent says it avoids a 15-25% sag; Apple's page claims only possible gains in very intensive workloads and cites video/3D examples, no LLM numbers. MacRumors forum comment: "performance is a bit of a wash, fan noise is considerably increased" (single comment). The 15-25% figure has no measurement behind it. — source: `asserted`
- Air sustained throttling severity. willitrunai: Air M4 sustained 60-70% of peak after ~8-12 min. modelfit: throttles after 20-30 min, figures labeled estimates. llmconfigurator: "halves", but its own table is labeled not measured. MindStudio: M5 Air CPU merge-sort stress test lost only ~6% first-to-last iteration (CPU, not GPU decode; M4/M3 ~5%, M2 ~3%). No GPU-decode sustained curve with a stated method was found. — source: `asserted`
- Battery: Jirak M1 Pro says 8B costs ~11%/hr in bursty use and ~38%/hr when hammered; deploy.live M5 Max under sustained load drained ~1%/min. — source: `asserted`
- Measured sustained tok/s vs time curves for M-series Air GPU decode (first-party or reproducible). — source: `asserted`
- Controlled tok/s and joules/token for Low vs Automatic vs High on one machine. — source: `asserted`
- MLX vs llama.cpp joules per token with wall-metered power; found only tooling (TokenWatt) and throughput comparisons, no head-to-head energy dataset. — source: `asserted`
- Whether Mac mini base-chip and M4 Pro minis hold clocks under hours of decode (llmconfigurator says "effectively unconstrained" but marks it not measured). — source: `asserted`
- Reddit threads on these topics could not be fetched. — source: `asserted`
- Apple Support documents three Energy Modes: Low Power, Automatic (default), High Power; settable separately for On battery and On power adapter. — [source](https://support.apple.com/en-us/101613)
- Low Power Mode gives reduced power consumption and reduced fan noise on supported models. — [source](https://support.apple.com/en-us/101613)
- High Power Mode lets fans run at higher speeds; extra cooling "may allow" higher performance in very intensive workloads, with more fan noise. — [source](https://support.apple.com/en-us/101613)
- Apple's High Power Mode support list covers Mac Studio introduced in 2026, Mac mini 2024+ with Pro chip, MacBook Pro 2024+ with Pro or Max chip, and 2021-2023 MacBook Pro with Max chip; MacBook Air is not on it. — [source](https://support.apple.com/en-us/101613)
- Apple recommends the 96 W USB-C adapter for High Power Mode while charging on the 14-inch MacBook Pro with M4 Pro or M5 Pro. — [source](https://support.apple.com/en-us/101613)
- In macOS Sequoia 15.1 and later, Energy Modes can also be selected from Control Center. — [source](https://support.apple.com/en-us/101613)
- High Power Mode first reached Pro-chip Macs (14/16-inch MacBook Pro and Mac mini M4 Pro) in Nov 2024; previously Max-only. — [source](https://forums.macrumors.com/threads/apple-expands-high-power-mode-to-macbook-pro-and-mac-mini-models-with-m4-pro-chip.2442254/)
- A MacRumors forum commenter said performance under High Power is "a bit of a wash" while fan noise is considerably increased. — [source](https://forums.macrumors.com/threads/apple-expands-high-power-mode-to-macbook-pro-and-mac-mini-models-with-m4-pro-chip.2442254/)
- Claim "High Power Mode avoids a 15-25% sag" originates from one blog (Jakub Jirak, thinkdifferent.blog, 3 Jun 2026) with no benchmark shown. — [source](https://www.thinkdifferent.blog/blog/the-macbook-pro-setting-that-doubles-your-ai-performance/)
- Claim "Low Power Mode halves speed" in that blog is hearsay ("I've seen people benchmark..."); the same author's other post measures a 20-30% drop (30 to 21 tok/s on 8B, M1 Pro). — [source](https://jakubjirak.medium.com/your-macbook-battery-can-handle-8-hours-of-local-ai-heres-how-132693bec6b7)
- Newport reports M2 Max MacBook Pro running Gemma-3-27B Q4 at 8.02 tok/s in low power and 16 tok/s in high power. — [source](https://medium.com/@billynewport/apples-m3-ultra-mac-studio-misses-the-mark-for-llm-inference-f57f1f10a56f)
- The same blog says a 40-50k-token context made inference about 10x slower than short prompts. — [source](https://medium.com/@billynewport/apples-m3-ultra-mac-studio-misses-the-mark-for-llm-inference-f57f1f10a56f)
- willitrunai lists Air M4 passive cooling with "~8-12 min before throttle" and sustained throttle to ~60-70% of peak; Llama 3.1 8B 42 tok/s burst vs 25-30 sustained. — [source](https://willitrunai.com/blog/macbook-air-m4-vs-pro-m4-for-local-llms)
- willitrunai's Air-vs-Pro burst tok/s are labeled approximate and no method is given. — [source](https://willitrunai.com/blog/macbook-air-m4-vs-pro-m4-for-local-llms)
- modelfit says the Air M4 throttles after 20-30 minutes of intensive GPU use and that its tok/s are estimates from chip bandwidth, not lab benchmarks. — [source](https://modelfit.io/blog/macbook-air-vs-pro-llm/)
- The Air M4 and base-chip MacBook Pro M4 have the same 120 GB/s bandwidth; the M4 Pro has 273 GB/s (per willitrunai). — [source](https://willitrunai.com/blog/macbook-air-m4-vs-pro-m4-for-local-llms)
- MindStudio measured 142 GB/s STREAM bandwidth on M5 Air vs 113 on M4 Air. — [source](https://www.mindstudio.ai/blog/m5-macbook-air-local-ai-performance)
- MindStudio's all-core CPU stress test (merge sort of 1B integers, 5 rounds) showed first-to-last drop of ~6% on M5 Air, ~5% on M4 and M3, ~3% on M2; it did not test GPU LLM decode. — [source](https://www.mindstudio.ai/blog/m5-macbook-air-local-ai-performance)
- llmconfigurator shows a user trace of 42.6 tok/s at minute 1 and 19.8 at minute 10 and attributes it to thermal limits; its chassis-expectation table is explicitly "our characterisation, not measured" with no sustained-load rows in its dataset. — [source](https://llmconfigurator.com/en/guides/troubleshooting/macbook-thermal-throttling-llm)
- llmconfigurator asserts macOS caps sustained performance on battery independent of temperature. — source: `asserted`
- llmconfigurator recommends keeping the lid open, off soft surfaces, and using a smaller model or shorter context for long jobs because lower work per token lowers power and raises the plateau. — [source](https://llmconfigurator.com/en/guides/troubleshooting/macbook-thermal-throttling-llm)
- Monitoring command: `sudo powermetrics --samplers smc,gpu_power -i 1000` reports GPU power and die temperature; a power peak then lower plateau with flat temperature indicates throttling. — [source](https://llmconfigurator.com/en/guides/troubleshooting/macbook-thermal-throttling-llm)
- Check power policy with `pmset -g batt` and `pmset -g | grep -i lowpowermode`. — [source](https://llmconfigurator.com/en/guides/troubleshooting/macbook-thermal-throttling-llm)
- powermetrics supports sampler selection (-s, "all"/"default" groups), -i interval in ms, plist output, --show-plimits and --show-pstates (hardware-dependent), and needs root. — [source](https://ss64.com/mac/powermetrics.html)
- asitop is a Python TUI for Apple Silicon CPU/GPU/ANE usage built on powermetrics; last commit Jan 2023 ("Updated for Ventura"), 4.6k stars. — [source](https://github.com/tlkh/asitop)
- macmon is a sudoless monitor reading CPU/GPU/ANE power, temperatures and memory through a private macOS API (IOReport), with JSON pipe output; active in 2026 (commits Jul 2026). — [source](https://github.com/vladkens/macmon)
- powermetrics can emit plist, which is useful for feeding monitoring systems. — [source](https://whatsuphome.fi/blog/part-111-some-handy-command-line-tools-plus-mac-gpu-monitoring)
- iStat Menus plus btop were used as lightweight monitors in a two-Mac Ollama comparison (M1 mini 16.5 tok/s vs M5 Pro 62.6 tok/s, Gemma 4). — [source](https://engineersmeetai.substack.com/p/measuring-local-llm-performance-across)
- TokenWatt (open-source proxy) estimates per-request marginal energy from IOReport SoC rails with a rolling idle baseline, calibrated to a Shelly Plug wall meter (stated error +/-2.6 to 4.5%). — [source](https://towardsdatascience.com/how-much-does-a-local-llm-actually-cost-to-run-i-measured-every-watt-on-apple-silicon/)
- M3 Ultra Mac Studio (96 GB), sustained decode, wall-calibrated: Qwen3.5-4B 133.8 tok/s $0.063/1M out; Qwen3.6-35B-A3B 76.0 tok/s $0.087; Qwen3-Coder-Next ~80B 65.0 tok/s $0.103; gpt-oss-120b 74.0 tok/s $0.109 (94 W); dense Qwen3.6-27B 21.5 tok/s 138 W $0.554 (at $0.31/kWh). — [source](https://towardsdatascience.com/how-much-does-a-local-llm-actually-cost-to-run-i-measured-every-watt-on-apple-silicon/)
- In that author's real 30-day traffic (uncalibrated estimate tier), the dense 27B cost roughly 10x more per token than MoE models, and sub-saturation use raises per-token cost. — [source](https://towardsdatascience.com/how-much-does-a-local-llm-actually-cost-to-run-i-measured-every-watt-on-apple-silicon/)
- Holding a big model resident adds standing overhead rising with size and flattening near 21 W on that machine. — [source](https://towardsdatascience.com/how-much-does-a-local-llm-actually-cost-to-run-i-measured-every-watt-on-apple-silicon/)
- On M1 Pro 14" (70 Wh), 8B inference package power peaks ~18-22 W and idles under 2 W between requests. — [source](https://jakubjirak.medium.com/your-macbook-battery-can-handle-8-hours-of-local-ai-heres-how-132693bec6b7)
- Same source, Ollama Q4_K_M, battery drain per hour (active/hammer): 3B ~7%/~26%; 8B ~11%/~38%; 14B ~15%/~47%; 32B ~22%/60%+; baseline without AI 5-6%/hr. — [source](https://jakubjirak.medium.com/your-macbook-battery-can-handle-8-hours-of-local-ai-heres-how-132693bec6b7)
- Same source recommends OLLAMA_KEEP_ALIVE=10m on battery (reload cost 8B ~4 s, 14B ~7 s) and measures battery with a pmset -g batt log every 5 minutes. — [source](https://jakubjirak.medium.com/your-macbook-battery-can-handle-8-hours-of-local-ai-heres-how-132693bec6b7)
- Single author, single machine (M1 Pro); the figures are self-measured with no method files. — source: `asserted`
- A 10-hour flight on an M5 Max MacBook Pro 16" 128 GB: roughly 1% battery per minute under sustained load, battery drained even on a 60 W adapter, and a cable change on the same adapter changed measured draw by 34 W (36%). — [source](https://deploy.live/blog/running-local-llms-offline-on-a-ten-hour-flight/)
- The same author wrote powermonitor, a CLI reading CPU/GPU/ANE/adapter/battery telemetry, and lmstats for LM Studio throughput. — [source](https://deploy.live/blog/running-local-llms-offline-on-a-ten-hour-flight/)
- Ollama 0.19 (30 Mar 2026) switched from llama.cpp Metal to MLX; decode on one machine rose from 58 to 112 tok/s. — [source](https://www.xda-developers.com/run-local-llms-one-worlds-priciest-energy-markets/)
- A 5-runtime study on M2 Ultra (MLX v0.26, llama.cpp b5963, Ollama 0.10.1, MLC-LLM, PyTorch MPS) found MLX highest sustained throughput and MLC-LLM lowest TTFT; it reports no energy figures in its abstract. — [source](https://www.alphaxiv.org/abs/2511.5502v1)
- With long prompts, MLX prefill can leave the GPU at ~10 W for minutes (3.5 min before first token on a 40k-token prompt), and llama.cpp with flash attention finished an 8.5K-context task 9 s faster on an M1 Max. — [source](https://pub.towardsai.net/apples-mlx-runs-local-llms-3x-faster-than-llama-cpp-until-your-context-hits-40k-715ec441afbb)
- Raising the GPU wired-memory ceiling uses `sudo sysctl iogpu.wired_limit_mb=<MB>` (default ~65-75% of RAM); out of scope here except that spill to CPU raised CPU fan load. — [source](https://www.thinkdifferent.blog/blog/the-macbook-pro-setting-that-doubles-your-ai-performance/)
- Closed-lid clamshell use reduces cooling on MacBook Pro. — source: `asserted`
- No source found showing MLX consistently uses fewer joules per token than llama.cpp; the inference is that higher tok/s at similar watts implies lower joules/token. — source: `asserted`
- Mac Studio and Mac mini are "effectively unconstrained" for sustained decode. — source: `asserted`
