Thermal, power and battery behavior during local inference
Parent: Mac local LLMs: Speed, bandwidth and prefill · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Decode is memory-bandwidth-bound: each token streams all active weights, so dense models draw more watts and give fewer tok/s than MoE models of larger parameter count (M3 Ultra measurements: 27B dense 138 W / 21.5 tok/s vs 120B MoE 94 W / 74 tok/s).
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Decode is memory-bandwidth-bound: each token streams all active weights, so dense models draw more watts and give fewer tok/s than MoE models of larger parameter count (M3 Ultra measurements: 27B dense 138 W / 21.5 tok/s vs 120B MoE 94 W / 74 tok/s). [source]
- Energy per token = watts / tokens-per-second; electricity cost per 1M output tokens followed that ratio ($0.554 dense 27B vs $0.063-0.109 for the others at $0.31/kWh). [source]
- Every chassis has a sustained thermal budget below its burst budget; when die temperature reaches target, clocks drop. A fanless Air sheds heat only through the case. [source]
- Apple's Energy Modes: Automatic default; Low Power reduces power and fan noise; High Power lets fans run faster so the chip may sustain performance in very intensive workloads. Modes can be set separately for battery and adapter. [source]
- Inference is bursty in interactive use (duty cycle roughly 5-10% by one author), so battery cost depends on workload shape far more than peak wattage. [source]
- High Power Mode started on Max-chip Macs (2021-2023 MacBook Pro Max), reached M4 Pro MacBook Pro and Mac mini in Nov 2024, and per Apple's page now covers 2026 Mac Studio, Pro/Max MacBook Pro 2024+ and Pro-chip Mac mini. [source]
- Ollama 0.19 (30 Mar 2026) replaced its llama.cpp Metal backend with MLX, so "MLX vs llama.cpp" now partly means "Ollama before/after". [source]
- Throttling on the Air arrives after minutes of continuous GPU load; first-minute benchmarks are burst numbers. [source]
- Slowness from the first token is a memory/placement problem, not thermal. [source]
- On battery macOS may cap sustained performance independent of temperature (single-source, llmconfigurator). [source]
- Clamshell or soft-surface use removes heat paths (single-source blogs). [source]
- Adapter wattage and cable can cap power: one M5 Max user drained battery while on a 60 W adapter and measured a 34 W (36%) gap between two cables on the same adapter. [source]
- A resident large model carries standing overhead (~21 W plateau on the M3 Ultra for big models, single-source) and, on battery, memory pressure drain (~1.5-2%/hr above baseline for a resident 14B, single-source). [source]
- Prefill on long context can sit at low GPU power for minutes (about 10 W for 3.5 minutes on a 40k-token prompt), so tok/s headlines hide wall-clock cost. [source]
- Low Power Mode cost. Side A (thinkdifferent.blog, Jakub Jirak): Low Power "will absolutely strangle inference", people report "half the expected speed". Side B (same author, Medium battery post): Low Power slows generation ~20-30% (8B from ~30 to ~21 tok/s) and cuts drain more than that. Side C (Newport, Medium): M2 Max 8.02 tok/s at low power vs 16 at high power (his own "2.5x" label does not match 16/8.02). Same author contradicts himself; no source is controlled. Treat "halves" as an upper-range anecdote. [source]
- High Power Mode benefit. thinkdifferent says it avoids a 15-25% sag; Apple's page claims only possible gains in very intensive workloads and cites video/3D examples, no LLM numbers. MacRumors forum comment: "performance is a bit of a wash, fan noise is considerably increased" (single comment). The 15-25% figure has no measurement behind it. [source]
- Air sustained throttling severity. willitrunai: Air M4 sustained 60-70% of peak after ~8-12 min. modelfit: throttles after 20-30 min, figures labeled estimates. llmconfigurator: "halves", but its own table is labeled not measured. MindStudio: M5 Air CPU merge-sort stress test lost only ~6% first-to-last iteration (CPU, not GPU decode; M4/M3 ~5%, M2 ~3%). No GPU-decode sustained curve with a stated method was found. [source]
- Battery: Jirak M1 Pro says 8B costs ~11%/hr in bursty use and ~38%/hr when hammered; deploy.live M5 Max under sustained load drained ~1%/min. [source]
- Measured sustained tok/s vs time curves for M-series Air GPU decode (first-party or reproducible). [source]
- Controlled tok/s and joules/token for Low vs Automatic vs High on one machine. [source]
- MLX vs llama.cpp joules per token with wall-metered power; found only tooling (TokenWatt) and throughput comparisons, no head-to-head energy dataset. [source]
- Whether Mac mini base-chip and M4 Pro minis hold clocks under hours of decode (llmconfigurator says "effectively unconstrained" but marks it not measured). [source]
- Reddit threads on these topics could not be fetched. [source]
- Apple Support documents three Energy Modes: Low Power, Automatic (default), High Power; settable separately for On battery and On power adapter. [source]
- Low Power Mode gives reduced power consumption and reduced fan noise on supported models. [source]
- High Power Mode lets fans run at higher speeds; extra cooling "may allow" higher performance in very intensive workloads, with more fan noise. [source]
- Apple's High Power Mode support list covers Mac Studio introduced in 2026, Mac mini 2024+ with Pro chip, MacBook Pro 2024+ with Pro or Max chip, and 2021-2023 MacBook Pro with Max chip; MacBook Air is not on it. [source]
- Apple recommends the 96 W USB-C adapter for High Power Mode while charging on the 14-inch MacBook Pro with M4 Pro or M5 Pro. [source]
- In macOS Sequoia 15.1 and later, Energy Modes can also be selected from Control Center. [source]
- High Power Mode first reached Pro-chip Macs (14/16-inch MacBook Pro and Mac mini M4 Pro) in Nov 2024; previously Max-only. [source]
- A MacRumors forum commenter said performance under High Power is "a bit of a wash" while fan noise is considerably increased. [source]
- Claim "High Power Mode avoids a 15-25% sag" originates from one blog (Jakub Jirak, thinkdifferent.blog, 3 Jun 2026) with no benchmark shown. [source]
- Claim "Low Power Mode halves speed" in that blog is hearsay ("I've seen people benchmark..."); the same author's other post measures a 20-30% drop (30 to 21 tok/s on 8B, M1 Pro). [source]
- Newport reports M2 Max MacBook Pro running Gemma-3-27B Q4 at 8.02 tok/s in low power and 16 tok/s in high power. [source]
- The same blog says a 40-50k-token context made inference about 10x slower than short prompts. [source]
- willitrunai lists Air M4 passive cooling with "~8-12 min before throttle" and sustained throttle to ~60-70% of peak; Llama 3.1 8B 42 tok/s burst vs 25-30 sustained. [source]
- willitrunai's Air-vs-Pro burst tok/s are labeled approximate and no method is given. [source]
- modelfit says the Air M4 throttles after 20-30 minutes of intensive GPU use and that its tok/s are estimates from chip bandwidth, not lab benchmarks. [source]
- The Air M4 and base-chip MacBook Pro M4 have the same 120 GB/s bandwidth; the M4 Pro has 273 GB/s (per willitrunai). [source]
- MindStudio measured 142 GB/s STREAM bandwidth on M5 Air vs 113 on M4 Air. [source]
- MindStudio's all-core CPU stress test (merge sort of 1B integers, 5 rounds) showed first-to-last drop of ~6% on M5 Air, ~5% on M4 and M3, ~3% on M2; it did not test GPU LLM decode. [source]
- llmconfigurator shows a user trace of 42.6 tok/s at minute 1 and 19.8 at minute 10 and attributes it to thermal limits; its chassis-expectation table is explicitly "our characterisation, not measured" with no sustained-load rows in its dataset. [source]
- llmconfigurator asserts macOS caps sustained performance on battery independent of temperature. [source]
- llmconfigurator recommends keeping the lid open, off soft surfaces, and using a smaller model or shorter context for long jobs because lower work per token lowers power and raises the plateau. [source]
- Monitoring command: `sudo powermetrics --samplers smc,gpu_power -i 1000` reports GPU power and die temperature; a power peak then lower plateau with flat temperature indicates throttling. [source]
- Check power policy with `pmset -g batt` and `pmset -g | grep -i lowpowermode`. [source]
- powermetrics supports sampler selection (-s, "all"/"default" groups), -i interval in ms, plist output, --show-plimits and --show-pstates (hardware-dependent), and needs root. [source]
- asitop is a Python TUI for Apple Silicon CPU/GPU/ANE usage built on powermetrics; last commit Jan 2023 ("Updated for Ventura"), 4.6k stars. [source]
- macmon is a sudoless monitor reading CPU/GPU/ANE power, temperatures and memory through a private macOS API (IOReport), with JSON pipe output; active in 2026 (commits Jul 2026). [source]
- powermetrics can emit plist, which is useful for feeding monitoring systems. [source]
- iStat Menus plus btop were used as lightweight monitors in a two-Mac Ollama comparison (M1 mini 16.5 tok/s vs M5 Pro 62.6 tok/s, Gemma 4). [source]
- TokenWatt (open-source proxy) estimates per-request marginal energy from IOReport SoC rails with a rolling idle baseline, calibrated to a Shelly Plug wall meter (stated error +/-2.6 to 4.5%). [source]
- M3 Ultra Mac Studio (96 GB), sustained decode, wall-calibrated: Qwen3.5-4B 133.8 tok/s $0.063/1M out; Qwen3.6-35B-A3B 76.0 tok/s $0.087; Qwen3-Coder-Next ~80B 65.0 tok/s $0.103; gpt-oss-120b 74.0 tok/s $0.109 (94 W); dense Qwen3.6-27B 21.5 tok/s 138 W $0.554 (at $0.31/kWh). [source]
- In that author's real 30-day traffic (uncalibrated estimate tier), the dense 27B cost roughly 10x more per token than MoE models, and sub-saturation use raises per-token cost. [source]
- Holding a big model resident adds standing overhead rising with size and flattening near 21 W on that machine. [source]
- On M1 Pro 14" (70 Wh), 8B inference package power peaks ~18-22 W and idles under 2 W between requests. [source]
- Same source, Ollama Q4_K_M, battery drain per hour (active/hammer): 3B ~7%/~26%; 8B ~11%/~38%; 14B ~15%/~47%; 32B ~22%/60%+; baseline without AI 5-6%/hr. [source]
- Same source recommends OLLAMA_KEEP_ALIVE=10m on battery (reload cost 8B ~4 s, 14B ~7 s) and measures battery with a pmset -g batt log every 5 minutes. [source]
- Single author, single machine (M1 Pro); the figures are self-measured with no method files. [source]
- A 10-hour flight on an M5 Max MacBook Pro 16" 128 GB: roughly 1% battery per minute under sustained load, battery drained even on a 60 W adapter, and a cable change on the same adapter changed measured draw by 34 W (36%). [source]
- The same author wrote powermonitor, a CLI reading CPU/GPU/ANE/adapter/battery telemetry, and lmstats for LM Studio throughput. [source]
- Ollama 0.19 (30 Mar 2026) switched from llama.cpp Metal to MLX; decode on one machine rose from 58 to 112 tok/s. [source]
- A 5-runtime study on M2 Ultra (MLX v0.26, llama.cpp b5963, Ollama 0.10.1, MLC-LLM, PyTorch MPS) found MLX highest sustained throughput and MLC-LLM lowest TTFT; it reports no energy figures in its abstract. [source]
- With long prompts, MLX prefill can leave the GPU at ~10 W for minutes (3.5 min before first token on a 40k-token prompt), and llama.cpp with flash attention finished an 8.5K-context task 9 s faster on an M1 Max. [source]
- Raising the GPU wired-memory ceiling uses `sudo sysctl iogpu.wired_limit_mb=<MB>` (default ~65-75% of RAM); out of scope here except that spill to CPU raised CPU fan load. [source]
- Closed-lid clamshell use reduces cooling on MacBook Pro. [source]
- No source found showing MLX consistently uses fewer joules per token than llama.cpp; the inference is that higher tok/s at similar watts implies lower joules/token. [source]
- Mac Studio and Mac mini are "effectively unconstrained" for sustained decode. [source]
Children
- No children recorded.