<!-- llms-explorer concept facts · https://llms-explorer.com/tree/sustained-load-thermal-throttling-ane-versus-gpu/ · pack 2026-10-05 · ~2282 tokens -->

# Sustained-load thermal throttling ANE versus GPU decode

> Power: the ANE path draws less package power (M4 Max whole-system powermetrics: 12.7 W for CoreML-LLM versus 24.7 W MLX-Swift, 24.5 W llama.cpp), so the SoC has less reason to throttle it; on iPhone, power counters are not exposed, so the ANE power figure comes from a Mac measurement and is only ...

Parent: [Mac local LLMs: Speed, bandwidth and prefill](https://llms-explorer.com/tree/mac-local-llms-speed-bandwidth-and-prefill/) · 1 facets · 33 facts · page: https://llms-explorer.com/tree/sustained-load-thermal-throttling-ane-versus-gpu/

## Facts

- Power: the ANE path draws less package power (M4 Max whole-system powermetrics: 12.7 W for CoreML-LLM versus 24.7 W MLX-Swift, 24.5 W llama.cpp), so the SoC has less reason to throttle it; on iPhone, power counters are not exposed, so the ANE power figure comes from a Mac measurement and is only assumed to carry over. — source: `asserted`
- Hardware: the ANE has independent DVFS and hard power gating to 0 mW idle, but under continuous dispatch it stays pinned at a high DVFS state (CoreML-LLM measured it busy 97% of each step); a response of 200 tokens at 14 tok/s keeps it busy 14 s. — source: `asserted`
- Heat path: CoreML-LLM's own hypotheses are small die area (ANE about 10 mm2 giving high heat-flux density), the ANE sitting off the vapor chamber centroid, time-integrated skin temperature, and LPDDR reads (about 42 GB/s sustained at 14 tok/s on a 3 GB model) heating regardless of engine. — source: `asserted`
- Confounder: the ANE bundle uses sliding-window attention (bounded context), which keeps per-token work constant while GPU runtimes' attention cost grows with generated length. — source: `asserted`
- 2026-04-18 CoreML-LLM refutes "CPU orchestration saturates a P-core" with E4B device data (CPU active 3.0 ms, 4% of the step). — source: `asserted`
- 2026-05-02 the bench repo starts; by 2026-09 it has a thermal-state guard (only thermal-nominal starts count, added to the leaderboard 2026-09-05), energy per token via battery-delta (iPhone) and powermetrics (Mac), and interleaved arms with 45 s cooldowns (rule 11, 2026-08). — source: `asserted`
- 2026-07 an MLX row of 112 tok/s is retracted as Debug-contaminated (MLX 178.8 warm on iPhone 17 Pro), and the iPhone energy ranking was re-run after a fair-start versus nominal-start artifact. — source: `asserted`
- Order effects dominate on Mac GPU arms: in a blocked-order Core AI vs MLX vs ExecuTorch run on M4 Max, ExecuTorch decayed 23.5 to 17.4 tok/s inside its own block and Core AI read 15.6-20.7 tok/s on the already-heated GPU against a true value of about 27.4 tok/s. — source: `asserted`
- Thermal start state changes results: the LiteRT sustained run started at `fair` thermal while others started `nominal`; a newer Core AI capture "started hot" and was excluded by the guard. — source: `asserted`
- Energy ranking flips with method: on Mac, the ANE has the lowest watts but the worst joules per token (0.48 vs 0.24); on iPhone in a battery-delta test a LiteRT wNa8o8 build beat MLX on energy. — source: `asserted`
- Skin-temperature perception differs from throttling: a user found the ANE build hotter to hold than a GPU LiteRT-LM app even at lower watts. — source: `asserted`
- Does the ANE throttle? The bench README text says the ANE "barely moves" and "holds its rate"; its own table shows CoreML/ANE retains 67% of burst (33 to 22 tok/s over 10 min), i.e. a one-third loss, and the text elsewhere says it "retains about 65%". The ANE loses less than the GPU runtimes (MLX 38%, LiteRT 48%) but is not flat. — source: `asserted`
- Which runtime throttles least on iPhone? In the Gemma 4 E2B 10-minute test: ANE 67%, LiteRT-LM GPU 48%, MLX 38%. In a separate energy campaign (600 s, deep protocol): LiteRT wNa8o8 76%, OptiQ 67%, MLX-PTQ 64%, Cactus 57%, Core AI 56%, llama.cpp 54%. The LiteRT build in the second campaign (QAT, w4a8) retains more than the ANE build in the first; the two tests use different models and builds. — source: `asserted`
- Is the ANE more energy efficient? Whole-package watts say yes (about half); joules per token on M4 Max say no (0.48 J/token vs 0.24 MLX), because the ANE decodes at 32 vs 185 tok/s; iPhone battery-delta numbers for the ANE build are not published in the quoted pages. — source: `asserted`
- Perceived heat: the repo's own investigation says the ANE is the heat source (busy 97% of step), yet its README frames the ANE as the cool, sustained option. — source: `asserted`
- Sustained ANE decode on a Mac: no throttling, power or thermal-state series exists for ANEMLL, CoreML-LLM or Forge on M-series (Forge says "no power/thermal conclusion is included"). — source: `asserted`
- How much of the 67% vs 38% gap is the ANE's lower power versus the sliding-window attention confounder? Not isolated. — source: `asserted`
- Does the ANE decode rate throttle on a fanless MacBook Air under 30+ minutes? Not measured. — source: `asserted`
- The sustained-throttle test is 600 s of continuous generation (tg128) from a cold `nominal` start, unplugged, on iPhone 17 Pro with Gemma 4 E2B 4-bit, decode rate from a rolling window; the LiteRT-LM run has no output cap and started at `fair` thermal. — [source](https://github.com/john-rocky/apple-silicon-llm-bench)
- In that test MLX crosses the 50%-lost line within about 60 s, LiteRT-LM holds about 53% at 1 minute and crosses 50% near 4 minutes, and the ANE retains about 65-67% at 10 minutes. — [source](https://github.com/john-rocky/apple-silicon-llm-bench)
- The bench README says CoreML-LLM uses sliding-window attention (bounded context), which is part of why its decode stays flat. — [source](https://github.com/john-rocky/apple-silicon-llm-bench)
- The bench README states iOS does not expose power counters to third-party apps, so the ANE power figure (about half of package power) was measured on Mac via powermetrics. — [source](https://github.com/john-rocky/apple-silicon-llm-bench)
- A separate bench campaign (600 s, deep protocol, unplugged iPhone) reports sustained-to-burst retention of LiteRT wNa8o8 76%, OptiQ 67%, MLX-PTQ 64%, Cactus 57%, Core AI 56% and llama.cpp 54%, and energy of 0.122 J/token (LiteRT wNa8o8) vs 0.151 (MLX-PTQ). — [source](https://github.com/john-rocky/apple-silicon-llm-bench)
- The bench re-measured iPhone energy in one nominal-start session and found MLX 0.164 J/token below Cactus (0.164-0.183), llama.cpp 0.213, LiteRT 0.232 and OptiQ 0.415 (n=1), saying the earlier fair-start ranking does not survive nominal. — [source](https://github.com/john-rocky/apple-silicon-llm-bench)
- llama.cpp draws about 9.8 W on the iPhone and is the energy floor at 0.483 J/token, keeping 54% of burst under sustained load. — [source](https://github.com/john-rocky/apple-silicon-llm-bench)
- On M4 Max Gemma 4 E2B energy per 512-token run is 123.0 J for MLX-Swift, 126.3 J for llama.cpp and 244.9 J for CoreML-LLM, with the bench noting whole-system powermetrics includes the idle baseline. — [source](https://github.com/john-rocky/apple-silicon-llm-bench)
- A later Mac decode-window energy table gives MLX PTQ 4-bit 0.090 J/token at 14.6 W and 177.8 tok/s, LiteRT 0.154, llama.cpp 0.170 at 20.5 W, and Core AI int4 about 0.33 J/token, and concludes MLX owns the Mac energy Pareto. — [source](https://github.com/john-rocky/apple-silicon-llm-bench)
- The bench's rule 11 interleaves arms after a blocked-order run decayed ExecuTorch from 23.5 to 17.4 tok/s inside its own block and made Core AI read 15.6-20.7 against about 27.4 on a cool GPU, with 45 s cooldowns between arms. — [source](https://github.com/john-rocky/apple-silicon-llm-bench)
- CoreML-LLM measured an E4B step of 70.6 ms with ANE wait 68.7 ms (97%) and CPU 3.0 ms (4%), and says what feels hot is the ANE pinned at high DVFS for about 14 s per response, not CPU orchestration. — [source](https://raw.githubusercontent.com/john-rocky/CoreML-LLM/main/docs/CPU_BOTTLENECK_INVESTIGATION.md)
- CoreML-LLM's mitigation menu ships `MLOptimizationHints.specializationStrategy = .fastPrediction` (default on, iOS 18+), rejects thermalState-keyed inter-dispatch sleeps as a throughput trade, rejects async ANE dispatch (more concurrency means more sustained heat) and rejects private `_ANEClient` `qos:21` for App Store risk. — [source](https://raw.githubusercontent.com/john-rocky/CoreML-LLM/main/docs/CPU_BOTTLENECK_INVESTIGATION.md)
- CoreML-LLM's untried levers are INT8 KV cache (about 50% bandwidth, hence DRAM heat) and W2 QAT weights, with post-training W2/W3 producing gibberish. — [source](https://raw.githubusercontent.com/john-rocky/CoreML-LLM/main/docs/CPU_BOTTLENECK_INVESTIGATION.md)
- maderix reports the ANE has its own DVFS channels and hard power gating to 0 mW idle, and CoreML-LLM cites this to argue the ANE drops to 0 mW between chunks yet stays at high DVFS under continuous dispatch. — [source](https://maderix.substack.com/p/inside-the-m4-apple-neural-engine)
- Forge states no power or thermal conclusion is included in its 27B ANE measurements. — [source](https://huggingface.co/anemll/anemll-forge-qwen3.8-27B)
