<!-- llms-explorer concept facts · https://llms-explorer.com/tree/ane-sram-32-mb-cliff-and-graph-depth-utilization/ · pack 2026-10-05 · ~765 tokens -->

# ANE SRAM 32 MB cliff and graph-depth utilization

> No source measures the cliff for 1x1 conv (the layout LLM layers use), for decode-shaped (seq 16) tensors, or on M1-M3/M5.

Parent: [Mac local LLMs: ANE and Core ML LLMs](https://llms-explorer.com/tree/mac-local-llms-ane-and-coreml-llm/) · 1 facets · 13 facts · page: https://llms-explorer.com/tree/ane-sram-32-mb-cliff-and-graph-depth-utilization/

## Facts

- No source measures the cliff for 1x1 conv (the layout LLM layers use), for decode-shaped (seq 16) tensors, or on M1-M3/M5. — source: `asserted`
- No source states whether compiled weights count against the 32 MB working set. — source: `asserted`
- maderix computed the matmul working set as 24 MB at 2048x2048 and 96 MB at 4096x4096 in FP16 (3 matrices), about 3x the inferred SRAM. — [source](https://maderix.substack.com/p/inside-the-m4-apple-neural-engine-615)
- maderix measured 4.0 TFLOPS at 4096x4096 against 5.7 TFLOPS at 2048x2048 (about 30% lower), and placed the SRAM at roughly 32 MB from the 24 MB (fast) to 96 MB (slow) gap. — [source](https://maderix.substack.com/p/inside-the-m4-apple-neural-engine-615)
- maderix says the SRAM transition degrades gradually, suggesting a cache-like hierarchy rather than a hard scratchpad. — [source](https://maderix.substack.com/p/inside-the-m4-apple-neural-engine-615)
- maderix's rule "Stay under 32 MB" says to keep per-tensor footprint within SRAM because spilling to DRAM kills throughput. — [source](https://maderix.substack.com/p/inside-the-m4-apple-neural-engine-615)
- maderix's 94% utilization at 32+ layer depth is described as 100% of a theoretical 19 TFLOPS peak (16 cores x ~1.2 TFLOPS/core), so single-op 5.7 TFLOPS is about 30% of peak. — [source](https://maderix.substack.com/p/inside-the-m4-apple-neural-engine-615)
- Orion's compiler has an "SRAM annotation" pass that estimates working-set size against the 32 MB budget and only emits a warning when exceeded (about 30% degradation); it does not split or reschedule graphs. — [source](https://arxiv.org/html/2603.06728v1)
- Orion's footnote attributes the ~30% over-budget drop to maderix (2026c), so there is a single measurement behind both sources. — [source](https://arxiv.org/html/2603.06728v1)
- maderix's repo probes SRAM with sram_bench.m (bandwidth) and sram_probe.m (size/layout), and reports INT8 activation caching between layers in L2 SRAM gives 1.88x throughput. — [source](https://github.com/maderix/ANE)
- maderix's repo reports ANE training utilization of only ~5-9% of peak despite deep graphs. — [source](https://github.com/maderix/ANE)
- Inference for LLM chunk sizing: the 32 MB limit applies to per-op activation/tensor working set; ANEMLL chunks (default max 950 MB) are sized by a per-file/cache limit, a different constraint, so the two must not be conflated. — source: `asserted`
- Mitigation for decode: chain a whole layer or several layers per program (ANEMLL four-layer chunks) and keep per-op tensors small, since over-budget ops lose ~30% and sub-1 ms ops lose to dispatch. — source: `asserted`
