ANE SRAM 32 MB cliff and graph-depth utilization
Parent: Mac local LLMs: ANE and Core ML LLMs · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
No source measures the cliff for 1x1 conv (the layout LLM layers use), for decode-shaped (seq 16) tensors, or on M1-M3/M5.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- No source measures the cliff for 1x1 conv (the layout LLM layers use), for decode-shaped (seq 16) tensors, or on M1-M3/M5. [source]
- No source states whether compiled weights count against the 32 MB working set. [source]
- maderix computed the matmul working set as 24 MB at 2048x2048 and 96 MB at 4096x4096 in FP16 (3 matrices), about 3x the inferred SRAM. [source]
- maderix measured 4.0 TFLOPS at 4096x4096 against 5.7 TFLOPS at 2048x2048 (about 30% lower), and placed the SRAM at roughly 32 MB from the 24 MB (fast) to 96 MB (slow) gap. [source]
- maderix says the SRAM transition degrades gradually, suggesting a cache-like hierarchy rather than a hard scratchpad. [source]
- maderix's rule "Stay under 32 MB" says to keep per-tensor footprint within SRAM because spilling to DRAM kills throughput. [source]
- maderix's 94% utilization at 32+ layer depth is described as 100% of a theoretical 19 TFLOPS peak (16 cores x ~1.2 TFLOPS/core), so single-op 5.7 TFLOPS is about 30% of peak. [source]
- Orion's compiler has an "SRAM annotation" pass that estimates working-set size against the 32 MB budget and only emits a warning when exceeded (about 30% degradation); it does not split or reschedule graphs. [source]
- Orion's footnote attributes the ~30% over-budget drop to maderix (2026c), so there is a single measurement behind both sources. [source]
- maderix's repo probes SRAM with sram_bench.m (bandwidth) and sram_probe.m (size/layout), and reports INT8 activation caching between layers in L2 SRAM gives 1.88x throughput. [source]
- maderix's repo reports ANE training utilization of only ~5-9% of peak despite deep graphs. [source]
- Inference for LLM chunk sizing: the 32 MB limit applies to per-op activation/tensor working set; ANEMLL chunks (default max 950 MB) are sized by a per-file/cache limit, a different constraint, so the two must not be conflated. [source]
- Mitigation for decode: chain a whole layer or several layers per program (ANEMLL four-layer chunks) and keep per-op tensors small, since over-budget ops lose ~30% and sub-1 ms ops lose to dispatch. [source]
Children
- No children recorded.