ANE dispatch and IOSurface round-trip overhead in small decode graphs
Parent: Mac local LLMs: ANE and Core ML LLMs · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
No primary per-dispatch or IOSurface figure exists for Core AI, M5 or M6; Core AI's effect on dispatch cost is unmeasured.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- No primary per-dispatch or IOSurface figure exists for Core AI, M5 or M6; Core AI's effect on dispatch cost is unmeasured. [source]
- Orion does not say whether the 2.3 ms is per layer, per program, or per token. [source]
- Maderix measured a 256x256 FP16 matmul at 0.101 ms total, of which ~0.095 ms is XPC+IOKit overhead and ~0.006 ms compute. [source]
- Maderix states a single matmul uses only ~30% of ANE peak and recommends chaining 16-64 ops in one MIL program; 94% utilization is reached at 32+ layer depth. [source]
- Maderix rule: operations under ~1 ms are dominated by the 0.095 ms dispatch overhead and should be avoided. [source]
- Maderix recommends ANE for prefill and large-batch work and SME (CPU) for single-token decode because SME has "zero dispatch overhead". [source]
- Orion Table 9 gives CPU decode 3.48 ms/token (283 tok/s) against ANE full-forward decode 5.76 ms/token (170 tok/s) on GPT-2 124M, M4 Max. [source]
- Orion ANE cached prefill reaches 165 tok/s, below CPU decode, and first-call prefill is 12 tok/s because it includes ~1015 ms compilation of 24 programs. [source]
- Orion says the IOSurface overhead is amortized in prefill (longer sequences) and dominates only single-token decode. [source]
- Orion decode pads the token to a minimum sequence dimension of 16 (constraint #4), so each decode dispatch carries 16 positions for one real token. [source]
- Orion lists ANE advantages despite slower decode: hard power-gating at idle, GPU/CPU left free, and large-vocab softmax 33.8x faster; it concedes MLX/Metal on GPU has higher absolute throughput. [source]
- Maderix measured Core ML adding 2-4x overhead over direct _ANEClient on small ops, so latency-sensitive token decode through Core ML pays a larger fixed cost. [source]
- Chunking mitigation in practice: ANEMLL Forge amortizes dispatch with 16 four-layer Core AI chunks and verify T=8 batches (speculative decoding) instead of one-token calls; per-chunk verification measured 6.0-7.7 ms and whole target verifier plans 106-141 ms in a bounded probe. [source]
- Forge attributes much of Core AI program overhead to bonded vs nonbonded compiled variants; undocumented compile mode 2 reduced it, and the analogous Core ML mode caused timeouts. [source]
- No source read gives a per-dispatch or IOSurface round-trip measurement for Core AI (macOS 27); whether Core AI lowers dispatch cost is unestablished. [source]
- Inference: batching tokens per dispatch (speculative verify, multi-token) is the shared lever, since compute at 0.006 ms is negligible against 0.095 ms fixed cost. [source]
Children
- No children recorded.