<!-- llms-explorer concept facts · https://llms-explorer.com/tree/ane-dispatch-and-iosurface-round-trip-overhead-i/ · pack 2026-10-05 · ~944 tokens -->

# ANE dispatch and IOSurface round-trip overhead in small decode graphs

> No primary per-dispatch or IOSurface figure exists for Core AI, M5 or M6; Core AI's effect on dispatch cost is unmeasured.

Parent: [Mac local LLMs: ANE and Core ML LLMs](https://llms-explorer.com/tree/mac-local-llms-ane-and-coreml-llm/) · 1 facets · 16 facts · page: https://llms-explorer.com/tree/ane-dispatch-and-iosurface-round-trip-overhead-i/

## Facts

- No primary per-dispatch or IOSurface figure exists for Core AI, M5 or M6; Core AI's effect on dispatch cost is unmeasured. — source: `asserted`
- Orion does not say whether the 2.3 ms is per layer, per program, or per token. — source: `asserted`
- Maderix measured a 256x256 FP16 matmul at 0.101 ms total, of which ~0.095 ms is XPC+IOKit overhead and ~0.006 ms compute. — [source](https://maderix.substack.com/p/inside-the-m4-apple-neural-engine-615)
- Maderix states a single matmul uses only ~30% of ANE peak and recommends chaining 16-64 ops in one MIL program; 94% utilization is reached at 32+ layer depth. — [source](https://maderix.substack.com/p/inside-the-m4-apple-neural-engine-615)
- Maderix rule: operations under ~1 ms are dominated by the 0.095 ms dispatch overhead and should be avoided. — [source](https://maderix.substack.com/p/inside-the-m4-apple-neural-engine-615)
- Maderix recommends ANE for prefill and large-batch work and SME (CPU) for single-token decode because SME has "zero dispatch overhead". — [source](https://maderix.substack.com/p/inside-the-m4-apple-neural-engine-615)
- Orion Table 9 gives CPU decode 3.48 ms/token (283 tok/s) against ANE full-forward decode 5.76 ms/token (170 tok/s) on GPT-2 124M, M4 Max. — [source](https://arxiv.org/html/2603.06728v1)
- Orion ANE cached prefill reaches 165 tok/s, below CPU decode, and first-call prefill is 12 tok/s because it includes ~1015 ms compilation of 24 programs. — [source](https://arxiv.org/html/2603.06728v1)
- Orion says the IOSurface overhead is amortized in prefill (longer sequences) and dominates only single-token decode. — [source](https://arxiv.org/html/2603.06728v1)
- Orion decode pads the token to a minimum sequence dimension of 16 (constraint #4), so each decode dispatch carries 16 positions for one real token. — [source](https://arxiv.org/html/2603.06728v1)
- Orion lists ANE advantages despite slower decode: hard power-gating at idle, GPU/CPU left free, and large-vocab softmax 33.8x faster; it concedes MLX/Metal on GPU has higher absolute throughput. — [source](https://arxiv.org/html/2603.06728v1)
- Maderix measured Core ML adding 2-4x overhead over direct _ANEClient on small ops, so latency-sensitive token decode through Core ML pays a larger fixed cost. — [source](https://maderix.substack.com/p/inside-the-m4-apple-neural-engine-615)
- Chunking mitigation in practice: ANEMLL Forge amortizes dispatch with 16 four-layer Core AI chunks and verify T=8 batches (speculative decoding) instead of one-token calls; per-chunk verification measured 6.0-7.7 ms and whole target verifier plans 106-141 ms in a bounded probe. — [source](https://github.com/Anemll/anemll-forge/blob/main/docs/SESSION_LESSONS.md)
- Forge attributes much of Core AI program overhead to bonded vs nonbonded compiled variants; undocumented compile mode 2 reduced it, and the analogous Core ML mode caused timeouts. — [source](https://github.com/Anemll/anemll-forge/blob/main/docs/SESSION_LESSONS.md)
- No source read gives a per-dispatch or IOSurface round-trip measurement for Core AI (macOS 27); whether Core AI lowers dispatch cost is unestablished. — source: `asserted`
- Inference: batching tokens per dispatch (speculative verify, multi-token) is the shared lever, since compute at 0.006 ms is negligible against 0.095 ms fixed cost. — source: `asserted`
