Core ML and Apple Neural Engine for LLMs

Parent: Mac local LLMs: ANE and Core ML LLMs · Published reference · snapshot 2026-10-05

↓ Facts as markdownall context files

Core ML LLM path: export with torch.jit.trace (torch.export beta), then three optimizations: fused SDPA op, KV cache as Core ML "state" (macOS Sequoia+) with flexible-shape inputs, and block-wise int4 weight quantization (block 32). Apple's reference run targeted the GPU, not the ANE, because dec...

These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.

Facts

Corrections and disagreements

Children

← the whole tree · 3D view· how to read this page