Orion ANE constraint catalog and direct-ANE programming
Parent: Mac local LLMs: ANE and Core ML LLMs · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Each constraint is recorded as constraint, symptom, workaround and source tag (P prior work by maderix or hollance, O discovered in Orion, starred if confirmed in maderix/ANEgpt code). Four categories: MIL IR restrictions (1, 6, 10, 12, 13, 16), memory and I/O layout (2, 3, 4, 8, 9, 11, 18, 19, 2...
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Each constraint is recorded as constraint, symptom, workaround and source tag (P prior work by maderix or hollance, O discovered in Orion, starred if confirmed in maderix/ANEgpt code). Four categories: MIL IR restrictions (1, 6, 10, 12, 13, 16), memory and I/O layout (2, 3, 4, 8, 9, 11, 18, 19, 20), compilation limits (5, 7, 14, 15), performance characteristics (16, 17). [source]
- Orion's compiler enforces them: its constraint-validation pass rejects banned ops (concat), too-small tensors, nil weight dictionaries and dead output references before calling `ANECCompile()`; a separate uniform-output-padding pass applies constraint 2. [source]
- The catalog is a property of the direct MIL-to-E5 path; the paper itself says many constraints "are likely artifacts of the ANE's microarchitecture and compiler, not fundamental". [source]
- 2022 hollance's neural-engine notes and 2026 maderix supply constraints 5, 6, 7, 15 and 17; Orion adds 14 more MIL/layout restrictions; constraints 18-20 were found late, during LoRA work. [source]
- 2026-03-06 v2.0 commit ("delta compilation, LoRA hot-swap") records the 3 LoRA constraints; the stated speedups were re-measured and corrected from 7.8x/3.6x to 8.5x/3.8x. [source]
- 2026-08-23 (latest commit) the repo adds a "blob budget" probe and extends a "#16" penalty rule, showing the catalog has diverged from the paper's numbering. [source]
- Immediate crashes (not silent): MIL text passed as `NSString` instead of `NSData` (9), weight dictionary `nil` instead of `@{}` (11). [source]
- Silent corruption: BLOBFILE offset is `uint64(64)` from the chunk header, not 128 or the file start (8); over-allocated input surfaces are read as a flat packed [1,C,1,S] buffer from byte 0 (20), so the surface's nominal shape is ignored. [source]
- Compile rejections: matmul transpose flags must be named const nodes (12); output variables must reference nodes that survive dead-code elimination (14). [source]
- Repo-side probes after the paper report a per-program weight "blob budget": each inline scalar fed into an elementwise op (add, mul, pow) costs budget once per program, and a norm already costs one. [source]
- Concat. Orion constraint 1: concat is rejected by the ANE compiler. CoreML-LLM's ANERMSNorm uses `cat([x,-x])` and its Gemma 4 E2B graphs run 99.78% on ANE through Core ML, and maderix's training kernels expose intermediates through concat outputs. Side by side: direct MIL path vs Core ML-compiled graphs; not reconciled. [source]
- Large-channel convs. Orion constraint 16 rejects 32K-channel convolutions and keeps the logits projection on CPU ("wte is 73 MB, too large for ANE SRAM" in the README); Orion's own Table 11 times the classifier forward on ANE (10.2x), maderix's repo has an `ane_classifier.h` with a 32K conv on ANE, and CoreML-LLM runs a 262,144-wide Conv2d on ANE via Core ML. Four sources, three behaviors. [source]
- Failure mode of nil/wrong-type inputs. Orion: nil weight dict and NSString MIL crash immediately. maderix: a string MIL "fails silently", and his bridge commit says fixing nil to `@{}` "prevents silent compile failure". Same inputs, different symptom reports. [source]
- IOSurface minimum. Orion's table gives a minimum of about 49 KB per surface (pad seq to at least 16), but its own worked example pads [1,768,1,1] (3,072 bytes) to [1,768,1,16], which is 24,576 bytes in fp16 (49,152 only if fp32). The rule and the example disagree by 2x. [source]
- SiLU lowering. Forge reports native SiLU is inaccurate on ANE and the compiler fuses x*sigmoid(x) back to it, so its recipe uses the tanh form; Orion's blob-budget probe says the tanh form costs one unit of blob ceiling (15 vs 16) while sigmoid form does not. Accuracy and budget pull in opposite directions; no source tests both on one stack. [source]
- What is the semantics and size of the "blob budget" (what is the ceiling counting, and is it per program or per process)? Only the commit message describes it. [source]
- Do the constraints hold on M1-M3, M5, M6 and macOS 27? The paper validates on M4 Max only. [source]
- Are rule 4 (49 KB) and its example reconcilable (fp16 vs fp32 I/O)? [source]
- Orion's catalog has 20 constraints; Table 3 tags five as prior work (5, 6*, 7, 15, 17) and the rest as discovered in Orion, while the text says six were first documented by maderix or hollance (5, 7, 15, 17 and partly 6). [source]
- Constraint 8 says the BLOBFILE offset is uint64(64) and not 128, with garbage weights as the symptom, and the text specifies the offset is 64 bytes from the chunk header, not from the file start. [source]
- Constraint 9 says MIL text must be NSData and not NSString, with an immediate crash, and constraint 11 says the weight dictionary must be `@{}` not nil, also an immediate crash. [source]
- Constraint 12 says matmul transpose flags need named const nodes or MIL rejects the program, and constraint 14 says output variables must reference live post-optimization nodes so references must be updated after dead-code elimination. [source]
- Constraint 15 puts exec() restart overhead at about 50 ms, and constraint 17 says a 1x1 convolution is about 3x faster than matmul (first measured by maderix and noted by hollance). [source]
- Constraint 20 says over-allocated multi-input surfaces are read by the ANE as a packed [1,C,1,S] buffer from byte 0 regardless of nominal dimensions, so padded data must be written at the buffer start. [source]
- The paper's worked example for constraint 4 pads a [1,768,1,1] tensor (3,072 bytes in fp16) to [1,768,1,16] (24,576 bytes) although the table states a roughly 49 KB minimum. [source]
- Orion states that many of its constraints are likely artifacts of the ANE's microarchitecture and compiler rather than fundamental NPU limits, and that results were validated on M4 Max only with INT8/INT4 quantization not yet supported. [source]
- Orion's Table 11 times the classifier forward at 10.77 ms on CPU and 1.06 ms on ANE (10.2x) and softmax over 32,000 at 81.11 ms versus 2.40 ms, while RMSNorm backward is about equal (0.18 vs 0.21 ms). [source]
- Orion's README says the GPT-2 logits projection stays on the CPU because the 73 MB wte is too large for ANE SRAM. [source]
- A latest-commit message (2026-08-23, co-authored by Claude Opus 5) reports a probe in which SiLU via sigmoid gives a blob ceiling of 16 and SiLU via tanh gives 15, because the tanh identity introduces scalar mul and add operands that trigger a "#16 penalty". [source]
- The same message extends the penalty: mul(scalar) costs the same as add(scalar) and pow(), the cost attaches to feeding an inline scalar into an elementwise op, and it does not stack, so a program that already contains pow(x,-0.5) from a norm already has ceiling 15. [source]
- The Orion repo lists 49 commits, 128 stars, 15 forks, 2 branches and no tags, with the last commit 2026-08-23. [source]
- Its README says ANE inference leaves the GPU and CPU free, the ANE is hard-power-gated at idle, and that the "~19 TFLOPS fp16" figure holds because the ANE dequantizes INT8 to fp16 before computing. [source]
- maderix's bridge commit "fix weight dict nil -> @{} (prevents silent compile failure)" conflicts in symptom with Orion's "immediate crash" for the same nil dictionary. [source]
- CoreML-LLM's RMSNorm for ANE is built as `cat([x,-x])` into LayerNorm then a slice, and the resulting graphs reach 99.78% ANE placement through Core ML. [source]
- Forge reports the compiler fuses `x*sigmoid(x)` back into the inaccurate native SiLU, so it uses `0.5*x*(1+tanh(x/2))` in DeltaNet and MLP. [source]
Children
- No children recorded.