Metal 4 tensor API and M5 Neural Accelerator prefill in llama.cpp

Parent: Mac local LLMs: Speed, bandwidth and prefill · Published reference · snapshot 2026-10-05

↓ Facts as markdownall context files

TensorOps is portable by design: Apple says the same code runs M1 to M5, and on GPUs without Neural Accelerators it falls back to optimised shader implementations. That is why upstream could ship one tensor kernel and then merely disable it by chip name on pre-M5 (it was slower on M2 Ultra), rath...

These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.

Facts

Children

← the whole tree · 3D view· how to read this page