llama.cpp MoE tensor-API port
Parent: Mac local LLMs: Speed, bandwidth and prefill · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
LM Studio issue 2040 root cause (minos 14.0 binary defaults the runtime shader compile to Metal 3.2) is already held; additionally the failing path logs only "error compiling source" without the NSError, so the real diagnostic never appears in logs.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- LM Studio issue 2040 root cause (minos 14.0 binary defaults the runtime shader compile to Metal 3.2) is already held; additionally the failing path logs only "error compiling source" without the NSError, so the real diagnostic never appears in logs. [source]
- Whether kernel_mul_mm_id received the #20962 layout is still unconfirmed. [source]
- ggml's tensor-API library init failure logs the no-detail variant "error compiling source" rather than the variant carrying the NSError, which hides the "use of undeclared identifier 'mpp'" compile errors seen at Metal language version 3.2. [source]
- A reporter's Swift probe showed the embedded matmul2d test kernel compiles at the default and explicit version4_0 language versions but fails at version3_2, and that ggml sets no explicit languageVersion (only setPreprocessorMacros). [source]
- BaseRT (arXiv 2607.19438, M5 Pro, 15 model configs) implements MoE expert GEMM and a fused gate/up projection with SiLU on tensor-core accumulators in one dispatch, removing a global memory round trip per expert, and routes only compute-bound phases through the tensor path while decode keeps GEMV kernels. [source]
- BaseRT reports decode gains over llama.cpp of 1.02-1.75x overall and 1.53-1.75x on 35B-A3B MoE models, and prefill leads over llama.cpp of 29-120% and over MLX of 15-288% across four MoE models. [source]
- BaseRT states M5 keeps advertising MTLGPUFamily.apple9, so the accelerators are reached only through the Metal 4 tensor API, not a capability flag. [source]
Children
- No children recorded.