<!-- llms-explorer concept facts · https://llms-explorer.com/tree/llama-cpp-moe-tensor-api-port/ · pack 2026-10-05 · ~580 tokens -->

# llama.cpp MoE tensor-API port

> LM Studio issue 2040 root cause (minos 14.0 binary defaults the runtime shader compile to Metal 3.2) is already held; additionally the failing path logs only "error compiling source" without the NSError, so the real diagnostic never appears in logs.

Parent: [Mac local LLMs: Speed, bandwidth and prefill](https://llms-explorer.com/tree/mac-local-llms-speed-bandwidth-and-prefill/) · 1 facets · 7 facts · page: https://llms-explorer.com/tree/llama-cpp-moe-tensor-api-port/

## Facts

- LM Studio issue 2040 root cause (minos 14.0 binary defaults the runtime shader compile to Metal 3.2) is already held; additionally the failing path logs only "error compiling source" without the NSError, so the real diagnostic never appears in logs. — source: `asserted`
- Whether kernel_mul_mm_id received the #20962 layout is still unconfirmed. — source: `asserted`
- ggml's tensor-API library init failure logs the no-detail variant "error compiling source" rather than the variant carrying the NSError, which hides the "use of undeclared identifier 'mpp'" compile errors seen at Metal language version 3.2. — [source](https://github.com/lmstudio-ai/lmstudio-bug-tracker/issues/2040)
- A reporter's Swift probe showed the embedded matmul2d test kernel compiles at the default and explicit version4_0 language versions but fails at version3_2, and that ggml sets no explicit languageVersion (only setPreprocessorMacros). — [source](https://github.com/lmstudio-ai/lmstudio-bug-tracker/issues/2040)
- BaseRT (arXiv 2607.19438, M5 Pro, 15 model configs) implements MoE expert GEMM and a fused gate/up projection with SiLU on tensor-core accumulators in one dispatch, removing a global memory round trip per expert, and routes only compute-bound phases through the tensor path while decode keeps GEMV kernels. — [source](https://arxiv.org/html/2607.19438v1)
- BaseRT reports decode gains over llama.cpp of 1.02-1.75x overall and 1.53-1.75x on 35B-A3B MoE models, and prefill leads over llama.cpp of 29-120% and over MLX of 15-288% across four MoE models. — [source](https://arxiv.org/html/2607.19438v1)
- BaseRT states M5 keeps advertising MTLGPUFamily.apple9, so the accelerators are reached only through the Metal 4 tensor API, not a capability flag. — [source](https://arxiv.org/html/2607.19438v1)
