MoE active-parameter decode on unified memory

Parent: Mac local LLMs: Speed, bandwidth and prefill · Published reference · snapshot 2026-10-05

↓ Facts as markdownall context files

Measured bandwidth efficiency is much lower for MoE than for dense on the same chip, and falls as chip bandwidth rises. Derived (active bytes are inferred from public active-parameter counts at about 4.25 bits): gpt-oss-20b MXFP4 llama.cpp tg128 reaches roughly 43% of the active-byte ceiling on M...

These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.

Facts

Children

← the whole tree · 3D view· how to read this page