UltraFusion die-to-die latency and irregular MoE access
Parent: Mac local LLMs: Speed, bandwidth and prefill · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Older commentary blamed M3 Ultra GPU compute for slow dense inference and favoured MoE for it; none attributes MoE slowness to UltraFusion.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Older commentary blamed M3 Ultra GPU compute for slow dense inference and favoured MoE for it; none attributes MoE slowness to UltraFusion. [source]
- No microbenchmark of cross-die access latency or expert-locality placement was found. [source]
- This concept is already covered by moe-active-parameter-decode-on-unified-memory.md; no source read in this batch measures UltraFusion latency directly. [source]
- An early M3 Ultra review concluded the GPU, not RAM, limits dense LLM inference and that a 128 GB model suits dense 32B-and-smaller or large MoE models; its author later revised the M3 Ultra estimate upward to about 4x his M2 Max. [source]
Children
- No children recorded.