<!-- llms-explorer concept facts · https://llms-explorer.com/tree/ultrafusion-die-to-die-latency-and-irregular-moe/ · pack 2026-10-05 · ~313 tokens -->

# UltraFusion die-to-die latency and irregular MoE access

> Older commentary blamed M3 Ultra GPU compute for slow dense inference and favoured MoE for it; none attributes MoE slowness to UltraFusion.

Parent: [Mac local LLMs: Speed, bandwidth and prefill](https://llms-explorer.com/tree/mac-local-llms-speed-bandwidth-and-prefill/) · 1 facets · 4 facts · page: https://llms-explorer.com/tree/ultrafusion-die-to-die-latency-and-irregular-moe/

## Facts

- Older commentary blamed M3 Ultra GPU compute for slow dense inference and favoured MoE for it; none attributes MoE slowness to UltraFusion. — source: `asserted`
- No microbenchmark of cross-die access latency or expert-locality placement was found. — source: `asserted`
- This concept is already covered by moe-active-parameter-decode-on-unified-memory.md; no source read in this batch measures UltraFusion latency directly. — source: `asserted`
- An early M3 Ultra review concluded the GPU, not RAM, limits dense LLM inference and that a 128 GB model suits dense 32B-and-smaller or large MoE models; its author later revised the M3 Ultra estimate upward to about 4x his M2 Max. — [source](https://medium.com/@billynewport/apples-m3-ultra-mac-studio-misses-the-mark-for-llm-inference-f57f1f10a56f)
