<!-- llms-explorer concept facts · https://llms-explorer.com/tree/hybrid-mamba-2-and-moe-model-support-across-mac/ · pack 2026-10-05 · ~889 tokens -->

# Hybrid Mamba-2 and MoE model support across Mac runtimes

> mlx-lm support for the 30B Nano hybrid was described as still maturing; the 9B and 12B v2 variants were described as fully supported.

Parent: [Mac local LLMs: Runtime selection and frontends](https://llms-explorer.com/tree/mac-local-llms-runtime-selection-and-frontends/) · 1 facets · 14 facts · page: https://llms-explorer.com/tree/hybrid-mamba-2-and-moe-model-support-across-mac/

## Facts

- mlx-lm support for the 30B Nano hybrid was described as still maturing; the 9B and 12B v2 variants were described as fully supported. — source: `asserted`
- Dense 4B and MoE 30B-A3B decode at nearly the same rate on llama.cpp Metal (88 vs 86 tok/s), so MoE active-parameter savings do not appear there. — source: `asserted`
- Same model, two engines: llama.cpp Metal 86.2 tok/s versus MLX 159.7 tok/s (1.85x) on M4 Max for Nemotron-3 Nano 30B-A3B, but the two quants differ (about 6.2 bpw versus 4-bit), so bytes moved per token are not equal. — source: `asserted`
- Whether llama.cpp Metal has a fused SSM (Mamba-2) kernel or runs it as separate ops is not stated in the sources read; the 88 vs 86 result suggests a fixed per-token cost. — source: `asserted`
- No batched-decode measurement of a Mamba-2 hybrid on Mac was found. — source: `asserted`
- NVIDIA Nemotron 3 Super is a 120B hybrid Transformer-Mamba MoE with 12B active parameters and a 1M-token context window. — [source](https://github.com/ggml-org/llama.cpp/discussions/20421)
- llama.cpp discussion 20421 (ggerganov, 2026-03-11) summarizes llama.cpp performance for the Nemotron 3 family. — [source](https://github.com/ggml-org/llama.cpp/discussions/20421)
- On an M4 Max 128 GB (macOS 27.0), Nemotron-3 Nano 30B-A3B decodes at 86.2 +/- 0.5 tok/s on llama.cpp Metal (b8680, unsloth Q4_K_M 24.6 GB) and 159.7 tok/s on mlx-lm 0.31.3 (mlx-community 4-bit, 17.8 GB). — [source](https://github.com/ggml-org/llama.cpp/discussions/20421)
- On the same M4 Max, Nemotron-3 Nano 4B decodes at 88.4 tok/s on llama.cpp Metal (Q4_K_M 2.84 GB), 176.8 tok/s on MLX (4-bit 2.24 GB) and 85.2 tok/s on Apple Core AI (int8-head export, 4.6 GB). — [source](https://github.com/ggml-org/llama.cpp/discussions/20421)
- Apple Core AI ran the Nano 4B at 16.0 tok/s on an iPhone 17 Pro (AOT, cooled). — [source](https://github.com/ggml-org/llama.cpp/discussions/20421)
- llama.cpp Metal prompt processing on the M4 Max: pp512 1212.8 tok/s for Nano 30B-A3B and 1308.8 tok/s for Nano 4B, with max RSS 24.4 GB and 3.0 GB; llama-bench used `-p 512 -n 256 -ngl 99 -fa 1 -r 3` and `-fa 0` was within noise. — [source](https://github.com/ggml-org/llama.cpp/discussions/20421)
- The contributor flagged that on llama.cpp the dense 4B decodes almost as slowly as the 30B-A3B despite a 9x smaller file, while MLX shows 177 versus 160 tok/s; he did not investigate the cause. — [source](https://github.com/ggml-org/llama.cpp/discussions/20421)
- Column caveats stated by the contributor: the unsloth Q4_K_M is about 6.2 bits per weight, the Core AI artifact is not a 4-bit run, and peak memory is max RSS for llama.cpp but Metal peak allocation for MLX. — [source](https://github.com/ggml-org/llama.cpp/discussions/20421)
- A guide to running Nemotron on a Mac with MLX says the 30B Nano hybrid Mamba2-Transformer is still maturing in mlx-lm and recommends the 9B or 12B v2 variants if problems appear, and quotes about 80-100 tok/s for the 30B in 4-bit on a 32-64 GB MacBook Pro. — [source](https://aplicar.ai/running-nvidias-nemotron-open-models-on-your-mac-with-mlx/)
