Hybrid Mamba-2 and MoE model support across Mac runtimes
Parent: Mac local LLMs: Runtime selection and frontends · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
mlx-lm support for the 30B Nano hybrid was described as still maturing; the 9B and 12B v2 variants were described as fully supported.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- mlx-lm support for the 30B Nano hybrid was described as still maturing; the 9B and 12B v2 variants were described as fully supported. [source]
- Dense 4B and MoE 30B-A3B decode at nearly the same rate on llama.cpp Metal (88 vs 86 tok/s), so MoE active-parameter savings do not appear there. [source]
- Same model, two engines: llama.cpp Metal 86.2 tok/s versus MLX 159.7 tok/s (1.85x) on M4 Max for Nemotron-3 Nano 30B-A3B, but the two quants differ (about 6.2 bpw versus 4-bit), so bytes moved per token are not equal. [source]
- Whether llama.cpp Metal has a fused SSM (Mamba-2) kernel or runs it as separate ops is not stated in the sources read; the 88 vs 86 result suggests a fixed per-token cost. [source]
- No batched-decode measurement of a Mamba-2 hybrid on Mac was found. [source]
- NVIDIA Nemotron 3 Super is a 120B hybrid Transformer-Mamba MoE with 12B active parameters and a 1M-token context window. [source]
- llama.cpp discussion 20421 (ggerganov, 2026-03-11) summarizes llama.cpp performance for the Nemotron 3 family. [source]
- On an M4 Max 128 GB (macOS 27.0), Nemotron-3 Nano 30B-A3B decodes at 86.2 +/- 0.5 tok/s on llama.cpp Metal (b8680, unsloth Q4_K_M 24.6 GB) and 159.7 tok/s on mlx-lm 0.31.3 (mlx-community 4-bit, 17.8 GB). [source]
- On the same M4 Max, Nemotron-3 Nano 4B decodes at 88.4 tok/s on llama.cpp Metal (Q4_K_M 2.84 GB), 176.8 tok/s on MLX (4-bit 2.24 GB) and 85.2 tok/s on Apple Core AI (int8-head export, 4.6 GB). [source]
- Apple Core AI ran the Nano 4B at 16.0 tok/s on an iPhone 17 Pro (AOT, cooled). [source]
- llama.cpp Metal prompt processing on the M4 Max: pp512 1212.8 tok/s for Nano 30B-A3B and 1308.8 tok/s for Nano 4B, with max RSS 24.4 GB and 3.0 GB; llama-bench used `-p 512 -n 256 -ngl 99 -fa 1 -r 3` and `-fa 0` was within noise. [source]
- The contributor flagged that on llama.cpp the dense 4B decodes almost as slowly as the 30B-A3B despite a 9x smaller file, while MLX shows 177 versus 160 tok/s; he did not investigate the cause. [source]
- Column caveats stated by the contributor: the unsloth Q4_K_M is about 6.2 bits per weight, the Core AI artifact is not a 4-bit run, and peak memory is max RSS for llama.cpp but Metal peak allocation for MLX. [source]
- A guide to running Nemotron on a Mac with MLX says the 30B Nano hybrid Mamba2-Transformer is still maturing in mlx-lm and recommends the 9B or 12B v2 variants if problems appear, and quotes about 80-100 tok/s for the 30B in 4-bit on a 32-64 GB MacBook Pro. [source]
Children
- No children recorded.