<!-- llms-explorer concept facts · https://llms-explorer.com/tree/mac-local-llms-mlx-kernels-numerics-and-internals/ · pack 2026-10-05 · ~3070 tokens -->

# Mac local LLMs: MLX kernels, numerics and internals

> `MLX_ENABLE_TF32` defaults to 1, is read once. It only lets NAX run fp32 GEMM at TF32; set `=0` in the launch env for fp32 reproducibility.

Parent: [Running LLM models locally on a Mac](https://llms-explorer.com/tree/running-llm-models-locally-on-mac/) · 11 facets · 61 facts · page: https://llms-explorer.com/tree/mac-local-llms-mlx-kernels-numerics-and-internals/

## Switches and checks (set before the first MLX op; they latch)

- `MLX_ENABLE_TF32` defaults to 1, is read once. It only lets NAX run fp32 GEMM at TF32; set `=0` in the launch env for fp32 reproducibility. — [source](https://github.com/ml-explore/mlx/issues/3860)
- M5 fp32 256x256 matmul error vs torch CPU 5.2e-2 (M4 3.8e-5) since mlx 0.30.0; TF32=0 fixes it. — [source](https://github.com/ml-explore/mlx/issues/3534)
- `MLX_METAL_GPU_ARCH=applegpu_g16s` on an M5 turns off every NAX gate but also drops qmv limits 33/25 to 13/15, disables the nvfp4 narrow qmv route and changes SDPA 2-pass routing. A malformed value such as `applegpu_g16` parses as gen 1 without error. — source: `asserted`
- Chips: M5 base `applegpu_g17g`; M5 Pro/Max `g17s`; M4 Pro `g16s`; M3 `g15s`; Check: `mx.device_info(mx.gpu)`. — [source](https://github.com/ml-explore/mlx/pull/4596)
- NAX needs macOS >= 26.2, gen >= 17 and a build with `MACOSX_DEPLOYMENT_TARGET=26.2` ; `MLX_METAL_NO_NAX` removes it. Otherwise it runs non-NAX silently. bf16 4096^3 GEMM, M5 Max: 2.4-2.6 ms vs ~9 ms without. — [source](https://github.com/ml-explore/mlx/pull/3838)

## M5 numerics

- Two causes: fp16/bf16 masked attention takes NAX attention (TF32 flag does nothing; only the arch override moves it); fp32 GEMM takes NAX at TF32. Padded batch-vs-single gap at D=64: M5 fp16 2^-11, bf16 2^-8, fp32 2^-11.7; M3 Max 2^-14.4, 0, 2^-22.1. Unpadded is exactly 0. — [source](https://github.com/ml-explore/mlx/issues/3897)
- mlx-lm test_generate on M5: 8 of 28 fail natively, 28 pass with TF32=0. mlx-lm PR 1595 pins TF32=0 in tests. — [source](https://github.com/ml-explore/mlx-lm/pull/1595)
- Matvec shapes (M=1 or N=1) stay exact fp32. Issue 3897 closed won't fix; docs now document the default. — [source](https://github.com/ml-explore/mlx/issues/3860)

## Kernel routing (what runs where)

- NAX serves prefill, large batch, qmm at or above the qmv limit. Decode (M=1) and verify windows use gemv/qmv/`qmv_wide`/`gemv_wide`, so NAX changes TTFT and batch throughput, not single-stream tok/s. — source: `asserted`
- `dispatch_qmv`: K 64/128 goes to `qmv_quad`; else M>=2 and (`mode != affine` or gen>=15) and no global scale goes to `qmv_wide`; else `qmv`. M1/M2 affine int4/int8 never get qmv_wide; fp modes do. `get_qmv_batch_limit` (K,N <=2048/<=4096/larger, transposed only): gen>=17 non-Ultra 33/25/13; gen 15-16 13/15/13; Ultra 32/18/12; gen 13-14 14/10/6. — [source](https://raw.githubusercontent.com/ml-explore/mlx/main/mlx/backend/metal/quantized.cpp)
- `qmv_wide` reads weights ceil(M/5) times: M=2-5 one stream, 6-10 two, 11-13 three. PR 3764 (v0.32.0) gains: int4 1.1-1.3x at M=2; nvfp4/mxfp8 up to 2.2x; — [source](https://github.com/ml-explore/mlx/pull/3764)
- `gemv_wide` (PR 3888, v0.32.1): bf16/fp16 `x @ W.T`, M=2-15 (passes=(M+4)/5, max 3), gen>=15, K%4==0, vec4-aligned; misaligned or fp32 silently falls back. Tiny in_proj M=2: 2.3-6.0x (M5 Max 6.0x). — [source](https://github.com/ml-explore/mlx/pull/3888)
- NAX qmm needs K%64==0 (group size irrelevant; K=2400 drops to classic qmm), plus TF32 clause for fp32 x. Non-transposed needs affine and N%64. No NAX for qmv, qmv_wide, qvm or `qmm_splitk`. — [source](https://raw.githubusercontent.com/ml-explore/mlx/main/mlx/backend/metal/quantized.cpp)
- `qmm_splitk` (PR 3120, 32x32 tiles, split_k=max(1,512/tiles)) pre-empts NAX while tiles <=256: M<=32 with N<=8192, M 33-64 with N<=4096. M5 cliff: N=6656 costs 252-258 us at M=24-32 vs 194-205 us at M=33-40. — [source](https://github.com/ml-explore/mlx/issues/4198)
- PR 4171 (merged 2026-08-21): 32-row `qmm_t_nax` for M<=32, 1.21-1.22x on base M5; forced BM=32 at M=128 is 13% slower. — [source](https://github.com/ml-explore/mlx/pull/4171)
- SDPA: vector kernel for qL<=8, full (NAX) kernel above. Main NAX head dims 64/96/128/256/512; 72/80 pad to 96 for half, qL,kL>=512, no mask. — [source](https://github.com/ml-explore/mlx/issues/3897)

## Allocator, memory limits, crashes

- Defaults: block_limit=min(1.5x recommended working set, 0.95x RAM); gc_limit=min(0.95x recommended, block_limit); cache limit=block_limit. `set_memory_limit` moves block and gc limits but not the cache limit. Error: `[metal::malloc] Resource limit (N) exceeded.` — [source](https://raw.githubusercontent.com/ml-explore/mlx/main/mlx/backend/metal/allocator.cpp)
- `[METAL] Command buffer execution failed: Insufficient Memory (kIOGPUCommandBufferCallbackErrorOutOfMemory)` kills the process. — [source](https://github.com/ml-explore/mlx-lm/issues/1015)
- Panic `completeMemory() prepare count underflow` at ~58k tokens, 80 GB wired on a 96 GB M3 Ultra; mlx_lm.server wires ~75% of RAM. Fix: `mx.metal.set_memory_limit(48 * 1024**3)` before `mlx_lm.server.main()` for a Python exception instead of a panic. — [source](https://github.com/ml-explore/mlx-lm/issues/883)

## Threads and streams (0.31.2+)

- Each thread gets its own default stream and encoder. Using a stream from another thread throws `There is no Stream(gpu, N) in current thread`; do not create streams at import. A lazy array built on thread A cannot be evaluated on B: `mx.eval` it on the owner first (oMLX 1558 crash). — [source](https://github.com/ml-explore/mlx/pull/3348)
- `new_thread_unsafe_stream` (PR 3578, merged 2026-06-17, 693ed59): a stream usable from any thread with no MLX locking; you add the locks. Built for Swift-style async eval (mlx-swift issue 407: completion-handler SIGABRT on iPhone 15 Pro); author calls it a stopgap until thread-safe streams. No mlx-swift adoption found. — [source](https://github.com/ml-explore/mlx/pull/3578)

## Dtypes and half precision

- Python scalars are weak: `x_fp16 * 0.5` stays fp16, `x_fp16 * mx.array(0.5)` or `mx.float32(...)` widens to fp32. A float above 65,504 becomes inf in fp16 without a warning. — [source](https://raw.githubusercontent.com/ml-explore/mlx/main/python/src/utils.cpp)

## Ternary Bonsai 2 27B (Hadamard-rotated weights)

- Weights are Hadamard-rotated (block 1024, fixed +-1 signs) before ternary assignment; the runtime must apply the matching activation transform. Ordinary MLX loaders skip it and return wrong output with no error; the pack declares the rotation in metadata. Stock llama.cpp cannot run the GGUF packs; use the PrismML-Eng/llama.cpp fork. — [source](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-mlx-2bit)
- Sizes: MLX 2-bit codes (scale s, bias -s) 2.25 bits/weight, 8.60 GB (7.67 GB LM + 0.92 GB FP16 vision); GGUF PTQ1_0 5.95 GB, PQ2_0 7.21 GB, neither uniformly faster. M5 Pro Metal PQ2_0: 28.1 tok/s TG128, 387 PP512. Vendor claim: 46.8 tok/s on M5 Max, 98.2% retention. — [source](https://prismml.com/news/bonsai-2-27b)

## JACCL (Thunderbolt RDMA)

- `RTR failed with errno 22` and the `[jaccl] No IPv4-mapped GID` error: the port has no IPv4 address. Fix: `networksetup -setmanual "<service>" 169.254.250.N 255.255.255.0`. errno 60/96: link-local or stale GID. — [source](https://github.com/ml-explore/mlx/issues/3467)
- Versions: 0.32.0 has the all_reduce race fix (3451) and ring recv deadlock fix (3654); ring `all_gather` is silently wrong in 0.32.1-0.32.2 (PR 4443, 0.32.3). — [source](https://github.com/ml-explore/mlx/releases/tag/v0.32.0)

## Corrections to earlier claims

- Earlier: TF32 explains all M5 divergence, "31" tests. Now 8 of 28; NAX attention is a separate cause. — source: `asserted`
- Earlier: cache limit defaults to memory limit (docs); only until `set_memory_limit`. — [source](https://ml-explore.github.io/mlx/build/html/python/_autosummary/mlx.core.set_cache_limit.html)
- Earlier: errno 22 fixed in code; #3451 ended reboot errors. Now: PR 4191 only adds a clearer error; 3451 is a wrong-sum race. — [source](https://github.com/ml-explore/mlx/pull/4191)

## Corrections and disagreements

- CONTRADICTS kv-cache-quantization-tradeoffs-on-apple-gpus.md: it attributes the eight failures to TF32 in fp32 GEMM and lists "31" tests; the issue author's counts are 8 of 28 (28 pass with TF32 off). It also folds the batched-attention divergence into one cause; the thread separates a NAX-attention cause (fp16/bf16, not TF32-controlled) from the TF32 GEMM cause (fp32 and apparently the quantized-matmul path). The count may differ because the suite changed. — source: `asserted`
- CONTRADICTS (docs vs source, no existing file): the set_cache_limit doc says the cache limit defaults to the memory limit; in source this holds only until set_memory_limit is called. — [source](https://ml-explore.github.io/mlx/build/html/python/_autosummary/mlx.core.set_cache_limit.html)
- CONTRADICTS: thunderbolt-bridge-bridge0-and-stp-fixes-for-rdm.md ("errno 22 ... a GID-selection regression in JACCL (#3467 ..., fix is code)"): the merged change does not alter which GID is chosen; it only throws a clearer error, and the working fix is giving the interface an IPv4 address. — [source](https://github.com/ml-explore/mlx/pull/4191)
- CONTRADICTS: thunderbolt-bridge-bridge0-and-stp-fixes-for-rdm.md "fix is code": PR 4191, merged 2026-08-12, only reports a missing IPv4-mapped GID with a remedy. — [source](https://github.com/ml-explore/mlx/pull/4191)
- CONTRADICTS (causal reading): distributed-inference-across-macs.md line "#3451 ... removed most reboot-needing RTR errno 22/60 failures": that is one operator's attribution in a five-patch stack; the PR describes only a wrong-sum race in `MeshImpl::all_reduce`. — [source](https://github.com/ml-explore/mlx/pull/3451)
- CONTRADICTS jaccl-gid-selection-regression-and-rtr-errno-22.md in framing only: errno 22 with no IPv4 address predates the April refactor (issue 1390, Feb 2026), so the refactor changed the symptom path (uninitialised GID instead of an absent index 1) and did not create the underlying missing-address condition. — [source](https://github.com/exo-explore/exo/issues/1390)

## Concepts in this cluster

- M5 GPU TF32 numerics and batched attention divergence — source: `asserted`
- M5 NAX tensor-unit kernels in MLX servers — source: `asserted`
- MLX NAX kernel gating and quantized matmul — source: `asserted`
- MLX allocator memory limit GC limit and buffer cache — source: `asserted`
- MLX batched decode scheduler crashes — source: `asserted`
- MLX qmv_wide kernel PR 3764 and affine gen-15+ gating — source: `asserted`
- MLX gemv_wide bf16 fp16 few-row matmul kernel PR 3888 — source: `asserted`
- MLX get_qmv_batch_limit per-generation tuning table — source: `asserted`
- MLX qmm_splitk small-M split-K quantized matmul — source: `asserted`
- Gemma 4 vision tower standardize overflow and mmproj dtype on Metal — source: `asserted`
- JACCL GID selection regression and RTR errno 22 (mlx 3467) — source: `asserted`
- JACCL MeshImpl all_reduce race fix mlx 3451 — source: `asserted`
- MLX_METAL_GPU_ARCH architecture override as NAX kill switch — source: `asserted`
- NAX-aware qmm_splitk for small-M quantized matmul on M5 — source: `asserted`
- Gemma4ClippableLinear input/output clamp scalars in the mmproj — source: `asserted`
- JACCL RTR sgid_index hard-coded to 1 — source: `asserted`
- MLX mixed fp16 and bf16 type promotion in arithmetic — source: `asserted`
- MLX thread-local streams and cross-thread lazy array evaluation crashes — source: `asserted`
- mlx 0.32.0 JACCL patch inventory — source: `asserted`
- FP16 attention partial-sum overflow in split attention kernels on M1 and M2 — source: `asserted`
- JACCL ring all_gather direction-1 slice bug (mlx PR 4443) — source: `asserted`
- MLX 0.31.2 thread-safety PRs and delivered levels — source: `asserted`
- MLX Python scalar weak typing with half-precision arrays — source: `asserted`
- Prism ML ternary Hadamard-rotated weights kernel on Metal — source: `asserted`
- MLX new_thread_unsafe_stream API and Swift concurrency use — source: `asserted`
- MLX ThreadLocalStream per-thread resolution and cross-thread synchronize no-op — source: `asserted`
