Mac local LLMs: KV cache sizing and quantization
Parent: Running LLM models locally on a Mac · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
KV/token = attn layers x 2 x KV heads x head_dim x bytes. Hybrids: count only full-attention layers; sliding layers add a bounded ring, DeltaNet a constant state.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Sizing (count only layers that hold KV)
- KV/token = attn layers x 2 x KV heads x head_dim x bytes. Hybrids: count only full-attention layers; sliding layers add a bounded ring, DeltaNet a constant state. [source]
- Qwen3.6-35B-A3B: 10 attn layers, 2 KV heads, head_dim 256 = 20 KiB/token fp16, ~5 GiB at 262K. [source]
- Qwen3.6-27B: 16 attn layers x 4 KV heads = 64 KiB/token, ~12.5 GiB at 200K. [source]
- Gemma 3/4 in mlx-lm: local layers use RotatingKVCache (bounded); global KVCache grows in 256-token steps, never evicts, so sliding window does not bound memory. [source]
- Gemma-4-26B-A4B global layers disputed (5 vs 10; ~10 vs >20 GiB at 262K): read config.json before sizing. [source]
- Gemma 4 E2B/E4B share KV: count only producer layers (15 of 35, 24 of 42); shared layers get doubled MLP, so weights do not shrink. 26B-A4B reports shared_kv_layers = 0. [source]
- Gemma 4 attention_k_eq_v: global layers reuse K as V (no v_proj); a quantizer expecting v_proj fails. [source]
- 32 GB: Qwen3.6-35B-A3B Q4 plus fp16 KV at 262K is tight; use ~128K or q8_0 KV. 24 GB: Qwen3.6-27B ceiling ~32K fp16, ~64K q8_0. [source]
- KV-head count beats parameter count for long-context sizing. [source]
Runtime knobs and defaults
- llama-server: --cache-prompt on, --cache-reuse 0, -ctxcp 32, -cms 8192, -cram 8192 MiB, --context-shift off, --swa-full false; --cache-type-k/v accepts f32 f16 bf16 q8_0 q4_0 q4_1 iq4_nl q5_0 q5_1. Drop --no-kv-offload on Apple Silicon. [source]
- --ctx-size is split across slots with explicit --parallel N (ctx/N each); set --parallel explicitly. --cache-ram competes with weights on unified memory. [source]
- --checkpoint-every-n-tokens was removed (PR #22929) and now fails to parse; --cache-disk does not exist. [source]
- Ollama: OLLAMA_KV_CACHE_TYPE global, default f16, q8_0 ~1/2 memory, q4_0 ~1/4; needs flash attention (auto since Oct 2025); the server reads it, so exporting in the client terminal does nothing; macOS app needs launchctl setenv; unsupported architectures silently fall back to f16. [source]
- Ollama MLX runner allocates context dynamically: ollama ps CONTEXT is the maximum, SIZE the live use; it keeps an 8 GB snapshot cache, so SIZE climbs 18 to 26 GB; OLLAMA_KEEP_ALIVE=0 stops accumulation. [source]
- mlx_lm.server has no --ctx flag; cap context in the client. MLX_LM_CACHE_LIMIT=0 stops allocator growth. [source]
- mlx-lm generate: --kv-bits, --kv-group-size 64, --quantized-kv-start 5000; server support merged 2026-09-09 (PR #1832). [source]
- Metal flash attention needs K type == V type (#25900), so K=q8_0/V=q4_0 may not run fused. [source]
Failure modes and fixes
- Silent full re-prefill each turn: HTTP 200, only log is "forcing full prompt re-processing due to lack of cache data" at -lv 4. Debug: grep "forcing full prompt|context checkpoint|n_past =", add --log-prompts-dir. [source]
- Hybrid/SWA models cannot truncate; reasoning models drop prior think tokens, breaking prefix reuse; -rea off does not fix it (--reasoning-preserve). [source]
- mlx_lm.server OOM: "kIOGPUCommandBufferCallbackErrorOutOfMemory" (M5 Pro 48 GB, Qwen3.6-27B 4-bit). mx.metal.set_memory_limit(48 GB) before import turns the panic into a Python exception. [source]
- fp16 KV at 63K panicked a 48 GB M5 Pro (30.8 GB peak, AppleARMWatchdogTimer); K8/V3 (28.7 GB) ran. [source]
- mlx-lm quantized KV fails on RotatingKVCache (NotImplementedError); layers lacking a quantized variant are skipped silently (Gemma 4 26B: 25 of 30). Quantized-KV serving is sequential. [source]
- mlx-lm prefill peak with KV4 is +13.2% at step 2048 but -16.8%/-24.0%/-26.2% at 1024/512/256: lower --prefill-step-size. [source]
- mlx-lm LRU prompt cache keeps pre-trim copies (10 GB x 10 = 100 GB). [source]
- M5: TF32 breaks 8 mlx-lm tests; MLX_ENABLE_TF32=0 passes all 31. [source]
- Metal JIT silently falls back to CPU if ggml-metal.metal includes custom headers; fork benchmarks may measure CPU. [source]
KV quantization decisions
- Rule: 8-16 GB, q8_0 K and V, never q4_0 under ~4B (Qwen3 0.6B q4_0 PPL 46.25 vs 13.67). 24-32 GB, q8_0 only if weights+KV pass ~75% of RAM. 48 GB+, keep fp16 unless past ~64K or avoiding the wired-limit panic. [source]
- llama.cpp PR #21038 rotates Q/K/V with a Hadamard transform: existing -ctk q4_0 improves with no new flag; load log shows attn_rot_k; LLAMA_ATTN_ROT_DISABLE=1 turns it off. Qwen3 8B q4_0 PPL 7.6451 to 7.5012 (f16 7.3203). [source]
- Multi-agent batching: skip --kv-bits; use prefix-cache-only quantization (vllm-mlx --kv-cache-bits). [source]
- TurboQuant on Mac: fork TheTom/llama-cpp-turboquant; safe default -ctk q8_0 -ctv turbo4 -fa on; turbo3 never on Qwen2.5 Q4_K_M; head_dim 64 falls back to q8_0. Apple mlx-swift-lm kvScheme turbo8v3 (2.72x vs affine8 1.88x). [source]
- QJL: paper keeps it; practitioners drop it (variance), though mlx_turboquant 3-bit ppl 77.5 without vs 4.82 with. [source]
- Compounding: K3/V3 on 3-bit weights collapsed past ~800 tokens; K8/V3 was clean. [source]
- Mixed precision on Macs is coarse: leave sliding layers and boundary layers high; no per-layer map in llama.cpp/mlx-lm. [source]
Tool calls and agents
Corrections to earlier claims
- "q4_0 costs 0.5-1.1% PPL" is model dependent (+4.4% before rotation). Rotated q4_0 still lost 16 AIME25 points (21.7% vs 37.9%). [source]
- "q4_0 92% slower at 64K" was retracted 2026-04-01. [source]
- "Gemma 3 q8_0 KV drops GPU to 20-30%" was a CUDA bug (#9683). [source]
- "~50-55% speed at 64K" is not general: hybrid MoE lost ~7% by 8K. [source]
Open
- Does MLX runner honor OLLAMA_KV_CACHE_TYPE; are TBQ types in llama.cpp master (PR #21089 closed); any Metal tool-call benchmark; does mlx-lm #1832 fail loudly. [source]
Corrections and disagreements
- CONTRADICTS on-device-local-llm-runtimes.md section 10 (generic formula only): for hybrid models the formula must use only full-attention layers, as in the two derivations above. [source]
- CONTRADICTS on-device-local-llm-runtimes.md 5.1 and anti-pattern 11 on context handling: on MLX-runner tags the VRAM-tier default and `num_ctx` do not pre-allocate or truncate the same way; the log still prints `vram-based default context total_vram="36.0 GiB" default_num_ctx=32768` on an M4 Pro. [source]
- CONTRADICTS on-device-local-llm-runtimes.md section 10 "at 64K depth decode ran at ~50-55% of short-context speed": that figure is not family-independent; the hybrid MoE above lost ~7% by 8K and the dense-attention M3 Ultra case lost 73% by 128K (sibling dossier). [source]
- CONTRADICTS on-device-local-llm-runtimes.md section 10 "GPU gets ~75% of RAM": the same author says llama.cpp on a 64 GB Mac uses about 56 GB by default (article truncated, mechanism unseen). [source]
- CONTRADICTS on-device-local-llm-runtimes.md section 10 ("q4_0 costs ~0.5-1.1% perplexity"): measured q4_0 cost is model dependent, +4.4% on Qwen3 8B before rotation and +2.5% after. [source]
- CONTRADICTS kv-cache-and-long-context-on-mac.md ("Q4_0 KV can be up to 92% slower at 64K"): the 92.5% prompt-throughput collapse and "q4_0 uses more memory than f16" came from a post that its author retracted on 2026-04-01 (silent request failures, RSS instead of GPU memory). [source]
- CONTRADICTS kv-cache-and-long-context-on-mac.md secondary claim "Gemma 3 with even q8_0 KV drops GPU use to 20-30% on Ollama": that was a CUDA kernel-coverage bug closed in October 2025, not an Apple Metal property. [source]
- CONTRADICTS botmonster claim above: llama.cpp metadata for 26B-A4B reports shared_kv_layers = 0, so for that model the 26B-A4B cache arithmetic must use all attention layers, not a reduced producer count (see hybrid-and-sliding-window-attention-kv-cache-rewinding.md). [source]
- CONTRADICTS kv-cache-quantization-tradeoffs-on-apple-gpus.md (rotation cannot apply to Gemma 4 mixed-precision cache): per-cache iSWA rotation exists upstream, so the attn_rot_k = 0 log in #21394 is a point-in-time build observation, not a design limit. [source]
- CONTRADICTS kv-cache-quantization-tradeoffs-on-apple-gpus.md (q4_0 near-lossless after rotation, +2.5% PPL): on a reasoning benchmark rotated Q4_0 still loses 16 points of AIME25 (21.7% vs 37.9% F16); perplexity understates agentic or reasoning damage. [source]
- TriAxialKV (arXiv 2605.17170) is the first source found that reports a tool-call benchmark (BFCL Memory) for a KV-cache quantization variant against a BF16 cache with weights held fixed. CONTRADICTS: kv-cache-quantization-degrading-tool-call-accura.md, which says no controlled measurement exists; the study is CUDA/SGLang, so the Metal gap remains. [source]
Concepts in this cluster
- KV cache and long context on Mac [source]
- KV cache quantization tradeoffs on Apple GPUs [source]
- TurboQuant and rotation-based KV quantization on Metal [source]
- Gemma 4 shared KV cache layers YOCO [source]
- Gemma 4 sliding-window layer exclusion from KV quantization [source]
- KV-cache quantization degrading tool-call accuracy [source]
- Metal flash attention K.type==V.type constraint [source]
- Mixed-precision KV-cache quantization per-layer sensitivity [source]
- Pre-M5 vs M5 Apple GPU decode behavior for compressed KV [source]
- QJL residual correction needed or harmful [source]
- Sparse V attention-gated dequant [source]
- llama.cpp attention rotation Hadamard for quantized KV cache [source]
- Cross-layer attention (CLA) KV sharing [source]
- KV-cache quantization effect on tool-call flips (int8 vs int4) [source]
- Gemma 4 attention_k_eq_v global-layer key-as-value sharing [source]
- KV cache quantization (4-bit and 8-bit) memory pricing and NaN failures in MTPLX [source]
- LCKV and YOCO cross-layer KV sharing variants [source]
- TriAxialKV per-role mixed-precision KV quantization for agentic prefills [source]
- TurboQuant KV cache compression (Hadamard, Lloyd-Max, QJL) [source]
Children
- Metal flash attention K.type==V.type constraint
- Mixed-precision KV-cache quantization per-layer sensitivity
- Pre-M5 vs M5 Apple GPU decode behavior for compressed KV
- QJL residual correction needed or harmful
- Sparse V attention-gated dequant
- TriAxialKV per-role mixed-precision KV quantization for agentic prefills
- TurboQuant and rotation-based KV quantization on Metal
- TurboQuant KV cache compression (Hadamard, Lloyd-Max, QJL)
- Cross-layer attention (CLA) KV sharing
- Gemma 4 attention_k_eq_v global-layer key-as-value sharing
- Gemma 4 shared KV cache layers YOCO
- Gemma 4 sliding-window layer exclusion from KV quantization
- KV cache and long context on Mac
- KV cache quantization (4-bit and 8-bit) memory pricing and NaN failures in MTPLX
- KV-cache quantization degrading tool-call accuracy
- KV-cache quantization effect on tool-call flips (int8 vs int4)
- KV cache quantization tradeoffs on Apple GPUs
- LCKV and YOCO cross-layer KV sharing variants
- llama.cpp attention rotation Hadamard for quantized KV cache