<!-- llms-explorer concept facts · https://llms-explorer.com/tree/q4-k-m-vs-q4-k-s-decode-on-moe-q6-k-ffn-down-exp/ · pack 2026-10-05 · ~1021 tokens -->

# Q4_K_M vs Q4_K_S decode on MoE (Q6_K ffn_down_exps share)

> For a non-Falcon model, Q4_K_M's ffn_down is Q6_K wherever use_more_bits(i_layer, n_layer) is true, else Q4_K; Q4_K_S's ffn_down is Q5_K only for i_layer < n_layer/8, else Q4_K.

Parent: [Mac local LLMs: Speed, bandwidth and prefill](https://llms-explorer.com/tree/mac-local-llms-speed-bandwidth-and-prefill/) · 1 facets · 16 facts · page: https://llms-explorer.com/tree/q4-k-m-vs-q4-k-s-decode-on-moe-q6-k-ffn-down-exp/

## Facts

- For a non-Falcon model, Q4_K_M's ffn_down is Q6_K wherever use_more_bits(i_layer, n_layer) is true, else Q4_K; Q4_K_S's ffn_down is Q5_K only for i_layer < n_layer/8, else Q4_K. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/src/llama-quant.cpp)
- Q4_K_M also sets the fused ATTENTION_QKV category to Q5_K and attn_v to Q6_K by use_more_bits; Q4_K_S gives Q5_K to attn_v only while i_attention_wv < 4. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/src/llama-quant.cpp)
- For n_expert > 1 the layer index comes from the tensor name, so merged ffn_down_exps tensors follow the same per-layer rule. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/src/llama-quant.cpp)
- Computed from the use_more_bits rule (first eighth, last eighth and every third layer between): the share of layers selected is 50% at 40, 48 and 60 layers and 48.9% at 94 layers. — source: `asserted`
- On a dense 8B model the llama.cpp quantize table (hardware not stated) lists Q4_K_S at 4.6672 bpw and 76.71 tg128 and Q4_K_M at 4.8944 bpw and 71.93: Q4_K_M is 4.9% larger and 6.2% slower, so decode lost slightly more than the byte increase. — [source](https://github.com/ggml-org/llama.cpp/blob/master/tools/quantize/README.md)
- On a MoE the active-expert bytes per token are a small fraction of the file, so the Q6_K share of per-token bytes is (6.5625 - 4.5) over 4.5 on the selected down-projection rows in about half the layers, an extra 2% to 8% of decoded expert bytes depending on the down share, not the file-size gap. — source: `asserted`
- Quality side: the same recipe difference is what moves KLD; the llama.cpp 8B table has q4_K_S 0.043136 and q4_K_M 0.031273 (no imatrix) at 4.37 and 4.58 GiB. — [source](https://github.com/ggml-org/llama.cpp/blob/master/tools/perplexity/README.md)
- 2023-06: both mixes introduced in PR 1684 (covering files). Rules read here are from current master. — source: `asserted`
- For Mixtral-style n_expert == 8, attn_v and attn_k special rules apply and attn_output goes to Q5_K for both types (existing dossier), narrowing the two recipes' attention difference. — source: `asserted`
- Unsloth and community "Q4_K_M" MoE files often are not stock llama-quantize output (UD recipes), so a published Q4_K_M versus Q4_K_S pair can differ in more than the two rules above. — source: `asserted`
- "Q4_K_S is the better local default for speed" (Artefact2-style advice, covering dossier) versus "Q4_K_M's extra Q6_K is cheap" (K-quant design intent): the dense 8B ratio (6.2% slower for 4.9% more bytes) supports the first mildly; no MoE number exists. — source: `asserted`
- llama-bench tg128 for stock Q4_K_S and Q4_K_M of one Qwen3-30B-A3B-class MoE on one Apple chip, same imatrix. No source found. — source: `asserted`
- In llama-quant.cpp Q4_K_M sets ffn_down to Q6_K by use_more_bits and Q4_K_S sets it to Q5_K only for the first eighth of layers. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/src/llama-quant.cpp)
- In llama-quant.cpp Q4_K_S sets attn_v to Q5_K only for the first four attention-v tensors, while Q4_K_M uses Q6_K by use_more_bits and Q5_K for fused QKV. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/src/llama-quant.cpp)
- llama.cpp's Llama-3.1-8B table gives Q4_K_S 76.71 and Q4_K_M 71.93 tokens per second at 128 tokens, with 4.6672 and 4.8944 bits per weight. — [source](https://github.com/ggml-org/llama.cpp/blob/master/tools/quantize/README.md)
- No MoE decode comparison of Q4_K_S and Q4_K_M was found. — source: `asserted`
