Q4_K_M vs Q4_K_S decode on MoE (Q6_K ffn_down_exps share)
Parent: Mac local LLMs: Speed, bandwidth and prefill · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
For a non-Falcon model, Q4_K_M's ffn_down is Q6_K wherever use_more_bits(i_layer, n_layer) is true, else Q4_K; Q4_K_S's ffn_down is Q5_K only for i_layer < n_layer/8, else Q4_K.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- For a non-Falcon model, Q4_K_M's ffn_down is Q6_K wherever use_more_bits(i_layer, n_layer) is true, else Q4_K; Q4_K_S's ffn_down is Q5_K only for i_layer < n_layer/8, else Q4_K. [source]
- Q4_K_M also sets the fused ATTENTION_QKV category to Q5_K and attn_v to Q6_K by use_more_bits; Q4_K_S gives Q5_K to attn_v only while i_attention_wv < 4. [source]
- For n_expert > 1 the layer index comes from the tensor name, so merged ffn_down_exps tensors follow the same per-layer rule. [source]
- Computed from the use_more_bits rule (first eighth, last eighth and every third layer between): the share of layers selected is 50% at 40, 48 and 60 layers and 48.9% at 94 layers. [source]
- On a dense 8B model the llama.cpp quantize table (hardware not stated) lists Q4_K_S at 4.6672 bpw and 76.71 tg128 and Q4_K_M at 4.8944 bpw and 71.93: Q4_K_M is 4.9% larger and 6.2% slower, so decode lost slightly more than the byte increase. [source]
- On a MoE the active-expert bytes per token are a small fraction of the file, so the Q6_K share of per-token bytes is (6.5625 - 4.5) over 4.5 on the selected down-projection rows in about half the layers, an extra 2% to 8% of decoded expert bytes depending on the down share, not the file-size gap. [source]
- Quality side: the same recipe difference is what moves KLD; the llama.cpp 8B table has q4_K_S 0.043136 and q4_K_M 0.031273 (no imatrix) at 4.37 and 4.58 GiB. [source]
- 2023-06: both mixes introduced in PR 1684 (covering files). Rules read here are from current master. [source]
- For Mixtral-style n_expert == 8, attn_v and attn_k special rules apply and attn_output goes to Q5_K for both types (existing dossier), narrowing the two recipes' attention difference. [source]
- Unsloth and community "Q4_K_M" MoE files often are not stock llama-quantize output (UD recipes), so a published Q4_K_M versus Q4_K_S pair can differ in more than the two rules above. [source]
- "Q4_K_S is the better local default for speed" (Artefact2-style advice, covering dossier) versus "Q4_K_M's extra Q6_K is cheap" (K-quant design intent): the dense 8B ratio (6.2% slower for 4.9% more bytes) supports the first mildly; no MoE number exists. [source]
- llama-bench tg128 for stock Q4_K_S and Q4_K_M of one Qwen3-30B-A3B-class MoE on one Apple chip, same imatrix. No source found. [source]
- In llama-quant.cpp Q4_K_M sets ffn_down to Q6_K by use_more_bits and Q4_K_S sets it to Q5_K only for the first eighth of layers. [source]
- In llama-quant.cpp Q4_K_S sets attn_v to Q5_K only for the first four attention-v tensors, while Q4_K_M uses Q6_K by use_more_bits and Q5_K for fused QKV. [source]
- llama.cpp's Llama-3.1-8B table gives Q4_K_S 76.71 and Q4_K_M 71.93 tokens per second at 128 tokens, with 4.6672 and 4.8944 bits per weight. [source]
- No MoE decode comparison of Q4_K_S and Q4_K_M was found. [source]
Children
- No children recorded.