Q6_K Metal mat-vec cost inside Q4_K_M files
Parent: Mac local LLMs: Speed, bandwidth and prefill · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Bytes: a Q6_K block is 128 B of low nibbles (ql), 64 B of high 2-bit pairs (qh), 16 B of scales and 2 B of d, 210 B per 256 weights or 6.5625 bpw, against 144 B and 4.5 bpw for Q4_K, so a Q6_K tensor moves 1.458 times the bytes of the same tensor in Q4_K.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Bytes: a Q6_K block is 128 B of low nibbles (ql), 64 B of high 2-bit pairs (qh), 16 B of scales and 2 B of d, 210 B per 256 weights or 6.5625 bpw, against 144 B and 4.5 bpw for Q4_K, so a Q6_K tensor moves 1.458 times the bytes of the same tensor in Q4_K. [source]
- Reconstruction work: the Q6_K kernel builds each weight from a low nibble and two high bits (mask-shift-or per weight), subtracts 32 as an integer, multiplies by the float activation, and applies one int8 sub-block scale per 16 weights and the fp16 d once per super-block. It has no per-row minimum term and does not pre-sum activations, unlike Q4_K. [source]
- Memory access style: the Q6_K kernel reads ql, qh and scales through single-byte pointers (uint8_t, int8_t) while the Q4_K kernel reads packed data and scales through uint16_t pointers. [source]
- Launch shape: both kernels use two output rows per simdgroup and two simdgroups per threadgroup (N_R0 2, N_SG 2), so a Q6_K threadgroup produces the same 4 rows as a Q4_K one but over 1.46 times the bytes; Q5_K is the odd one at N_R0 1. [source]
- Lane mapping: two lanes (tiisg % 2) split a super-block between them and stride by 2 over blocks; each lane keeps 16 activations in registers and, for each of the nr0 rows, rebuilds 16 weights per block. [source]
- MoE path: expert matmuls use kernel_mul_mv_id_q6_K_f32, a thin wrapper that calls the same kernel_mul_mv_q6_K_f32_impl with N_R0_Q6_K, so a Q6_K routed expert costs per selected expert what a dense Q6_K matrix of that shape costs. [source]
- Share of a MoE Q4_K_M file that is Q6_K: for MoE models (n_expert > 1) the "more bits" test is by layer, and it is true for the first eighth, the last eighth and every third layer between, about half the layers. It applies to every ffn_down-category tensor, so in those layers the merged `ffn_down_exps` and the shared `ffn_down_shexp` are Q6_K, together with attn_v-like tensors and output.weight. [source]
- Order of magnitude for that share: if the down projection holds about one third of expert parameters, about half the layers at Q6_K add roughly (6.5625 - 4.5) x 1/3 x 1/2, about 0.34 bits per expert weight, or about 7% more expert bytes than an all-Q4_K file before dilution by non-expert tensors. [source]
- Per-token cost view at batch 1: decode reads only the routed experts' rows, so extra Q6_K cost scales with the active down-projection bytes (a fraction of active parameters), not with total file size; prefill uses matrix-matrix kernels and is dominated by compute, where Q6_K and Q4_K both pay a dequantize into threadgroup memory. [source]
- 2023-06: K-quants (PR 1684) introduce Q4_K_M as Q6_K for half of attention.wv and feed_forward.w2 and publish ms per token by type on M2 Max, RTX 4080 and Ryzen 7950X. [source]
- Later Metal refactor: per-type N_R0 and N_SG constants move to ggml-metal-impl.h and the host picks nsg and nr0 per type in ggml-metal-device.cpp; MoE expert kernels reuse the dense implementations through kernel_mul_mv_id. [source]
- The only per-type timing table in the sources (PR 1684) reports ms per token at 4 and 8 threads on an M2 Max and does not name Metal; K-quants predate Metal K-quant kernels, so it is a CPU-thread measurement and says nothing direct about the Metal Q6_K kernel. [source]
- On that table, time tracks bytes closely on the Apple CPU (7B at 4 threads: Q4_K_S 50 ms, Q4_K_M 55 ms, Q6_K 75 ms against 3.56, 3.80 and 5.15 GB) but not on the RTX 4080 (15.5, 16.0, 18.3 ms) where Q6_K costs 18% more for 45% more bytes. A bandwidth-bound Apple GPU should sit nearer the first pattern. [source]
- A Q6_K tensor in a layer costs more than its byte ratio only if reconstruction ALU or single-byte loads become the limit; on Ultra and Max chips with high bandwidth that is plausible, on base M-series chips it is bandwidth-first. No measurement was found either way. [source]
- Mixed files mix two pipelines in one token: Q4_K and Q6_K tensors differ in nsg alignment and register use, and each distinct type is a separate compiled pipeline, so a Q4_K_M model keeps more kernels hot than a Q4_K_S one. [source]
- "Q6_K is nearly free because it covers few tensors" (K-quant design intent) versus "half of ffn_down plus all of output.weight is not few on a MoE" (by-layer rule applied to merged expert tensors): on a dense 7B the Q6_K share is about 1/6 of weights, on a MoE it is the down-projection share of experts in half the layers; the two situations are not the same cost. [source]
- Unsloth-style UD recipes drop or reshape the use-more-bits rule for experts while keeping attention high, versus stock Q4_K_M raising routed `ffn_down_exps`; neither side reports tokens per second per rule. [source]
- block_q6_K stores ql (QK_K/2 bytes), qh (QK_K/4 bytes), 16 int8 scales and one ggml_half d, and a static_assert fixes its size at sizeof(ggml_half) + QK_K/16 + 3*QK_K/4. [source]
- The k-quants discussion lists Q6_K as "type-0" 6-bit with 16 blocks of 16 weights and 8-bit scales at 6.5625 bpw, and Q4_K as "type-1" 4-bit with 8 blocks of 32 weights, 6-bit scales and mins, at 4.5 bpw. [source]
- kernel_mul_mv_q6_K_f32_impl rebuilds each weight as `(q & 0xF) | ((qh & mask) << 4)` (or the high-nibble variant) minus 32, accumulates four partial sums per group against float activations, and multiplies by the int8 scales sc[0], sc[2], sc[4], sc[6] and the fp16 d, with no minimum term. [source]
- In the Q6_K kernel, two lanes per simdgroup (ix = tiisg % 2) stride over super-blocks by 2, each lane holds a 16-element yl array, and quant data is read through `device const uint8_t *` pointers and scales through `int8_t *`. [source]
- The Q4_K kernel reads its nibbles and packed scales through `uint16_t *` pointers and pre-sums activations (sumy) for its minimum term, which the Q6_K kernel does not need. [source]
- ggml-metal-impl.h sets N_R0 and N_SG to 2 and 2 for Q4_K and for Q6_K, 1 and 2 for Q5_K, 2 and 4 for Q8_0 and 4 and 2 for Q4_0. [source]
- ggml-metal-device.cpp selects nsg and nr0 per type for both the dense and the id (MoE) mat-vec pipelines, using N_SG_Q4_K and N_R0_Q4_K for Q4_K and N_SG_Q6_K and N_R0_Q6_K for Q6_K. [source]
- mul_mv.metal registers kernel_mul_mv_id_q4_K_f32 and kernel_mul_mv_id_q6_K_f32 as wrappers over the dense kernel_mul_mv_q4_K_f32_impl and kernel_mul_mv_q6_K_f32_impl, so MoE expert tensors use the same per-type implementations. [source]
- For n_expert > 1, llama-quantize parses the layer from the tensor name and uses use_more_bits per layer, and for Q4_K_M sets Q6_K for every ffn_down-category tensor in qualifying layers. [source]
- PR 1684's table (7B, M2 Max, ms per token at 4 threads) lists Q4_K_S 50, Q4_K_M 55, Q5_K_S 70, Q6_K 75 and F16 116, with file sizes 3.56, 3.80, 4.33, 5.15 and 13.0 GB; at 8 threads the same types are 36, 40, 44, 51 and 111 ms. [source]
- PR 1684's 13B rows on M2 Max give Q4_K_S 95, Q4_K_M 102, Q6_K 142 ms at 4 threads (6.80, 7.32 and 9.95 GB) and 68, 73 and 95 ms at 8 threads. [source]
- PR 1684's RTX 4080 and Ryzen 7950X rows (7B) give Q4_K_S 15.5 and 68 ms, Q4_K_M 16.0 and 71 ms, Q6_K 18.3 and 93 ms, and its perplexities at 7B are 6.0215 (Q4_K_S), 5.9601 (Q4_K_M) and 5.9110 (Q6_K) against 5.9066 for F16. [source]
- On the M2 Max rows, Q4_K_M costs about 10% more time than Q4_K_S (55 versus 50 ms) for 6.7% more bytes, and Q6_K costs 50% more (75 ms) for 45% more bytes, so on that hardware time scaled with bytes and the Q6_K kernel carried no large per-byte penalty. [source]
- Because Q6_K adds 1.458 times the bytes of Q4_K and the kernels share launch shape, a bandwidth-bound Metal decode predicts a Q6_K tensor costs roughly 1.46 times a Q4_K tensor of the same shape, and any measured excess over that is kernel arithmetic. [source]
- In a Q4_K_M MoE the Q6_K fraction of expert bytes is the down-projection share of experts in about half the layers (roughly one-sixth of expert parameters), so the decode penalty of the Q6_K tensors is the byte increase on that sixth, about 7% of expert bytes (less of the whole file), plus any kernel excess. [source]
Children
- No children recorded.