<!-- llms-explorer concept facts · https://llms-explorer.com/tree/isolated-metal-microbenchmark-of-q6-k-vs-q4-k-ma/ · pack 2026-10-05 · ~740 tokens -->

# Isolated Metal microbenchmark of Q6_K vs Q4_K mat-vec at matched bytes

> Dispatch gate: ggml_metal_op_mul_mat picks the small-batch mat-vec kernels first when src1 is F32 and ne00 % 128 == 0, for batch (ne11) 2 to 8 on F32, F16, BF16, Q4_0, Q4_1, Q5_0, Q5_1, Q8_0, MXFP4 and IQ4_NL, but only for ne11 4 to 8 on Q2_K to Q6_K, using nsg 2.

Parent: [Mac local LLMs: Speed, bandwidth and prefill](https://llms-explorer.com/tree/mac-local-llms-speed-bandwidth-and-prefill/) · 1 facets · 10 facts · page: https://llms-explorer.com/tree/isolated-metal-microbenchmark-of-q6-k-vs-q4-k-ma/

## Facts

- Dispatch gate: ggml_metal_op_mul_mat picks the small-batch mat-vec kernels first when src1 is F32 and ne00 % 128 == 0, for batch (ne11) 2 to 8 on F32, F16, BF16, Q4_0, Q4_1, Q5_0, Q5_1, Q8_0, MXFP4 and IQ4_NL, but only for ne11 4 to 8 on Q2_K to Q6_K, using nsg 2. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/ggml/src/ggml-metal/ggml-metal-ops.cpp)
- So at batch 2 or 3 (a two-token speculative verify) a Q6_K or Q4_K tensor misses the small-batch kernel and takes the general mat-mat path, while a Q8_0 tensor of the same shape stays on the mat-vec path. — source: `asserted`
- Unsloth's per-type table for Qwen3.5-35B-A3B (hardware not stated) gives tg128 of 90.38 for q4_k, 90.50 for q6_k and 90.33 for q8_0, within 0.2%, though q8_0 moves 1.9 times the bytes of q4_k. That flatness shows decode in that setup is not scaling with weight bytes, so a byte-ratio prediction (1.46 for Q6_K over Q4_K) cannot be checked from it. — [source](https://unsloth.ai/docs/models/qwen3.5/gguf-benchmarks)
- An end-to-end tg number on a 3B-active MoE can sit on a launch or overhead floor, so equal tg for Q4_K, Q6_K and Q8_0 does not prove equal kernel cost. — source: `asserted`
- A valid microbenchmark must use shapes with ne00 % 256 == 0, F32 activations, and a batch of 1 (the mat-vec path) separately from 4 to 8. — source: `asserted`
- "Q6_K costs its byte ratio on a bandwidth-bound GPU" (covering file) versus a flat tg across q4_k, q6_k and q8_0 (Unsloth table): the two predict different outcomes and no source reconciles them. — source: `asserted`
- A per-type kernel timing (Metal capture or a test-backend-ops style op perf run) of MUL_MAT for Q4_K, Q6_K, Q8_0 at fixed shape on one M-series chip. — source: `asserted`
- The Metal small-batch mat-vec gate admits Q8_0 and the legacy types at batch 2 to 8 and K-quants only at batch 4 to 8, and requires F32 src1 and ne00 divisible by 128. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/ggml/src/ggml-metal/ggml-metal-ops.cpp)
- Unsloth's Qwen3.5-35B-A3B tg128 column shows q4_k 90.38, q6_k 90.50, q8_0 90.33, mxfp4 90.67 and q2_k 90.77, while iq3_xxs, iq2_xxs and iq3_s sit near 84 to 86. — [source](https://unsloth.ai/docs/models/qwen3.5/gguf-benchmarks)
- No isolated Metal Q6_K versus Q4_K microbenchmark was found in any cached source. — source: `asserted`
