Isolated Metal microbenchmark of Q6_K vs Q4_K mat-vec at matched bytes

Parent: Mac local LLMs: Speed, bandwidth and prefill · Published reference · snapshot 2026-10-05

↓ Facts as markdownall context files

Dispatch gate: ggml_metal_op_mul_mat picks the small-batch mat-vec kernels first when src1 is F32 and ne00 % 128 == 0, for batch (ne11) 2 to 8 on F32, F16, BF16, Q4_0, Q4_1, Q5_0, Q5_1, Q8_0, MXFP4 and IQ4_NL, but only for ne11 4 to 8 on Q2_K to Q6_K, using nsg 2.

These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.

Facts

Children

← the whole tree · 3D view· how to read this page