Q6_K Metal mat-vec cost inside Q4_K_M files

Parent: Mac local LLMs: Speed, bandwidth and prefill · Published reference · snapshot 2026-10-05

↓ Facts as markdownall context files

Bytes: a Q6_K block is 128 B of low nibbles (ql), 64 B of high 2-bit pairs (qh), 16 B of scales and 2 B of d, 210 B per 256 weights or 6.5625 bpw, against 144 B and 4.5 bpw for Q4_K, so a Q6_K tensor moves 1.458 times the bytes of the same tensor in Q4_K.

These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.

Facts

Children

← the whole tree · 3D view· how to read this page