Q4_K_M dequantization cost vs MLX affine g64 on dense decode

Parent: Mac local LLMs: Speed, bandwidth and prefill · Published reference · snapshot 2026-10-05

↓ Facts as markdownall context files

Q4_K Metal mat-vec: per 256-weight super-block a simdgroup loads 4 groups of activations (32 weights each) pre-summed for the minimum correction; per row it unpacks the 12 scale bytes with three masks (0x3f3f, 0x0f0f, 0xc0c0) into 8 scale and min values, accumulates nibble products without shifti...

These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.

Facts

Children

← the whole tree · 3D view· how to read this page