Quantization formats for Apple silicon (GGUF vs MLX)

Parent: Mac local LLMs: Quantization formats and methods · Published reference · snapshot 2026-10-05

↓ Facts as markdownall context files

Bits per weight (arithmetic, [asserted] where not cited): MLX affine 4-bit with default group size 64 and fp16/bf16 scale+bias = 4 + 32/64 = 4.5 bpw; group size 32 = 5.0 bpw. llama.cpp Q4_K = 4.5 bpw (super-block of 8x32, 6-bit scales and mins); Q4_K_M actual file = 4.89 bpw because Q6_K is used ...

These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.

Facts

Corrections and disagreements

Children

← the whole tree · 3D view· how to read this page