Mac local LLMs: Quantization formats and methods

Parent: Running LLM models locally on a Mac · Published reference · snapshot 2026-10-05

↓ Facts as markdownall context files

MLX affine 4-bit group 64 = ~4.5 bpw; Q4_K = 4.5 bpw but Q4_K_M files are 4.89 bpw (Q6_K on half of attn v/ffn down). MXFP4 4.25.

These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.

Format basics and choice

mlx-lm recipes and flags

fp16 vs bf16 on M1/M2

Unsloth Dynamic and GGUF recipes

MoE protection and data-driven MLX quants

Gemma 4 QAT

Ternary Bonsai (Prism) GGUF

Re-uploads and cache hygiene

Corrections to earlier claims

Open questions

Corrections and disagreements

Concepts in this cluster

Children

← the whole tree · 3D view· how to read this page