Mac local LLMs: Quantization evaluation

Parent: Running LLM models locally on a Mac · Published reference · snapshot 2026-10-05

↓ Facts as markdownall context files

Run reference once: `llama-perplexity -m ref --kl-divergence-base f.kld -f corpus`; then each quant with `--kl-divergence-base f.kld --kl-divergence` (ignores `-f`, same tokens). Reuse the reference `-c`; keep `-b` equal.

These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.

llama.cpp KLD workflow

MLX harnesses

Size and bits confounds

Calibration and imatrix

Tail and flip metrics

Determinism and cross-engine parity

BFCL and tool calls

localbench top-40 benchmarks (Gemma 4 31B, Qwen3.6 27B)

Corrections

Corrections and disagreements

Concepts in this cluster

Children

← the whole tree · 3D view· how to read this page