Gemma 4 31B GGUF quants ranked by KL divergence
Parent: Mac local LLMs: Quantization evaluation · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
localbench (oobabooga, 7 Apr 2026) benchmarked 53 quants: unsloth 21, bartowski 27, lmstudio-community 3, ggml-org 2, against the unsloth BF16 GGUF.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- localbench (oobabooga, 7 Apr 2026) benchmarked 53 quants: unsloth 21, bartowski 27, lmstudio-community 3, ggml-org 2, against the unsloth BF16 GGUF. [source]
- The dataset is about 250,000 tokens over coding, general chat, tool calling, science, non-Latin scripts and long documents, rendered through the model's chat template; KL is computed on prompt tokens. [source]
- Q8_0 is identical across all four uploaders. [source]
- ggml-org and lmstudio-community quants never reach the Pareto frontier except Q8_0, so the post says to avoid them. [source]
- 8 of 9 unsloth UD quants sit on the frontier; UD-IQ3_XXS is dominated by UD-Q2_K_XL at the same size with lower KL. [source]
- The frontier is split between unsloth and bartowski; neither dominates at all sizes. [source]
- At Q8_0 the per-category KL is 0.466 for long documents, 0.222 for non-Latin scripts, 0.078 for tool calling and 0.069 for science; the ordering holds at every quant size. [source]
- Unsloth's QAT table for 31B (reference: QAT BF16): UD Q4 17.29 GB with mean KLD 0.01403, 99.9% KLD 1.3659, top-1 96.67%; naive llama.cpp Q4_0 17.65 GB with 0.09349, 3.0030, 87.91%. [source]
- Bartowski's 28 Gemma 4 31B GGUF file sizes include Q8_0 30.87 GB, Q6_K 24.89, Q5_K_M 21.06, Q4_K_M 18.25, IQ4_XS 15.98, Q3_K_M 14.82, IQ3_XXS 12.09 and IQ1_M 9.42 (probed 2026-09-02). [source]
- 5 Jun 2026: Unsloth released Gemma 4 QAT GGUFs converted by a different route because a direct llama.cpp Q4_0 conversion loses accuracy from the BF16 QAT weights. [source]
- The conversion loss is F16 scales in llama.cpp against BF16 scales in QAT; naive conversion reaches 24.77% byte exactness and Unsloth reports 99.96%. [source]
- The localbench per-quant plot and Pareto table are images, so no per-quant KL value for the ranking was retrievable as text. [source]
- localbench KL values run higher than Wikipedia 2048-context ones because of long inputs up to about 30k tokens; the author says Q4_K_M at 0.01-0.03 elsewhere is the Wikipedia regime. [source]
- The QAT KLD is measured against the QAT BF16, not the original BF16, so it cannot be set beside the post-training-quant tables; Unsloth says KLD between non-QAT and QAT BF16 is "vastly different" and suggests MMLU-style benchmarks. [source]
- The post asks readers not to share its plots and tables, only the URL. [source]
- The localbench post calls ggml-org and lmstudio-community quants to be avoided, while ModelFit simply recommends bartowski Q4_K_M as the default with no KL data; the two use different evidence. [source]
- Per-quant localbench KL for Gemma 4 31B, and the Q4_K_M and IQ4_XS values needed for a size-to-KLD elasticity. [source]
- Whether the Gemma 4 31B elasticity matches the Qwen3.5-35B-A3B value of about -3.3 in log-log-size-to-kld-elasticity-of-k-quants-acros.md. [source]
- localbench ranked 53 Gemma 4 31B GGUF quants from four uploaders by KL against BF16 (Apr 2026). [source]
- On that ranking unsloth UD quants dominate the frontier (8 of 9) and ggml-org and lmstudio-community quants are off it apart from Q8_0. [source]
- Gemma 4 31B Q8_0 KL on the localbench set is 0.466 for long documents and 0.069 for science. [source]
- Unsloth's Gemma 4 31B QAT UD Q4 (17.29 GB) has mean KLD 0.01403 against the QAT BF16 and the naive Q4_0 has 0.09349. [source]
- QAT-reference KLD and PTQ-reference KLD are not comparable. [source]
- The reddit thread "Gemma 4 31B GGUF quants ranked by KL divergence" is not fetchable through the helper. [source]
Children
- No children recorded.