<!-- llms-explorer concept facts · https://llms-explorer.com/tree/gemma-4-31b-gguf-quants-ranked-by-kl-divergence/ · pack 2026-10-05 · ~1349 tokens -->

# Gemma 4 31B GGUF quants ranked by KL divergence

> localbench (oobabooga, 7 Apr 2026) benchmarked 53 quants: unsloth 21, bartowski 27, lmstudio-community 3, ggml-org 2, against the unsloth BF16 GGUF.

Parent: [Mac local LLMs: Quantization evaluation](https://llms-explorer.com/tree/mac-local-llms-quantization-evaluation/) · 1 facets · 24 facts · page: https://llms-explorer.com/tree/gemma-4-31b-gguf-quants-ranked-by-kl-divergence/

## Facts

- localbench (oobabooga, 7 Apr 2026) benchmarked 53 quants: unsloth 21, bartowski 27, lmstudio-community 3, ggml-org 2, against the unsloth BF16 GGUF. — [source](https://localbench.substack.com/p/gemma-4-31b-gguf-kl-divergence)
- The dataset is about 250,000 tokens over coding, general chat, tool calling, science, non-Latin scripts and long documents, rendered through the model's chat template; KL is computed on prompt tokens. — [source](https://localbench.substack.com/p/gguf-benchmark-methodology)
- Q8_0 is identical across all four uploaders. — [source](https://localbench.substack.com/p/gemma-4-31b-gguf-kl-divergence)
- ggml-org and lmstudio-community quants never reach the Pareto frontier except Q8_0, so the post says to avoid them. — [source](https://localbench.substack.com/p/gemma-4-31b-gguf-kl-divergence)
- 8 of 9 unsloth UD quants sit on the frontier; UD-IQ3_XXS is dominated by UD-Q2_K_XL at the same size with lower KL. — [source](https://localbench.substack.com/p/gemma-4-31b-gguf-kl-divergence)
- The frontier is split between unsloth and bartowski; neither dominates at all sizes. — [source](https://localbench.substack.com/p/gemma-4-31b-gguf-kl-divergence)
- At Q8_0 the per-category KL is 0.466 for long documents, 0.222 for non-Latin scripts, 0.078 for tool calling and 0.069 for science; the ordering holds at every quant size. — [source](https://localbench.substack.com/p/gemma-4-31b-gguf-kl-divergence)
- Unsloth's QAT table for 31B (reference: QAT BF16): UD Q4 17.29 GB with mean KLD 0.01403, 99.9% KLD 1.3659, top-1 96.67%; naive llama.cpp Q4_0 17.65 GB with 0.09349, 3.0030, 87.91%. — [source](https://unsloth.ai/docs/models/gemma-4/qat)
- Bartowski's 28 Gemma 4 31B GGUF file sizes include Q8_0 30.87 GB, Q6_K 24.89, Q5_K_M 21.06, Q4_K_M 18.25, IQ4_XS 15.98, Q3_K_M 14.82, IQ3_XXS 12.09 and IQ1_M 9.42 (probed 2026-09-02). — [source](https://modelfit.io/quant-compare/gemma4-31b/)
- 5 Jun 2026: Unsloth released Gemma 4 QAT GGUFs converted by a different route because a direct llama.cpp Q4_0 conversion loses accuracy from the BF16 QAT weights. — [source](https://huggingface.co/unsloth/gemma-4-31B-it-qat-GGUF/discussions/1)
- The conversion loss is F16 scales in llama.cpp against BF16 scales in QAT; naive conversion reaches 24.77% byte exactness and Unsloth reports 99.96%. — [source](https://unsloth.ai/docs/models/gemma-4/qat)
- The localbench per-quant plot and Pareto table are images, so no per-quant KL value for the ranking was retrievable as text. — [source](https://localbench.substack.com/p/gemma-4-31b-gguf-kl-divergence)
- localbench KL values run higher than Wikipedia 2048-context ones because of long inputs up to about 30k tokens; the author says Q4_K_M at 0.01-0.03 elsewhere is the Wikipedia regime. — [source](https://localbench.substack.com/p/gguf-benchmark-methodology)
- The QAT KLD is measured against the QAT BF16, not the original BF16, so it cannot be set beside the post-training-quant tables; Unsloth says KLD between non-QAT and QAT BF16 is "vastly different" and suggests MMLU-style benchmarks. — [source](https://huggingface.co/unsloth/gemma-4-31B-it-qat-GGUF/discussions/1)
- The post asks readers not to share its plots and tables, only the URL. — [source](https://localbench.substack.com/p/qwen-3-6-27b-gguf-quality-benchmark)
- The localbench post calls ggml-org and lmstudio-community quants to be avoided, while ModelFit simply recommends bartowski Q4_K_M as the default with no KL data; the two use different evidence. — [source](https://modelfit.io/quant-compare/gemma4-31b/)
- Per-quant localbench KL for Gemma 4 31B, and the Q4_K_M and IQ4_XS values needed for a size-to-KLD elasticity. — source: `asserted`
- Whether the Gemma 4 31B elasticity matches the Qwen3.5-35B-A3B value of about -3.3 in log-log-size-to-kld-elasticity-of-k-quants-acros.md. — source: `asserted`
- localbench ranked 53 Gemma 4 31B GGUF quants from four uploaders by KL against BF16 (Apr 2026). — [source](https://localbench.substack.com/p/gemma-4-31b-gguf-kl-divergence)
- On that ranking unsloth UD quants dominate the frontier (8 of 9) and ggml-org and lmstudio-community quants are off it apart from Q8_0. — [source](https://localbench.substack.com/p/gemma-4-31b-gguf-kl-divergence)
- Gemma 4 31B Q8_0 KL on the localbench set is 0.466 for long documents and 0.069 for science. — [source](https://localbench.substack.com/p/gemma-4-31b-gguf-kl-divergence)
- Unsloth's Gemma 4 31B QAT UD Q4 (17.29 GB) has mean KLD 0.01403 against the QAT BF16 and the naive Q4_0 has 0.09349. — [source](https://unsloth.ai/docs/models/gemma-4/qat)
- QAT-reference KLD and PTQ-reference KLD are not comparable. — [source](https://huggingface.co/unsloth/gemma-4-31B-it-qat-GGUF/discussions/1)
- The reddit thread "Gemma 4 31B GGUF quants ranked by KL divergence" is not fetchable through the helper. — [source](https://www.reddit.com/r/LocalLLaMA/comments/1seua77/gemma_4_31b_gguf_quants_ranked_by_kl_divergence/)
