Gemma 4 QAT GGUF conversion: F16 versus BF16 scales in llama.cpp Q4_0
Parent: Mac local LLMs: Quantization formats and methods · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Unsloth names two causes: the scale format differs (F16 versus BF16) and "the scales are not determined optimally in llama.cpp land".
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Unsloth names two causes: the scale format differs (F16 versus BF16) and "the scales are not determined optimally in llama.cpp land". [source]
- Its remedy applies the Unsloth dynamic method to force agreement between the llama.cpp Q4_0 lattice and the true BF16 QAT lattice; embeddings then did not need Q6_K, so files shrank. [source]
- A llama.cpp discussion (#26788, 9 Aug 2026) asks for optional BF16 scales in the Q4_0 GGUF format and inference with them. [source]
- The only upstream movement cited is PR 29675 (29 Sep 2026), which adds BF16 variants of unary, GLU, binary and scale ops on CPU and CUDA so activations can later stay in BF16; it adds no BF16 Q4_0 scale support. [source]
- 9 Aug 2026: discussion #26788 opened by Craftyawesome quoting Unsloth's analysis and noting the workaround still leaves mean KLD up to 0.13288 and top-1 as low as 85.63%. [source]
- 2 Sep 2026: a second commenter reports Gemma 4 QAT 26B degrading against regular Q4_K on private tests, suspects token embeddings and attention tensors held in Q4_0, and asks whether inference in BF16 can be offered; BF16 KV cache already works on hardware without BF16 support. [source]
- 2 Oct 2026: the author says the issue is "partially addressed by #29675". [source]
- The two worst figures quoted in the discussion come from different rows: 0.13288 mean KLD is the 12B Unsloth file and 85.63% top-1 is the 26B A4B Unsloth file. [source]
- 12B has the highest 99.9% KLD of the table at 9.2740 even after the workaround. [source]
- The discussion's reading that mean KLD of 0.13 means "not lossless" ignores that the reference is the QAT BF16, not the pre-QAT model. [source]
- Unsloth says the remedy yields quants that are "more accurate"; the discussion participants say residual KLD and low top-1 show the conversion is still lossy and want a format change. Both are consistent with the table: improvement over naive Q4_0 is large, yet absolute values for 12B and 26B stay high. [source]
- Unsloth's QAT table (KLD against QAT BF16): E2B 2.62 GB mean KLD 0.00173 top-1 98.16% versus naive Q4_0 3.35 GB 0.05109 89.29%. [source]
- E4B: Unsloth 4.22 GB 0.00121 and 98.54% versus Q4_0 5.15 GB 0.03778 and 90.94%. [source]
- 26B A4B: Unsloth 14.25 GB 0.09788 and 85.63% versus Q4_0 14.44 GB 0.36094 and 70.20%. [source]
- 12B: Unsloth 6.72 GB 0.13288 and 88.76% with 99.9% KLD 9.2740. [source]
- Unsloth says E2B's mean KLD is 29x lower than naive Q4_0 and its file 22% smaller. [source]
- Google released mobile-mixture QAT checkpoints for E2B and E4B; Unsloth converts the 2-bit layers with TQ2_0 and "a negative scaler" into UD-Q2_K_XL files. [source]
- E2B mobile: 2.19 GB, 61 TQ2_0 tensors including deep MLP, mean KLD 0.00409; E4B mobile: 3.22 GB, 2 TQ2_0 tensors (embeddings), mean KLD 0.00102. [source]
- Unsloth recommends 8-bit for the small models and Dynamic 4-bit for larger ones, and publishes 12B, 26B-A4B and 31B QAT GGUF repos besides E2B and E4B. [source]
- PR 29675 adds BF16 elementwise ops only and carries an AI-usage disclosure ("paired with claude"). [source]
- llama.cpp discussion #26788 is in the Ideas category with two comments and no maintainer reply as of 4 Oct 2026. [source]
Children
- No children recorded.