size-matched KLD reporting for local quants
Parent: Mac local LLMs: Quantization evaluation · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
llama.cpp's own scoreboard lists model size in GiB beside KLD for 40+ GGUF types of LLaMA 3 8B against an FP16 reference, sorted by KLD. A size-matched read needs no extra run.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- llama.cpp's own scoreboard lists model size in GiB beside KLD for 40+ GGUF types of LLaMA 3 8B against an FP16 reference, sorted by KLD. A size-matched read needs no extra run. [source]
- Equal-size pairs from that table (no imatrix unless stated): q5_0 and q5_K_S are both 5.21 GiB with KLD 0.022239 and 0.016595, so legacy q5_0 is 1.34 times worse at identical size. [source]
- q4_0 is 4.34 GiB with KLD 0.071940 while q4_K_S is 4.37 GiB with 0.043136: q4_0 is 0.7% smaller and 67% worse (q4_K_S is 40% lower). [source]
- Size ordering can invert quality ordering: q4_1 (4.78 GiB, KLD 0.071683) is 4.4% larger than q4_K_M (4.58 GiB, 0.031273) and 2.29 times worse; q5_1 (5.65 GiB, 0.018045) is 6.0% larger than q5_K_M (5.33 GiB, 0.010762) and 1.68 times worse. [source]
- A smaller file can tie a larger one: q3_K_L with the WT 10m imatrix (4.03 GiB, 0.073077) is 7.1% smaller than q4_0 (4.34 GiB, 0.071940) and within 1.6% of its KLD. [source]
- The importance matrix is a second axis at fixed size: q4_K_S at 4.37 GiB drops from 0.043136 (none) to 0.031951 (WT 10m), and q2_K at 2.96 GiB drops from 0.445132 (none) to 0.332223 (WT 10m), about 25%. Size-matching without matching imatrix status compares recipes, not formats. [source]
- The size of the imatrix corpus barely matters there: q2_K with 1k to 10m Wikitext tokens spans KLD 0.3321 to 0.3371, and the README says it sees no consistent improvement from more tokens. [source]
- Elasticity rule (derived): from q5_K_M to q4_K_M (5.33 to 4.58 GiB, -14%) KLD rises 2.9 times, and from q4_K_M to q3_K_M (4.58 to 3.74 GiB, -18%) it rises 3.3 times. The log-log slope is about -6 to -7, so on LLaMA 3 8B a 1% size gap corresponds to roughly 6 to 7% in KLD between 3- and 5-bit K-quants. [source]
- Consequence: a 3% size gap can masquerade as a 20% KLD gap, which is the size of difference vendors report between their quants and others. Report KLD together with size difference in percent, and treat a KLD gap smaller than the elasticity times the size gap as explained by size. [source]
- Baselines are large: the reference logits file for one 14B model on a 1.6 MiB corpus was about 55 GiB, and the guide suggests Q8_0 as the baseline if BF16 cannot be run. A Q8_0 baseline adds its own divergence floor (LLaMA 3 8B q8_0 mean KLD 0.001355), which is small beside 4-bit gaps but not beside 6-bit and 8-bit comparisons. [source]
- Record size in the llama.cpp record format: 2*((n_vocab+1)/2)+4 uint16 per scored position, about 256 KB per position at a 128,256-token vocabulary and about 304 KB at 151,936. A 55 GiB file is therefore about 190,000 scored positions at the larger vocabulary. [source]
- 2024-01 onward: llama.cpp adds the KL-divergence mode (PR 5076, per the existing methodology dossier) and publishes LLaMA 2 and 3 scoreboards with sizes. [source]
- 2025-05: ik_llama.cpp quant-cooking guide documents a two-pass KLD workflow on a custom corpus. [source]
- 2026-02 to 2026-09: community and vendor tables move to size-adjusted scores (existing dossiers). [source]
- Elasticity is model-, family- and corpus-specific; the -6 to -7 figure comes from one 8B dense model on Wikitext and is a guide to magnitude, not a constant. [source]
- A table sorted by KLD with sizes in a separate column invites reading rank as quality; sort by size, plot the frontier, and mark dominated points. [source]
- Matching "within 1%" by file size can still mismatch loaded bytes: MLX files carry unquantized embeddings or vision towers (existing dossier), and GGUF mmproj files are separate. [source]
- Mean KLD and tail KLD can order the same two quants differently at equal size (existing Unsloth versus bartowski case), so a size-matched report needs both. [source]
- Equal size or equal label: vendors who tune tensor allocation argue for equal-size comparison (Unsloth), while critics compare same-name quants where the vendor file is smaller (existing dossier). The README data shows both can be right because a few percent in bytes moves KLD by tens of percent. [source]
- Mean KLD versus a size-adjusted composite score: a composite normalized against one baseline changes when the baseline changes; a frontier plot does not. [source]
- A KLD-versus-GiB frontier for GGUF and MLX quants of one BF16 base on one token stream on Apple silicon. None found. [source]
- Whether the log-log elasticity is stable across dense and MoE models and across 2- to 8-bit. [source]
- Whether KLD from Wikitext and KLD on agent traces rank equal-size quants the same way. [source]
- The llama.cpp LLaMA 3 8B scoreboard lists model size in GiB and KLD for each GGUF type against an FP16 reference, with imatrix variants. [source]
- In that scoreboard q5_0 and q5_K_S are both 5.21 GiB with KLD 0.022239 and 0.016595. [source]
- In that scoreboard q4_0 is 4.34 GiB at KLD 0.071940 and q4_K_S is 4.37 GiB at 0.043136. [source]
- In that scoreboard q4_1 is 4.78 GiB at KLD 0.071683, larger and worse than q4_K_M at 4.58 GiB and 0.031273. [source]
- In that scoreboard q5_1 is 5.65 GiB at KLD 0.018045, larger and worse than q5_K_M at 5.33 GiB and 0.010762. [source]
- In that scoreboard q3_K_L with the WT 10m imatrix is 4.03 GiB at KLD 0.073077, 7.1% smaller than q4_0 at nearly equal KLD. [source]
- In that scoreboard q4_K_S improves from KLD 0.043136 to 0.031951 and q2_K from 0.445132 to 0.332223 when an imatrix is added at unchanged file size. [source]
- The README states the f16 row shows only the difference from casting reference logits to 16-bit integers, with KLD 0.000551 on LLaMA 3 8B. [source]
- The ik_llama.cpp guide reports a KLD baseline file of about 55 GiB for Qwen3-14B BF16 on a 1.6 MiB corpus and suggests Q8_0 as baseline if BF16 cannot run. [source]
- On LLaMA 3 8B, KLD rises about 2.9 times from q5_K_M to q4_K_M and about 3.3 times from q4_K_M to q3_K_M, for size drops of 14% and 18%. [source]
- A size mismatch of a few percent between two 3- to 5-bit K-quants can explain a KLD difference of 20% or more. [source]
- A size-matched report should give size difference in percent, imatrix status, reference precision, corpus, and both mean and tail KLD. [source]
Children
- No children recorded.