Tail-risk quant metrics 99.9% KLD max KLD top-1 agreement
Parent: Mac local LLMs: Quantization evaluation · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Top-1 disagreement is a discontinuous winner test: a tiny KL change can flip a 50.1/49.9 decision, and a much larger KL change can leave a dominant winner unchanged. KL and top-1 therefore rank quants differently at the margin.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Top-1 disagreement is a discontinuous winner test: a tiny KL change can flip a 50.1/49.9 decision, and a much larger KL change can leave a dominant winner unchanged. KL and top-1 therefore rank quants differently at the margin. [source]
- KL is computed over the whole distribution, including movement that leaves the winner unchanged; top-1 only sees the winner. [source]
- Tail statistics can be taken globally or per output range. A "worst-range p95 KLD" takes the p95 within each range and reports the maximum across ranges, which exposes a bad span that a global mean hides. [source]
- Conditioning flips on baseline confidence separates harmless flips (reference was uncertain) from damaging ones (reference was confident); the second kind drives literal-copy errors. [source]
- Qwen3.5 era (2026-02/03): Unsloth selected quants by 99.9% KLD and then Max KLD (existing file). [source]
- Qwen3.8 Dynamic v3.0 era (2026-08-19): Unsloth reports top-1 and mean KLD for all providers plus Divergence-300 @32, and cites Kimi-K3 Dynamic 1-bit at ~78.9% top-1 while 62% smaller. The headline metric moved from tail KLD to top-1 agreement within five months. [source]
- 2026-08-20: an independent Level1Techs analysis reports worst-range p95 KLD and top-1 flip rate for four Qwen3.8-27B derivatives. [source]
- Max KLD is a single-token statistic; a quant with a lower max can still have a worse 99.9% value, and the reverse. Existing numbers show both moving, but no source reports the correlation across many quants. [source]
- Reporting only mean top-1 hides clustering: flip rates vary from under 1% to several percent depending on the workstream captured. [source]
- High-confidence flips on exact literals (ports, hostnames, argument names) are the damaging tail, and they are invisible to KLD averages. [source]
- A quantized model can diverge from BF16 and be semantically better on a given answer; BF16 is a fidelity reference, not a correctness label. [source]
- Top-1 versus KLD as the primary tail signal: a forum commenter argues greedy top-1 is the stricter and more output-relevant measure; the harness author says top-1 and KL measure different things and keeps both. Unsloth's v3 page foregrounds top-1 agreement at equal size and treats KLD mean as secondary. No source tests which one predicts task pass rate on a Mac. [source]
- Unsloth v3 headline ">10% higher top-1 at the same size versus every other provider" is vendor-run; a user in the pinned thread asked for the evaluation dataset and methodology behind the top-1 and KLD figures. [source]
- Which tail statistic (max, 99.9%, worst-range p95, top-1 flip rate on confident tokens) best predicts failed tool calls for Mac quants (GGUF or MLX). The only data is CUDA derivative models. [source]
- Whether a confidence-conditioned flip rate can be computed from llama-perplexity's saved logits file without a new harness. [source]
- Top-1 disagreement is changed-winner positions divided by evaluated positions, and a 50.1/49.9 decision can flip on an extremely small KL change while a much larger KL change can leave a dominant winner unchanged. [source]
- A forum reply states greedy top-1 is stricter than KL and more relevant to what a user receives as output. [source]
- The Level1Techs harness retains reverse KL and Jensen-Shannon divergence alongside KL(P_BF16 || P_candidate) rather than one direction only. [source]
- Worst-range p95 KLD for four Qwen3.8-27B derivatives on SP06 was 0.03661 (Heretic-ARA), 0.04477 (Huihui), 0.31235 (Blackfrost) and 0.68507 (AEON). [source]
- Most flips occur where the reference was less confident; the two worst derivatives also overturned confident reference tokens (AEON: 24 strongly preferred stock decisions). [source]
- The author says BF16 is a numerical-fidelity reference, not an oracle, and a quantized model can diverge and still be semantically better. [source]
- Unsloth's Dynamic v3.0 page says it generally reports top-1 accuracy and that top-1 "is an argmax on 1 prediction, so it's not really effective on gauging actual inference". [source]
- Unsloth's Dynamic v3.0 page reports Top-1 and mean KLD for all providers, and cites Kimi-K3 Dynamic 1-bit at about 78.9% top-1 while 62% smaller. [source]
- Unsloth's Dynamic v3.0 page says its calibration data was kept disjoint from held-out test data and that it uses pure PTQ, not QAD or QAT, so overfitting is less of a concern. [source]
- A Hugging Face user asked Unsloth to share the evaluation dataset and methodology used for top-1 agreement and KLD in the Qwen3.8-27B v3 thread. [source]
- Unsloth's headline metric shifted from 99.9% and Max KLD (Qwen3.5, March 2026) to top-1, mean KLD and Divergence-300 @32 (Qwen3.8 v3, August 2026). [source]
- Across four weight-edited derivatives, flip rate, worst-range p95 KLD and invalid-branch count share one ranking, which is weak evidence (n=4, not quants) that a per-range tail KLD tracks task damage. [source]
- For Mac quant selection, report 99.9% KLD or worst-range p95 together with top-1 flips on tokens where the reference probability exceeds 0.9, since those flips are the literal-copy failures. [source]
Children
- No children recorded.