<!-- llms-explorer concept facts · https://llms-explorer.com/tree/tail-risk-quant-metrics-99-9-kld-max-kld-top-1-a/ · pack 2026-10-05 · ~1658 tokens -->

# Tail-risk quant metrics 99.9% KLD max KLD top-1 agreement

> Top-1 disagreement is a discontinuous winner test: a tiny KL change can flip a 50.1/49.9 decision, and a much larger KL change can leave a dominant winner unchanged. KL and top-1 therefore rank quants differently at the margin.

Parent: [Mac local LLMs: Quantization evaluation](https://llms-explorer.com/tree/mac-local-llms-quantization-evaluation/) · 1 facets · 28 facts · page: https://llms-explorer.com/tree/tail-risk-quant-metrics-99-9-kld-max-kld-top-1-a/

## Facts

- Top-1 disagreement is a discontinuous winner test: a tiny KL change can flip a 50.1/49.9 decision, and a much larger KL change can leave a dominant winner unchanged. KL and top-1 therefore rank quants differently at the margin. — source: `asserted`
- KL is computed over the whole distribution, including movement that leaves the winner unchanged; top-1 only sees the winner. — source: `asserted`
- Tail statistics can be taken globally or per output range. A "worst-range p95 KLD" takes the p95 within each range and reports the maximum across ranges, which exposes a bad span that a global mean hides. — source: `asserted`
- Conditioning flips on baseline confidence separates harmless flips (reference was uncertain) from damaging ones (reference was confident); the second kind drives literal-copy errors. — source: `asserted`
- Qwen3.5 era (2026-02/03): Unsloth selected quants by 99.9% KLD and then Max KLD (existing file). — source: `asserted`
- Qwen3.8 Dynamic v3.0 era (2026-08-19): Unsloth reports top-1 and mean KLD for all providers plus Divergence-300 @32, and cites Kimi-K3 Dynamic 1-bit at ~78.9% top-1 while 62% smaller. The headline metric moved from tail KLD to top-1 agreement within five months. — source: `asserted`
- 2026-08-20: an independent Level1Techs analysis reports worst-range p95 KLD and top-1 flip rate for four Qwen3.8-27B derivatives. — source: `asserted`
- Max KLD is a single-token statistic; a quant with a lower max can still have a worse 99.9% value, and the reverse. Existing numbers show both moving, but no source reports the correlation across many quants. — source: `asserted`
- Reporting only mean top-1 hides clustering: flip rates vary from under 1% to several percent depending on the workstream captured. — source: `asserted`
- High-confidence flips on exact literals (ports, hostnames, argument names) are the damaging tail, and they are invisible to KLD averages. — source: `asserted`
- A quantized model can diverge from BF16 and be semantically better on a given answer; BF16 is a fidelity reference, not a correctness label. — source: `asserted`
- Top-1 versus KLD as the primary tail signal: a forum commenter argues greedy top-1 is the stricter and more output-relevant measure; the harness author says top-1 and KL measure different things and keeps both. Unsloth's v3 page foregrounds top-1 agreement at equal size and treats KLD mean as secondary. No source tests which one predicts task pass rate on a Mac. — source: `asserted`
- Unsloth v3 headline ">10% higher top-1 at the same size versus every other provider" is vendor-run; a user in the pinned thread asked for the evaluation dataset and methodology behind the top-1 and KLD figures. — source: `asserted`
- Which tail statistic (max, 99.9%, worst-range p95, top-1 flip rate on confident tokens) best predicts failed tool calls for Mac quants (GGUF or MLX). The only data is CUDA derivative models. — source: `asserted`
- Whether a confidence-conditioned flip rate can be computed from llama-perplexity's saved logits file without a new harness. — source: `asserted`
- Top-1 disagreement is changed-winner positions divided by evaluated positions, and a 50.1/49.9 decision can flip on an extremely small KL change while a much larger KL change can leave a dominant winner unchanged. — [source](https://forum.level1techs.com/t/why-your-local-llm-feels-dumber-than-it-is/253917)
- A forum reply states greedy top-1 is stricter than KL and more relevant to what a user receives as output. — [source](https://forum.level1techs.com/t/why-your-local-llm-feels-dumber-than-it-is/253917)
- The Level1Techs harness retains reverse KL and Jensen-Shannon divergence alongside KL(P_BF16 || P_candidate) rather than one direction only. — [source](https://forum.level1techs.com/t/why-your-local-llm-feels-dumber-than-it-is/253917)
- Worst-range p95 KLD for four Qwen3.8-27B derivatives on SP06 was 0.03661 (Heretic-ARA), 0.04477 (Huihui), 0.31235 (Blackfrost) and 0.68507 (AEON). — [source](https://forum.level1techs.com/t/why-your-local-llm-feels-dumber-than-it-is/253917)
- Most flips occur where the reference was less confident; the two worst derivatives also overturned confident reference tokens (AEON: 24 strongly preferred stock decisions). — [source](https://forum.level1techs.com/t/why-your-local-llm-feels-dumber-than-it-is/253917)
- The author says BF16 is a numerical-fidelity reference, not an oracle, and a quantized model can diverge and still be semantically better. — [source](https://forum.level1techs.com/t/why-your-local-llm-feels-dumber-than-it-is/253917)
- Unsloth's Dynamic v3.0 page says it generally reports top-1 accuracy and that top-1 "is an argmax on 1 prediction, so it's not really effective on gauging actual inference". — [source](https://unsloth.ai/docs/basics/dynamic-3.0-ggufs)
- Unsloth's Dynamic v3.0 page reports Top-1 and mean KLD for all providers, and cites Kimi-K3 Dynamic 1-bit at about 78.9% top-1 while 62% smaller. — [source](https://unsloth.ai/docs/basics/dynamic-3.0-ggufs)
- Unsloth's Dynamic v3.0 page says its calibration data was kept disjoint from held-out test data and that it uses pure PTQ, not QAD or QAT, so overfitting is less of a concern. — [source](https://unsloth.ai/docs/basics/dynamic-3.0-ggufs)
- A Hugging Face user asked Unsloth to share the evaluation dataset and methodology used for top-1 agreement and KLD in the Qwen3.8-27B v3 thread. — [source](https://huggingface.co/unsloth/Qwen3.8-27B-GGUF/discussions/74)
- Unsloth's headline metric shifted from 99.9% and Max KLD (Qwen3.5, March 2026) to top-1, mean KLD and Divergence-300 @32 (Qwen3.8 v3, August 2026). — source: `asserted`
- Across four weight-edited derivatives, flip rate, worst-range p95 KLD and invalid-branch count share one ranking, which is weak evidence (n=4, not quants) that a per-range tail KLD tracks task damage. — source: `asserted`
- For Mac quant selection, report 99.9% KLD or worst-range p95 together with top-1 flips on tokens where the reference probability exceeds 0.9, since those flips are the literal-copy failures. — source: `asserted`
