<!-- llms-explorer concept facts · https://llms-explorer.com/tree/localbench-top-40-union-kl-estimator-versus-kld/ · pack 2026-10-05 · ~903 tokens -->

# localbench top-40 union KL estimator versus .kld full-vocab KLD

> Requests go to `/v1/completions` with `echo: true` and `logprobs: 40`, which returns top-40 logprobs for every prompt token.

Parent: [Mac local LLMs: Quantization evaluation](https://llms-explorer.com/tree/mac-local-llms-quantization-evaluation/) · 1 facets · 16 facts · page: https://llms-explorer.com/tree/localbench-top-40-union-kl-estimator-versus-kld/

## Facts

- Requests go to `/v1/completions` with `echo: true` and `logprobs: 40`, which returns top-40 logprobs for every prompt token. — [source](https://localbench.substack.com/p/gguf-benchmark-methodology)
- The chat template comes from the official model release, not GGUF metadata, and is the same for every quant of a model; prompts are rendered through `/v1/internal/chat-prompt`. — [source](https://localbench.substack.com/p/gguf-benchmark-methodology)
- The author justifies top-40 as covering "virtually all the probability mass" with negligible approximation error; no measurement is given. — [source](https://localbench.substack.com/p/gguf-benchmark-methodology)
- Per-token KL is averaged over all tokens and prompts for the headline number; top-1 agreement is the share of tokens where both models' argmax match. — [source](https://localbench.substack.com/p/gguf-benchmark-methodology)
- The inference stack is TextGen with a patched llama.cpp fork that returns prompt logprobs. — [source](https://localbench.substack.com/p/gguf-benchmark-methodology)
- Missing tokens use a floor, so the two truncated distributions need not sum to 1 over the union; whether the author renormalises is not stated. — source: `asserted`
- 13 Apr 2026: methodology post; the 7 Apr Gemma 4 31B post and the 25 Apr Qwen 3.6 27B post link to it as their full methodology. — [source](https://localbench.substack.com/p/gguf-benchmark-methodology)
- Values run higher than Wikipedia-based KLD because inputs reach about 30k tokens across six task types; Q4_K_M numbers of 0.01-0.03 elsewhere are not comparable. — [source](https://localbench.substack.com/p/gguf-benchmark-methodology)
- An estimator with a "min minus 2" floor penalises a quant that moves mass into tokens the reference ranks below 40 differently from a full-vocabulary sum, so ranking can differ in the tail even when means agree. — source: `asserted`
- The patched llama.cpp is a fork, so a mainline build cannot reproduce the numbers without the patch. — [source](https://localbench.substack.com/p/gguf-benchmark-methodology)
- localbench: top-40 error is negligible. The existing clamp dossier: the two estimators are different and absolute values are not interchangeable. Neither has a head-to-head measurement. — [source](https://localbench.substack.com/p/gguf-benchmark-methodology)
- A reader asked on 16 Apr 2026 whether the harness or dataset is downloadable; the visible reply is not shown in the cached page. — [source](https://localbench.substack.com/p/gguf-benchmark-methodology)
- localbench requests top-40 prompt logprobs via `/v1/completions` with `echo: true`. — [source](https://localbench.substack.com/p/gguf-benchmark-methodology)
- localbench uses the official release's Jinja2 template for every quant, not the GGUF's embedded template. — [source](https://localbench.substack.com/p/gguf-benchmark-methodology)
- localbench's top-1 agreement is the fraction of positions where quant and reference pick the same most likely token. — [source](https://localbench.substack.com/p/gguf-benchmark-methodology)
- localbench does not publish a measured error bound for the top-40 truncation. — source: `asserted`
