localbench top-40 union KL estimator versus .kld full-vocab KLD
Parent: Mac local LLMs: Quantization evaluation · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Requests go to `/v1/completions` with `echo: true` and `logprobs: 40`, which returns top-40 logprobs for every prompt token.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Requests go to `/v1/completions` with `echo: true` and `logprobs: 40`, which returns top-40 logprobs for every prompt token. [source]
- The chat template comes from the official model release, not GGUF metadata, and is the same for every quant of a model; prompts are rendered through `/v1/internal/chat-prompt`. [source]
- The author justifies top-40 as covering "virtually all the probability mass" with negligible approximation error; no measurement is given. [source]
- Per-token KL is averaged over all tokens and prompts for the headline number; top-1 agreement is the share of tokens where both models' argmax match. [source]
- The inference stack is TextGen with a patched llama.cpp fork that returns prompt logprobs. [source]
- Missing tokens use a floor, so the two truncated distributions need not sum to 1 over the union; whether the author renormalises is not stated. [source]
- 13 Apr 2026: methodology post; the 7 Apr Gemma 4 31B post and the 25 Apr Qwen 3.6 27B post link to it as their full methodology. [source]
- Values run higher than Wikipedia-based KLD because inputs reach about 30k tokens across six task types; Q4_K_M numbers of 0.01-0.03 elsewhere are not comparable. [source]
- An estimator with a "min minus 2" floor penalises a quant that moves mass into tokens the reference ranks below 40 differently from a full-vocabulary sum, so ranking can differ in the tail even when means agree. [source]
- The patched llama.cpp is a fork, so a mainline build cannot reproduce the numbers without the patch. [source]
- localbench: top-40 error is negligible. The existing clamp dossier: the two estimators are different and absolute values are not interchangeable. Neither has a head-to-head measurement. [source]
- A reader asked on 16 Apr 2026 whether the harness or dataset is downloadable; the visible reply is not shown in the cached page. [source]
- localbench requests top-40 prompt logprobs via `/v1/completions` with `echo: true`. [source]
- localbench uses the official release's Jinja2 template for every quant, not the GGUF's embedded template. [source]
- localbench's top-1 agreement is the fraction of positions where quant and reference pick the same most likely token. [source]
- localbench does not publish a measured error bound for the top-40 truncation. [source]
Children
- No children recorded.