llama.cpp .kld reference file binary layout and external reader
Parent: Mac local LLMs: Quantization evaluation · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
The encoder clamps the low end of the stored range: min_logit = max(min_logit, max_logit - 16), so the 65535 steps cover at most 16 nats and the step is at most 16/65535, about 2.4e-4 nats.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- The encoder clamps the low end of the stored range: min_logit = max(min_logit, max_logit - 16), so the 65535 steps cover at most 16 nats and the step is at most 16/65535, about 2.4e-4 nats. [source]
- Any logit at or below the clamped minimum is stored as 0 and decodes to min_log_prob, so every token more than 16 nats under the top token reads back as the same floor value; tail mass below that is not represented per token. [source]
- If max_logit equals min_logit (scale 0), the encoder writes zeros for the whole record. [source]
- min_log_prob is min_logit - max_logit - log_sum_exp, so a decoder needs only the two floats and the uint16 values, with no separate normalizer. [source]
- The saved token ids are the original text tokens for every chunk, written as n_chunk*n_ctx llama_token values; BOS is substituted at the first position of each chunk only inside the decode batch and then restored, so the file holds no BOS. [source]
- The reference run requires at least 2*n_ctx tokens of text and sets n_chunk to min(requested chunks, tokens / n_ctx). [source]
- The reader checks the magic and fails on a bad one; on n_ctx larger than the current context and on n_vocab differing from the model's vocabulary it only logs an error and keeps going, so a mismatched vocabulary does not stop the run and can produce wrong KLD. [source]
- The reader asserts the vocabulary does not add EOS and uses the same add_bos setting as the writer through llama_vocab_get_add_bos. [source]
- A pure-Python reader therefore needs: 8-byte magic, <I n_ctx, <i n_vocab, <i n_chunk, n_ctx*n_chunk int32 token ids, then per chunk (n_ctx - 1 - n_ctx/2) records of nv uint16, each decoded as scale*u + min_log_prob. [source]
- PR 5076 introduced the mode (existing dossiers); the 16-nat clamp is in the current master source read here, and its introduction date was not checked. [source]
- Tail KLD terms that depend on tokens more than 16 nats below the top are quantized to one value, so reference-side precision for extreme tails is limited by the clamp, separate from the 16-bit step. [source]
- A reader that decodes with the wrong n_vocab misaligns every later record because the record stride is nv uint16. [source]
- None found; the file format is documented only by the source. [source]
- Whether the 16-nat clamp changes KLD for models with very peaked distributions enough to matter at the 4e-6 floor quoted in cross-engine-reference-logits-precision-and-toke.md. [source]
- The .kld encoder clamps the stored logit range to 16 nats below the maximum logit. [source]
- Logits at or below the clamped minimum are stored as 0 and decode to min_log_prob. [source]
- A record with scale 0 is written as all zeros. [source]
- The file stores the original chunk tokens and restores them after temporarily setting BOS at each chunk start. [source]
- The reference pass needs at least 2*n_ctx tokens and n_chunk is capped by tokens divided by n_ctx. [source]
- The KLD reader logs but does not abort on an n_vocab mismatch or an n_ctx larger than the current context. [source]
- The remaining layout is already covered in confidence-conditioned-top-1-flip-rate-for-liter.md. [source]
Children
- No children recorded.