<!-- llms-explorer concept facts · https://llms-explorer.com/tree/generation-drift-metrics-beyond-prefill-kld/ · pack 2026-10-05 · ~2490 tokens -->

# Generation-drift metrics beyond prefill KLD

> Branch-and-follow: when engines disagree at a position, neither is pulled back to the teacher. Both continue and the pair of futures is compared. Requirement: the engines share one token dictionary (the Level1Techs work stays inside the Qwen3.x family).

Parent: [Mac local LLMs: Quantization evaluation](https://llms-explorer.com/tree/mac-local-llms-quantization-evaluation/) · 1 facets · 40 facts · page: https://llms-explorer.com/tree/generation-drift-metrics-beyond-prefill-kld/

## Facts

- Branch-and-follow: when engines disagree at a position, neither is pulled back to the teacher. Both continue and the pair of futures is compared. Requirement: the engines share one token dictionary (the Level1Techs work stays inside the Qwen3.x family). — source: `asserted`
- A teacher-forced top-1 flip is a counterfactual root: it says the candidate would have chosen a different greedy token at that position, not that the final answer differs. — source: `asserted`
- Flip outcomes are classified as recovers, changes meaning, or malformed/incorrect tool call. "Invalid branch futures" counts structurally invalid outputs among the flip roots the author chose to explore; it is not a general tool-call failure rate. — source: `asserted`
- Per-range summaries: the sample is split into output ranges (prose, exact CLI/SQL/code, multi-tool calls, recovery actions, architecture text) and a statistic such as p95 KLD is taken per range; the "worst-range p95 KLD" is the maximum over ranges. — source: `asserted`
- Capturing 100% of logits for tool-call spans is storage-heavy; the author moved from a 3% sample (one probe every 32 tokens, 250 probes per 8k window) to full capture only inside selected output ranges because gen5 NVMe space ran out. — source: `asserted`
- 2026-08-15 (Part 1): 3% sampling of logits, teacher-forced top-1 disagreement by 8k-token window on a ~100k-token agent prompt. — source: `asserted`
- 2026-08-17 (Part 2): full-logit capture during tool calls, forked decode paths, visualiser of parallel token streams. — source: `asserted`
- 2026-08-20 (Part 3.11): derivative models (abliterated or ablated fine-tunes) compared against stock Qwen3.8-27B with flip rate, worst-range p95 KLD and invalid-branch counts. — source: `asserted`
- 2026-08-24 (Part 3.14): 22 ranges and 9,060 positions across H200, B200 and SM120. — source: `asserted`
- Teacher-forced disagreement is not monotonic in context length: it clusters by content, spikes in some ranges and falls back to about 10% (the author says that is still "somewhere between bad and terrible"). — source: `asserted`
- A commenter read the dip at 50-70k context as the quants "converging back"; the author replied they are not converging and that a 3% sample leaves 97% of positions unobserved. — source: `asserted`
- KV-cache quantization: after the same flips in a tool call, BF16 KV was fine, int8 KV eventually recovered, int4 KV did not. — source: `asserted`
- Five-way weight comparison on one agent prompt: NVFP4 reached about 50% top-1 flips by 88k context and, with AWQ W4A16, failed to close the tool call and ran `show run` instead of `show arp`; first-party FP8 and a W8A16 INT8 completed the correct call. — source: `asserted`
- Failures concentrate in literal-copy tokens even when the baseline is confident: stock chose the final "2" of port 5432 with p=0.9991 and a derivative chose "ql" with p=0.9158, yielding 543ql; stock chose "enant" in "tenant" with p=0.99996 and the derivative chose "-" with p=0.99747; a psycopg `page_size=100` became `size=100`. — source: `asserted`
- A single flip can propagate: in one trace the wrong interface (GigabitEthernet0/1/4 instead of 0/0/1.201) caused the model to rerun a wrong command in two diverging calls. — source: `asserted`
- Different weights show different flip profiles even when both are "full BF16": see claims for four Qwen3.8-27B derivatives. — source: `asserted`
- Does top-1 at one position matter: Unsloth says top-1 is "an argmax on 1 prediction, so it's not really effective on gauging actual inference" and prefers Divergence-300 @32 (existing file). The Level1Techs author keeps top-1 but adds branching, arguing a single flip can be a tool-call failure. The two sides agree top-1 alone is insufficient and differ on the fix: fixed-length greedy rollout versus branch-and-follow from selected roots. — source: `asserted`
- Sample size: the author calls his own 3% random sample "not a correct overall methodology" while still using it for a high-level view; a commenter's "convergence" reading depends on that sample. — source: `asserted`
- No Mac harness runs branch-and-follow or Divergence-300-style rollouts across GGUF and MLX quants. — source: `asserted`
- How teacher-forced flip rate maps to end-task pass rate is unmeasured; the sources give only n=4 derivative models and one agent workstream. — source: `asserted`
- Whether sampled (temperature > 0) generation, which the sources do not measure, amplifies or damps drift versus greedy is unreported. — source: `asserted`
- The Level1Techs harness sampled one probe every 32 prompt tokens, giving 250 probes per 8k-token window, in Part 1. — [source](https://forum.level1techs.com/t/why-your-local-llm-feels-dumber-than-it-is/253917)
- A "top-1 flip" in the harness means a candidate would have chosen a different greedy next token than the reference at that position under a forced shared token history. — [source](https://forum.level1techs.com/t/why-your-local-llm-feels-dumber-than-it-is/253917)
- The author calls a teacher-forced top-1 flip a counterfactual root, not automatically a different complete answer, and branches from selected flip positions to see whether the continuation recovers, changes meaning or produces malformed tool calls. — [source](https://forum.level1techs.com/t/why-your-local-llm-feels-dumber-than-it-is/253917)
- In the branch visualiser, when two runtimes differ the display branches and follows both, never pulling the wrong one back to the teacher; the only requirement is that the token dictionary matches. — [source](https://forum.level1techs.com/t/why-your-local-llm-feels-dumber-than-it-is/253917)
- With BF16 weights and BF16 KV as baseline, int8 KV cache eventually recovered from a tool-call flip and int4 KV did not. — [source](https://forum.level1techs.com/t/why-your-local-llm-feels-dumber-than-it-is/253917)
- In a five-way Qwen3.6-27B comparison, NVFP4 (NVIDIA checkpoint) reached about 50% top-1 flips by 88k context, the worst of the five. — [source](https://forum.level1techs.com/t/why-your-local-llm-feels-dumber-than-it-is/253917)
- NVFP4 and AWQ W4A16 both failed to close their tool calls and ran `show run` where the correct command was `show arp`; FP8 and INT8 W8A16 completed the correct calls. — [source](https://forum.level1techs.com/t/why-your-local-llm-feels-dumber-than-it-is/253917)
- The author says teacher-forced disagreement "falls back to 10%" in late context, that this is still bad, and that disagreement is highly prompt dependent. — [source](https://forum.level1techs.com/t/why-your-local-llm-feels-dumber-than-it-is/253917)
- Part 3.14 samples 22 output ranges and 9,060 positions at depths from 6,335 to 122,863 tokens, covering prose, parallel tool calls, a mutating Podman call and a SQLite-to-PostgreSQL migration script. — [source](https://forum.level1techs.com/t/why-your-local-llm-feels-dumber-than-it-is/253917)
- Derivative comparison (SP06 top-1 flips / worst-range p95 KLD / invalid branch futures): Heretic-ARA 1.337% / 0.03661 / 0 of 58; Huihui 1.406% / 0.04477 / 0 of 61; Blackfrost 4.978% / 0.31235 / 1 of 216; AEON 5.831% / 0.68507 / 36 of 253. — [source](https://forum.level1techs.com/t/why-your-local-llm-feels-dumber-than-it-is/253917)
- SP04 covers 1,535 assistant-output tokens in six ranges and SP06 covers 4,339 tokens in seven ranges. — [source](https://forum.level1techs.com/t/why-your-local-llm-feels-dumber-than-it-is/253917)
- For Heretic-ARA, 57 of 58 SP06 flips occurred where stock Qwen was already uncertain, and it overturned no strongly preferred stock token. — [source](https://forum.level1techs.com/t/why-your-local-llm-feels-dumber-than-it-is/253917)
- AEON overturned 24 strongly preferred stock decisions on SP06 and accounted for all 8 SP04 invalid futures. — [source](https://forum.level1techs.com/t/why-your-local-llm-feels-dumber-than-it-is/253917)
- At SP06 token position 42,950 stock chose "2" (p=0.9991) in PostgreSQL port 5432 and AEON chose "ql" (p=0.9158), and the continuation produced 543ql and failed to close the tool envelope. — [source](https://forum.level1techs.com/t/why-your-local-llm-feels-dumber-than-it-is/253917)
- The author states refusal benchmarks and model-card KLD numbers do not characterize collateral changes to tool use, exact literals, code or long-context agentic work. — [source](https://forum.level1techs.com/t/why-your-local-llm-feels-dumber-than-it-is/253917)
- The author moved from 3% sampling to 100% capture inside selected ranges because storing full log-probs for hundreds of thousands of tokens and branching paths exhausted fast NVMe storage. — [source](https://forum.level1techs.com/t/why-your-local-llm-feels-dumber-than-it-is/253917)
- Unsloth states that its Divergence-300 @32 sharp drop between UD-Q2_K_XL and UD-IQ2_S means tool calling and non-thinking modes break down. — [source](https://unsloth.ai/docs/basics/dynamic-3.0-ggufs)
- For derivative models of one base, ranking by SP06 top-1 flips, by worst-range p95 KLD and by invalid-branch count gives the same order (Heretic-ARA, Huihui, Blackfrost, AEON), but n=4 and the derivatives are weight edits, not quants. — source: `asserted`
- A Mac drift check can reuse the method without CUDA: capture reference token streams from one engine, force them through each quant with llama-server or mlx-lm, then branch only at flip roots to keep storage small. — source: `asserted`
