Calibration-set contamination and held-out evaluation corpora
Parent: Mac local LLMs: Quantization evaluation · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
No published ablation of DWQ/AWQ/GPTQ calibration corpus on code or non-English quality on Apple Silicon.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- No published ablation of DWQ/AWQ/GPTQ calibration corpus on code or non-English quality on Apple Silicon. [source]
- mlx-community's default DWQ dataset was not confirmed from LEARNED_QUANTS.md (page only gives sample-count defaults). [source]
- bartowski's old calibration_datav5.txt was plain text, mostly English with a few languages, with no chat-formatted text. [source]
- bartowski's v6 imatrix dataset is two files, prose and conversations; conversations come from interstellarninja/hermes_reasoning_tool_use rendered through each model's chat template, with the rendered file uploaded beside the GGUF (e.g. Qwen3.8-27B-calibration-v6.txt). [source]
- The v6 prose-to-conversation ratio is 2:3, with 400-800 chunks tested; larger sets did not help and may drown important signal. [source]
- Above about 4 bits per weight (Q4_0/IQ4_XS and up), imatrix dataset choice showed no predictable difference, even with random-noise calibration. [source]
- At Q2_K, Qwen3.6-35B without any imatrix lost about 28 BFCL points (about 82% to about 54%); best vs worst corpus differed about 10%. [source]
- On dense models the corpus differences were hard to find; most low-bit imatrix benefit comes from MoE expert coverage. [source]
- Q2_K MMLU-Pro and GSM8K accuracy was nearly identical across all imatrix datasets including random words. [source]
- For MoE, adding chat-templated calibration data was the largest single change, implying some experts activate mainly on chat tokens. [source]
- A pure English tool corpus left 18 of 512 experts in Qwen3-Next-80B-A3B untouched even at 3126 chunks; multilingual and code content was what covered them. [source]
- bartowski's per-chunk attribution found some experts are fed only by specific prose (code, French, fiction, science), so the v6 prose was chosen for expert coverage; share of Latin-script text fell versus v5. [source]
- bartowski found imatrix context 2048 gave no gain over 512 and sometimes cost expert coverage. [source]
- bartowski found imatrix computed from Q8_0 had over 99.99% cosine similarity to one computed from bf16, so bf16 is likely unnecessary. [source]
- bartowski's v5-vs-v6 comparison on Qwen3.8-27B used held-to-different evals: wikitext, HuggingFaceH4/no_robots and Salesforce/xlam-function-calling-60k, 100 chunks each; on wikitext mean KLD differences were within about 4% (v6 slightly better at Q2_K -2.6%, worse at Q4_K_S +3.6%). [source]
- bartowski calls no_robots numbers only a proxy because bf16 PPL on it is 19.05. [source]
- bartowski validated v6 with KLD vs bf16, BFCL, a canary prompt set, MoE expert-coverage checks, and subsets of MMLU-Pro and GSM8K across Qwen3.6-27B, Qwen3.6-35B-A3B, Mistral-Small-4-119B, Qwen3-Next-80B-A3B and Qwen3.5-397B-A17B. [source]
- eaddario/imatrix-calibration on Hugging Face is a public collection of calibration datasets that bartowski used for corpus comparisons. [source]
- mlx-lm LEARNED_QUANTS.md lists DWQ, AWQ, dynamic quantization and GPTQ as all calibration-data driven; mlx_lm.dwq defaults to 1024 samples and mlx_lm.awq to 32 samples with n-grid 10. [source]
- Because bartowski's v6 mixes chat-templated data, a WikiText KLD eval is mostly held-out for it but a no_robots or xlam eval is nearer in distribution; recommended held-out eval is a corpus no quantizer's calibration set contains. [source]
Children
- No children recorded.