<!-- llms-explorer concept facts · https://llms-explorer.com/tree/calibration-set-contamination-and-held-out-evalu/ · pack 2026-10-05 · ~1197 tokens -->

# Calibration-set contamination and held-out evaluation corpora

> No published ablation of DWQ/AWQ/GPTQ calibration corpus on code or non-English quality on Apple Silicon.

Parent: [Mac local LLMs: Quantization evaluation](https://llms-explorer.com/tree/mac-local-llms-quantization-evaluation/) · 1 facets · 20 facts · page: https://llms-explorer.com/tree/calibration-set-contamination-and-held-out-evalu/

## Facts

- No published ablation of DWQ/AWQ/GPTQ calibration corpus on code or non-English quality on Apple Silicon. — source: `asserted`
- mlx-community's default DWQ dataset was not confirmed from LEARNED_QUANTS.md (page only gives sample-count defaults). — source: `asserted`
- bartowski's old calibration_datav5.txt was plain text, mostly English with a few languages, with no chat-formatted text. — [source](https://huggingface.co/blog/bartowski/imatrix-dataset)
- bartowski's v6 imatrix dataset is two files, prose and conversations; conversations come from interstellarninja/hermes_reasoning_tool_use rendered through each model's chat template, with the rendered file uploaded beside the GGUF (e.g. Qwen3.8-27B-calibration-v6.txt). — [source](https://huggingface.co/blog/bartowski/imatrix-dataset)
- The v6 prose-to-conversation ratio is 2:3, with 400-800 chunks tested; larger sets did not help and may drown important signal. — [source](https://huggingface.co/blog/bartowski/imatrix-dataset)
- Above about 4 bits per weight (Q4_0/IQ4_XS and up), imatrix dataset choice showed no predictable difference, even with random-noise calibration. — [source](https://huggingface.co/blog/bartowski/imatrix-dataset)
- At Q2_K, Qwen3.6-35B without any imatrix lost about 28 BFCL points (about 82% to about 54%); best vs worst corpus differed about 10%. — [source](https://huggingface.co/blog/bartowski/imatrix-dataset)
- On dense models the corpus differences were hard to find; most low-bit imatrix benefit comes from MoE expert coverage. — [source](https://huggingface.co/blog/bartowski/imatrix-dataset)
- Q2_K MMLU-Pro and GSM8K accuracy was nearly identical across all imatrix datasets including random words. — [source](https://huggingface.co/blog/bartowski/imatrix-dataset)
- For MoE, adding chat-templated calibration data was the largest single change, implying some experts activate mainly on chat tokens. — [source](https://huggingface.co/blog/bartowski/imatrix-dataset)
- A pure English tool corpus left 18 of 512 experts in Qwen3-Next-80B-A3B untouched even at 3126 chunks; multilingual and code content was what covered them. — [source](https://huggingface.co/blog/bartowski/imatrix-dataset)
- bartowski's per-chunk attribution found some experts are fed only by specific prose (code, French, fiction, science), so the v6 prose was chosen for expert coverage; share of Latin-script text fell versus v5. — [source](https://huggingface.co/blog/bartowski/imatrix-dataset)
- bartowski found imatrix context 2048 gave no gain over 512 and sometimes cost expert coverage. — [source](https://huggingface.co/blog/bartowski/imatrix-dataset)
- bartowski found imatrix computed from Q8_0 had over 99.99% cosine similarity to one computed from bf16, so bf16 is likely unnecessary. — [source](https://huggingface.co/blog/bartowski/imatrix-dataset)
- bartowski's v5-vs-v6 comparison on Qwen3.8-27B used held-to-different evals: wikitext, HuggingFaceH4/no_robots and Salesforce/xlam-function-calling-60k, 100 chunks each; on wikitext mean KLD differences were within about 4% (v6 slightly better at Q2_K -2.6%, worse at Q4_K_S +3.6%). — [source](https://huggingface.co/blog/bartowski/imatrix-dataset)
- bartowski calls no_robots numbers only a proxy because bf16 PPL on it is 19.05. — [source](https://huggingface.co/blog/bartowski/imatrix-dataset)
- bartowski validated v6 with KLD vs bf16, BFCL, a canary prompt set, MoE expert-coverage checks, and subsets of MMLU-Pro and GSM8K across Qwen3.6-27B, Qwen3.6-35B-A3B, Mistral-Small-4-119B, Qwen3-Next-80B-A3B and Qwen3.5-397B-A17B. — [source](https://huggingface.co/blog/bartowski/imatrix-dataset)
- eaddario/imatrix-calibration on Hugging Face is a public collection of calibration datasets that bartowski used for corpus comparisons. — [source](https://huggingface.co/blog/bartowski/imatrix-dataset)
- mlx-lm LEARNED_QUANTS.md lists DWQ, AWQ, dynamic quantization and GPTQ as all calibration-data driven; mlx_lm.dwq defaults to 1024 samples and mlx_lm.awq to 32 samples with n-grid 10. — [source](https://raw.githubusercontent.com/ml-explore/mlx-lm/main/mlx_lm/LEARNED_QUANTS.md)
- Because bartowski's v6 mixes chat-templated data, a WikiText KLD eval is mostly held-out for it but a no_robots or xlam eval is nearer in distribution; recommended held-out eval is a corpus no quantizer's calibration set contains. — source: `asserted`
