MoE expert coverage in imatrix calibration
Parent: Mac local LLMs: Quantization evaluation · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Collection: the callback intercepts GGML_OP_MUL_MAT_ID, copies the routing ids to host, and for each token adds the squared activations only to the experts in that token's top-k; a per-expert counts vector records tokens seen. Non-finite values abort the run, checked only for touched experts.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Collection: the callback intercepts GGML_OP_MUL_MAT_ID, copies the routing ids to host, and for each token adds the squared activations only to the experts in that token's top-k; a per-expert counts vector records tokens seen. Non-finite values abort the run, checked only for touched experts. [source]
- Saving: the collector scans the counts; a tensor with some zero counts logs `entry '<name>' has partial data (xx.xx%)`; a tensor with all zero counts logs `has no data - skipping`; the summary line `storing only N out of M entries` appears if any were dropped. [source]
- Legacy .dat format: zero-count experts are written with value 1 and count 1, a neutral entry; GGUF format stores `*.in_sum2` and `*.counts` per tensor so files can be merged with --in-file. [source]
- Quantize time (GGUF imatrix): per-expert sums are divided by per-expert counts; an expert with count 0 receives weight 1 for every channel, which is the flat (non-imatrix) weighting for that expert. The imatrix slice for each expert is applied separately and chunks never cross an expert boundary. [source]
- Required-imatrix types: IQ1_S, IQ1_M, IQ2_XXS, IQ2_XS, IQ2_S, IQ3_XXS and Q2_K inside a Q2_K_S file refuse to quantize without an imatrix entry for a tensor and abort with 'Missing importance matrix for tensor ... in a very low-bit quantization'. A tensor skipped for 'no data' therefore aborts a low-bit run rather than falling back silently. [source]
- 2024-01: ikawrakow's PR 4861 introduces imatrix (CPU only at first; ~10 minutes for 50k tokens on a 7B on a 32-core CPU). [source]
- 2024-09: compilade's PR 9400 moves imatrix files to GGUF with in_sum2 and counts, adds merge support, and adds the partial-data warning 'to help guess dataset coverage'; the legacy writer is made to emit neutral values for missing data to match read-time behavior of the new format. [source]
- 2026: Nota AI describes expert-coverage-balanced synthetic calibration (PASCAL-MoE) for NVIDIA's Nemotron hackathon; vendor blog only. [source]
- An all-zero tensor is dropped from the imatrix file; at IQ2/IQ1 file types the quantizer then aborts, but at Q4_K/IQ4 types it quantizes that tensor with no imatrix. [source]
- Partial data is accepted silently in the file and only logged as a warning; partial coverage cannot be detected after the fact from the file alone except by inspecting per-expert counts in the GGUF imatrix. [source]
- Old Mixtral imatrix files made for split per-expert tensors do not match the merged 3D layout; llama-quantize throws 'imatrix size ... is different from tensor size' and stops (since the error otherwise hides most of the model being quantized without an imatrix). [source]
- process-output defaults to false; the output.weight is normally quantized without importance data. [source]
- Unsloth (chat/tool-call/long-context calibration, 300K-1.5M tokens) vs Wikipedia-style imatrix sets: Unsloth argues text-only calibration is ineffective for instruct models; existing dossier covers this. New here: neither side reports per-expert coverage numbers. [source]
- Whether a flat weighting for an unrouted expert is harmless or harmful is unmeasured: a rarely routed expert is by definition rarely used, yet 'rarely activated' experts are exactly those the MoE-quant literature calls sensitive. [source]
- What fraction of experts per tensor is left at count zero by a standard 300K-token chat imatrix on fine-grained many-expert models such as Qwen3.6-35B-A3B? No published per-expert coverage histogram was found. [source]
- Whether KLD loss concentrates in under-covered experts (needs a per-expert KLD attribution, not found). [source]
- llama-imatrix collects data for MoE through GGML_OP_MUL_MAT_ID, copies the routing ids to host, and keeps one sums vector and one token-count entry per expert for each merged expert tensor. [source]
- llama-imatrix adds a token's squared activations only to the experts in that token's top-k and checks non-finite values only for experts that were touched. [source]
- When saving, llama-imatrix logs `entry '<name>' has partial data (xx.xx%)` for any tensor with some zero-count experts and `has no data - skipping` for tensors with all-zero counts, then `storing only N out of M entries`. [source]
- In the legacy imatrix.dat writer, zero-count experts are stored with value 1 and count 1 (a neutral entry), matching what the new format does at read time. [source]
- The GGUF imatrix format stores `<tensor>.in_sum2` (per-channel sums of squared activations, F32) and `<tensor>.counts` (activation counts) so that imatrix files can be merged with --in-file. [source]
- llama-quantize loading a GGUF imatrix divides each expert's sums by that expert's count and sets every channel weight to 1 for an expert whose count is 0. [source]
- llama-quantize applies the imatrix slice per expert (global row index divided by rows per expert), and chunks never cross an expert boundary. [source]
- llama-quantize aborts with 'Missing importance matrix for tensor X in a very low-bit quantization' when a tensor targeted at IQ1_S, IQ1_M, IQ2_XXS, IQ2_XS, IQ2_S or IQ3_XXS (or Q2_K inside Q2_K_S) has no imatrix entry. [source]
- IQ4_NL, IQ4_XS and the K-quants other than Q2_K_S do not require an imatrix, so zero-data or missing entries only lose quality for those types and do not stop the run. [source]
- llama-quantize throws 'imatrix size ... is different from tensor size' for a size mismatch, except for token_embd; the code comment says this happened when quantizing an old Mixtral with split tensors against a newer incompatible imatrix. [source]
- PR 9400's commit message says imatrix now 'warns when writing partial data, to help guess dataset coverage'; the legacy format is made to store neutral values for missing data to give the same quality. [source]
- The llama-imatrix README documents --process-output (default false, because the importance matrix is usually not used for output.weight), --chunk/--chunks, --no-ppl, --show-statistics, --in-file merging and --parse-special. [source]
- ikawrakow's original imatrix PR argues per-row diagonal weights w_j proportional to squared activations are enough, ignoring off-diagonal terms, and notes that instruct models may need instruction-formatted data split by entry instead of fixed chunks. [source]
- Nota AI states that experts rarely activated by calibration data receive sparse activation statistics and unreliable quantization parameters, and its PASCAL-MoE pipeline generates synthetic calibration samples to balance expert coverage (a Nemotron-3-Super-120B agent finds the tokens that trigger under-used experts). [source]
- For a 256-expert model with top-8 routing, a calibration set of T tokens gives an average of 8T/256 = T/32 tokens per expert, so a 300K-token set averages about 9.4K tokens per expert with large variance by expert popularity. [source]
- Because an unrouted expert is quantized with flat weights, its quality equals non-imatrix quantization of that expert, which at 4 bits is usually close (Q4_K/IQ4 do not need an imatrix) and at 2 bits can be badly worse. [source]
Children
- No children recorded.