Expert overlap between draft tokens in MoE speculative verification
Parent: Mac local LLMs: Speculative decoding and MTP · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
A MoE layer's arithmetic intensity is T(k+S)/(N+S) for T tokens, k routed experts per token, S shared experts and N routed experts. Low k/N keeps the layer bandwidth-bound at larger batch sizes.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- A MoE layer's arithmetic intensity is T(k+S)/(N+S) for T tokens, k routed experts per token, S shared experts and N routed experts. Low k/N keeps the layer bandwidth-bound at larger batch sizes. [source]
- Verification cost ratio T(K+1)/T(1) decides speedup because draft cost is small. The cost splits into routed-expert weight loading (fraction f, scaling with unique experts) and everything else (fixed at small batch). [source]
- Routing statistics are compared against a uniform baseline (overlap k/N), an independence baseline (empirical popularity, no temporal correlation) and the empirical measurement; the gap between the last two isolates temporal correlation. [source]
- Neighboring draft tokens share experts, so a drafter can predict which experts verification will need (SP-MoE prefetches them from the draft pass). [source]
- 2025: SP-MoE (arXiv 2510.10302, 11 Oct 2025) uses draft-token expert overlap to prefetch experts on offloaded GPU setups. [source]
- 2026: Cohere's write-up measures unique experts per verification window and the overlap decay with token distance (the data below). [source]
- Weight-traffic savings from overlap exist only if the kernel reads each unique expert once. The MLX `gather_qmv` path in the existing MLX dossier reads per-token blocks, so the GPU measurements below bound what a deduplicating kernel could save, not what current MLX gains. [source]
- At BS=1 a second effect, amortized fixed overhead, adds speedup that expert overlap does not explain. [source]
- A long verify block approaches full expert activation, which removes the overlap benefit. [source]
- llama.cpp DFlash on gpt-oss-20b ran slower than baseline (0.61x to 1.27x) because the verify block activates more experts than one decode step. [source]
- Overlap size: Cohere measures 38.1% adjacent-token overlap for k=8 of N=128 experts with shared experts on MT-Bench. The flash-moe paper measured 8 to 34% per layer at K=4 on a different model and 200 tokens. Both are held side by side; the models, k and sampling differ. [source]
- Verify-window expert overlap for the Mac-relevant MoE targets (Qwen3.5/3.6-35B-A3B, Flash-Next): no source measures it. [source]
- Whether any MLX or llama.cpp Metal kernel deduplicates experts across a verify window. [source]
- Whether overlap measured on natural-text decode holds for tokens drafted by a block-diffusion drafter, which are not sampled autoregressively. [source]
- Cohere measured expert routing on a MoE with 128 routed experts, k=8 and shared experts, using MT-Bench and a modified vLLM `enable_return_routed_experts` API. [source]
- Adjacent-token expert overlap in that model is 0.381 at step 1, 0.329 at step 2, 0.301 at step 3 and 0.299 at step 4, against 0.118 for the independence baseline and 0.0625 for the uniform baseline. [source]
- Mid-network layers show about 0.50 step-1 overlap, early layers about 0.32 and final layers about 0.25. [source]
- Expected active experts for a batch of 4 tokens is 25.4 measured against 29.1 uniform, and the measured-to-uniform ratio falls from 1.00 at one token to 0.71 at 32 tokens and recovers to 0.85 at 128. [source]
- Verifying four tokens (K=3) activates 20.36 unique experts measured, 25.4 under the independence baseline and 29.1 under uniform, which Cohere reads as about 2.5 times top-k instead of 3.2 to 3.6 times. [source]
- Cohere reports temporal correlation cuts unique experts by 20 to 31% against the independence baseline. [source]
- Step-1 overlap is 0.377 to 0.385 across 13 Spec-Bench categories and 0.378 to 0.386 across seven languages, against 0.381 on MT-Bench. [source]
- At BS=1 the measured target cost ratio T(1)/T(4) is 0.80 and at BS=2 it is 0.65, and expert routing does not explain the gap because concentration is slightly better at BS=2. [source]
- Cohere attributes the BS=1 gap to fixed overhead (attention, norms, communication, kernel launches) amortized over four verified tokens. [source]
- An Amdahl decomposition with routed-expert fraction f=0.30 predicts T(4)/T(1) of 1.46x against 1.25x measured, and implied speedup 1.87x against 2.18x implied by measurement and 1.95x observed at BS=1. [source]
- The measured verification cost scaling is 1.25 to 1.42x of one decode step at BS=1, not the (K+1)x worst case. [source]
- Cohere's dense baseline (Command A, 111B) shows speculative speedup decaying monotonically with batch size, while its MoE shows a non-monotonic curve with a sweet spot at moderate batch size. [source]
- The sweet spot moves to larger batch sizes for sparser models, and shared experts lower verification cost at low batch size while raising effective k/N and bringing the compute-bound regime earlier. [source]
- SP-MoE reports that the number of activated experts does not grow linearly with the number of neighboring draft tokens and that many token pairs share experts, measured on Mixtral-style and DeepSeek-style models over two datasets. [source]
- SP-MoE predicts which experts verification needs by combining the draft model's attention outputs with the target model's gating network, which gave lower activation entropy than history-based prefetching. [source]
- SP-MoE reports 1.07x to 3.5x lower time per output token than offloading baselines and a U-shaped effect of its prefetch cutoff layer, with the best cutoff near layer 20 on its main models. [source]
- Loading one Mixtral 8x7B layer over PCIe 4.0 takes about 80 ms against about 3 ms to compute it on an RTX 4090, so expert transfer, not compute, limits offloaded MoE verification. [source]
- llama.cpp's DFlash PR measured gpt-oss-20b DFlash at 0.61x to 1.27x and attributes the weak MoE result to more experts activating during parallel verification. [source]
Corrections and disagreements
- CONTRADICTS: moe-gatherqmm-small-m-verify-cost.md lines 30 and 37, which say no source reports overlap between draft tokens and that none was found for a measured verify cost against window size. Cohere's blog reports overlap by step, unique experts per window, and a measured T(1)/T(4) of 0.80 at BS=1, though on vLLM and a GPU, not MLX. [source]
Children
- No children recorded.