<!-- llms-explorer concept facts · https://llms-explorer.com/tree/expert-overlap-between-draft-tokens-in-moe-specu/ · pack 2026-10-05 · ~2041 tokens -->

# Expert overlap between draft tokens in MoE speculative verification

> A MoE layer's arithmetic intensity is T(k+S)/(N+S) for T tokens, k routed experts per token, S shared experts and N routed experts. Low k/N keeps the layer bandwidth-bound at larger batch sizes.

Parent: [Mac local LLMs: Speculative decoding and MTP](https://llms-explorer.com/tree/mac-local-llms-speculative-decoding-and-mtp/) · 2 facets · 33 facts · page: https://llms-explorer.com/tree/expert-overlap-between-draft-tokens-in-moe-specu/

## Facts

- A MoE layer's arithmetic intensity is T(k+S)/(N+S) for T tokens, k routed experts per token, S shared experts and N routed experts. Low k/N keeps the layer bandwidth-bound at larger batch sizes. — source: `asserted`
- Verification cost ratio T(K+1)/T(1) decides speedup because draft cost is small. The cost splits into routed-expert weight loading (fraction f, scaling with unique experts) and everything else (fixed at small batch). — source: `asserted`
- Routing statistics are compared against a uniform baseline (overlap k/N), an independence baseline (empirical popularity, no temporal correlation) and the empirical measurement; the gap between the last two isolates temporal correlation. — source: `asserted`
- Neighboring draft tokens share experts, so a drafter can predict which experts verification will need (SP-MoE prefetches them from the draft pass). — source: `asserted`
- 2025: SP-MoE (arXiv 2510.10302, 11 Oct 2025) uses draft-token expert overlap to prefetch experts on offloaded GPU setups. — source: `asserted`
- 2026: Cohere's write-up measures unique experts per verification window and the overlap decay with token distance (the data below). — source: `asserted`
- Weight-traffic savings from overlap exist only if the kernel reads each unique expert once. The MLX `gather_qmv` path in the existing MLX dossier reads per-token blocks, so the GPU measurements below bound what a deduplicating kernel could save, not what current MLX gains. — source: `asserted`
- At BS=1 a second effect, amortized fixed overhead, adds speedup that expert overlap does not explain. — source: `asserted`
- A long verify block approaches full expert activation, which removes the overlap benefit. — source: `asserted`
- llama.cpp DFlash on gpt-oss-20b ran slower than baseline (0.61x to 1.27x) because the verify block activates more experts than one decode step. — source: `asserted`
- Overlap size: Cohere measures 38.1% adjacent-token overlap for k=8 of N=128 experts with shared experts on MT-Bench. The flash-moe paper measured 8 to 34% per layer at K=4 on a different model and 200 tokens. Both are held side by side; the models, k and sampling differ. — source: `asserted`
- Verify-window expert overlap for the Mac-relevant MoE targets (Qwen3.5/3.6-35B-A3B, Flash-Next): no source measures it. — source: `asserted`
- Whether any MLX or llama.cpp Metal kernel deduplicates experts across a verify window. — source: `asserted`
- Whether overlap measured on natural-text decode holds for tokens drafted by a block-diffusion drafter, which are not sampled autoregressively. — source: `asserted`
- Cohere measured expert routing on a MoE with 128 routed experts, k=8 and shared experts, using MT-Bench and a modified vLLM `enable_return_routed_experts` API. — [source](https://cohere.com/blog/mixture-of-experts-models-get-more-from-speculative-decoding)
- Adjacent-token expert overlap in that model is 0.381 at step 1, 0.329 at step 2, 0.301 at step 3 and 0.299 at step 4, against 0.118 for the independence baseline and 0.0625 for the uniform baseline. — [source](https://cohere.com/blog/mixture-of-experts-models-get-more-from-speculative-decoding)
- Mid-network layers show about 0.50 step-1 overlap, early layers about 0.32 and final layers about 0.25. — [source](https://cohere.com/blog/mixture-of-experts-models-get-more-from-speculative-decoding)
- Expected active experts for a batch of 4 tokens is 25.4 measured against 29.1 uniform, and the measured-to-uniform ratio falls from 1.00 at one token to 0.71 at 32 tokens and recovers to 0.85 at 128. — [source](https://cohere.com/blog/mixture-of-experts-models-get-more-from-speculative-decoding)
- Verifying four tokens (K=3) activates 20.36 unique experts measured, 25.4 under the independence baseline and 29.1 under uniform, which Cohere reads as about 2.5 times top-k instead of 3.2 to 3.6 times. — [source](https://cohere.com/blog/mixture-of-experts-models-get-more-from-speculative-decoding)
- Cohere reports temporal correlation cuts unique experts by 20 to 31% against the independence baseline. — [source](https://cohere.com/blog/mixture-of-experts-models-get-more-from-speculative-decoding)
- Step-1 overlap is 0.377 to 0.385 across 13 Spec-Bench categories and 0.378 to 0.386 across seven languages, against 0.381 on MT-Bench. — [source](https://cohere.com/blog/mixture-of-experts-models-get-more-from-speculative-decoding)
- At BS=1 the measured target cost ratio T(1)/T(4) is 0.80 and at BS=2 it is 0.65, and expert routing does not explain the gap because concentration is slightly better at BS=2. — [source](https://cohere.com/blog/mixture-of-experts-models-get-more-from-speculative-decoding)
- Cohere attributes the BS=1 gap to fixed overhead (attention, norms, communication, kernel launches) amortized over four verified tokens. — [source](https://cohere.com/blog/mixture-of-experts-models-get-more-from-speculative-decoding)
- An Amdahl decomposition with routed-expert fraction f=0.30 predicts T(4)/T(1) of 1.46x against 1.25x measured, and implied speedup 1.87x against 2.18x implied by measurement and 1.95x observed at BS=1. — [source](https://cohere.com/blog/mixture-of-experts-models-get-more-from-speculative-decoding)
- The measured verification cost scaling is 1.25 to 1.42x of one decode step at BS=1, not the (K+1)x worst case. — [source](https://cohere.com/blog/mixture-of-experts-models-get-more-from-speculative-decoding)
- Cohere's dense baseline (Command A, 111B) shows speculative speedup decaying monotonically with batch size, while its MoE shows a non-monotonic curve with a sweet spot at moderate batch size. — [source](https://cohere.com/blog/mixture-of-experts-models-get-more-from-speculative-decoding)
- The sweet spot moves to larger batch sizes for sparser models, and shared experts lower verification cost at low batch size while raising effective k/N and bringing the compute-bound regime earlier. — [source](https://cohere.com/blog/mixture-of-experts-models-get-more-from-speculative-decoding)
- SP-MoE reports that the number of activated experts does not grow linearly with the number of neighboring draft tokens and that many token pairs share experts, measured on Mixtral-style and DeepSeek-style models over two datasets. — [source](https://arxiv.org/html/2510.10302v1)
- SP-MoE predicts which experts verification needs by combining the draft model's attention outputs with the target model's gating network, which gave lower activation entropy than history-based prefetching. — [source](https://arxiv.org/html/2510.10302v1)
- SP-MoE reports 1.07x to 3.5x lower time per output token than offloading baselines and a U-shaped effect of its prefetch cutoff layer, with the best cutoff near layer 20 on its main models. — [source](https://arxiv.org/html/2510.10302v1)
- Loading one Mixtral 8x7B layer over PCIe 4.0 takes about 80 ms against about 3 ms to compute it on an RTX 4090, so expert transfer, not compute, limits offloaded MoE verification. — [source](https://arxiv.org/html/2510.10302v1)
- llama.cpp's DFlash PR measured gpt-oss-20b DFlash at 0.61x to 1.27x and attributes the weak MoE result to more experts activating during parallel verification. — [source](https://github.com/ggml-org/llama.cpp/pull/22105)

## Corrections and disagreements

- CONTRADICTS: moe-gatherqmm-small-m-verify-cost.md lines 30 and 37, which say no source reports overlap between draft tokens and that none was found for a measured verify cost against window size. Cohere's blog reports overlap by step, unique experts per window, and a measured T(1)/T(4) of 0.80 at BS=1, though on vLLM and a GPU, not MLX. — source: `asserted`
