<!-- llms-explorer concept facts · https://llms-explorer.com/tree/moesd-analysis-of-moe-speculative-decoding-batch/ · pack 2026-10-05 · ~1839 tokens -->

# MoESD analysis of MoE speculative decoding batch-size regimes

> Speedup is `S/R` over a denominator of three terms: drafter-to-target forward ratio `gamma * T_D(B,1) / T_T(B,1)`, multi-token to single-token target ratio `T_T(B,gamma) / T_T(B,1)`, and a rejection-sampling term; the second term is the largest and drives the regime changes.

Parent: [Mac local LLMs: Speculative decoding and MTP](https://llms-explorer.com/tree/mac-local-llms-speculative-decoding-and-mtp/) · 1 facets · 31 facts · page: https://llms-explorer.com/tree/moesd-analysis-of-moe-speculative-decoding-batch/

## Facts

- Speedup is `S/R` over a denominator of three terms: drafter-to-target forward ratio `gamma * T_D(B,1) / T_T(B,1)`, multi-token to single-token target ratio `T_T(B,gamma) / T_T(B,1)`, and a rejection-sampling term; the second term is the largest and drives the regime changes. — [source](https://arxiv.org/html/2505.19645v3)
- Two separate mechanisms raise `T_T(B,gamma) / T_T(B,1)`: compute-boundness (the ratio goes from about 1 when memory-bound to `gamma` when compute-bound, hurting dense and MoE at large batch) and extra expert weight loads (hurting MoE at small batch). — [source](https://arxiv.org/html/2505.19645v3)
- The paper defines "target efficiency" as `T_T(B,1) / T_T(B,gamma)`, which isolates target-side system effects from the acceptance rate. — [source](https://arxiv.org/html/2505.19645v3)
- With i.i.d. uniform routing, the expected number of activated experts for `t` tokens is `N(t) = E * (1 - ((E - K) / E)^t)` for `E` experts and `K` experts per token. — [source](https://arxiv.org/html/2505.19645v3)
- Defining sparsity `rho = K / E` and "almost full activation" as `N(t) >= tau * E` (tau usually 0.95) gives a token threshold `T_thres = ceil(log_(1-rho)(1 - tau))`; above it the `B * gamma` verify tokens add only marginal weight traffic. — [source](https://arxiv.org/html/2505.19645v3)
- Average tokens per expert is `rho * t / (1 - (1 - rho)^t)`, which falls as the MoE gets sparser, so sparser MoEs stay memory-bound to larger batch and keep a wider SD window. — [source](https://arxiv.org/html/2505.19645v3)
- Worked values of the threshold at tau = 0.95 under uniform routing: rho = 1/8 gives 23 tokens, rho = 1/16 gives 47, rho = 1/32 gives 95. — source: `asserted`
- The arXiv identifier 2505.19645 dates the first version to May 2025; the cached v3 is dated 6 Oct 2025. — [source](https://arxiv.org/html/2505.19645v3)
- The paper builds on MagicDec's analysis of the KV-dominated regime and says the two can be combined for a fuller view across batch sizes. — [source](https://arxiv.org/html/2505.19645v3)
- Very sparse variants with K = 1 or 2 on Qwen2-57B-A14B showed speedup that only decreased with batch, which the authors attribute to attention dominating a model whose FFN share was artificially cut by editing `num_experts_per_token`, not to a flaw in the theory. — [source](https://arxiv.org/html/2505.19645v3)
- The sparsity sweep changes K without retraining, so the authors rescale raw speedup by `sigma_(K=8) / sigma_K` to remove the acceptance change. — [source](https://arxiv.org/html/2505.19645v3)
- The paper assumes KV-cache traffic is small next to weight traffic; at long context the regime shifts and MagicDec's analysis applies instead. — [source](https://arxiv.org/html/2505.19645v3)
- Temperature 1.0 and open-ended chat cut the gain (MT-Bench peak 1.19x to 1.29x on Qwen2 at 2xGPU-A across gamma of 2 to 4) against code at temperature 0 (1.63x to 2.18x). — [source](https://arxiv.org/html/2505.19645v3)
- Mac single-stream decode is batch 1, below every threshold in the paper, so the extra-expert-load regime applies there; the paper's win regime needs tens of concurrent requests. — source: `asserted`
- Routing model: MoESD assumes uniform i.i.d. routing and reports that measured activated-expert counts on DeepSeek-V2-Lite-Chat and Qwen1.5-MoE-Chat match `N(t)`. Cohere measured fewer unique experts than uniform for consecutive tokens (20.36 against 29.1 for a 4-token window, a ratio of 0.71 at 32 tokens), which MoESD's formula does not model, so the true saturation threshold for natural text is above the uniform value. The two are held side by side; they use different models and token sources. — [source](https://arxiv.org/html/2505.19645v3)
- Hardware: MoESD found GPUs with a higher roofline ridge point give larger speedups; no source in this batch gives the Apple-silicon ridge point for the same analysis. — [source](https://arxiv.org/html/2505.19645v3)
- Whether `T_T(B,gamma) / T_T(B,1)` follows the paper's two-factor shape on an M-series GPU served by oMLX or vllm-metal at batch 8 to 32. No source measures it. — source: `asserted`
- Whether a Mac multi-agent workload that batches tens of requests reaches the saturation threshold given routing correlation. — source: `asserted`
- MoESD (arXiv 2505.19645) claims SD can help sparse MoEs more than dense models at moderate batch sizes, and that the useful batch range widens as the MoE gets sparser. — [source](https://arxiv.org/html/2505.19645v3)
- The paper reports a peak speedup of 2.29x for Qwen2-57B-A14B on one of its hardware platforms. — [source](https://arxiv.org/html/2505.19645v3)
- Its two target and draft pairs are Qwen2-57B-A14B-Instruct with Qwen2-0.5B-Instruct, and Mixtral-8x7B-Instruct with an EAGLE head; the dense control is OPT-30B with OPT-350M. — [source](https://arxiv.org/html/2505.19645v3)
- Experiments ran in vLLM on 2xGPU-A, 2xGPU-B, 4xGPU-A and 4xGPU-C (hardware names anonymized), on HumanEval and MT-Bench prompts of 38 to 391 and 5 to 356 tokens. — [source](https://arxiv.org/html/2505.19645v3)
- On 2xGPU-A, Qwen2 peak speedup on HumanEval at temperature 0 was 1.63x, 1.96x and 2.18x for draft length gamma of 2, 3 and 4, and Mixtral 1.67x, 1.69x and 1.79x. — [source](https://arxiv.org/html/2505.19645v3)
- On 2xGPU-B at temperature 0 on HumanEval, Qwen2 reached 1.71x, 2.01x and 2.29x for gamma of 2, 3 and 4. — [source](https://arxiv.org/html/2505.19645v3)
- Speedup rose then fell as batch size grew, and target efficiency tracked the same trend while the acceptance rate fluctuated only in a small range. — [source](https://arxiv.org/html/2505.19645v3)
- Target efficiency for the MoE first increased then decreased with batch size, while the dense model's decreased continuously. — [source](https://arxiv.org/html/2505.19645v3)
- The paper says MoE SD speedups become more pronounced than dense when the batch size exceeds 16. — [source](https://arxiv.org/html/2505.19645v3)
- Lowering sparsity (smaller K) moved the batch size of maximal speedup to larger values and widened the batch range above `x / sqrt(2)`. — [source](https://arxiv.org/html/2505.19645v3)
- The analytic model used 21 GPU measurements to fit its parameters and matched the sparsity and draft-length sweeps. — [source](https://arxiv.org/html/2505.19645v3)
- The authors name private serving with moderate batches of tens of requests, latency-critical settings and memory-restricted settings as the target use cases. — [source](https://arxiv.org/html/2505.19645v3)
- Cohere's separate GPU measurement reports the same non-monotonic shape and a sweet spot at moderate batch size that moves right for sparser models. — [source](https://cohere.com/blog/mixture-of-experts-models-get-more-from-speculative-decoding)
