<!-- llms-explorer concept facts · https://llms-explorer.com/tree/moe-phds-post-hoc-declared-sparsity-fine-tune/ · pack 2026-10-05 · ~1331 tokens -->

# MoE-PHDS post-hoc declared sparsity fine-tune

> arXiv 2509.23012 is the preprint; Apple's research page is dated December 2025.

Parent: [Mac local LLMs: MoE streaming and offload](https://llms-explorer.com/tree/mac-local-llms-moe-streaming-and-offload/) · 1 facets · 25 facts · page: https://llms-explorer.com/tree/moe-phds-post-hoc-declared-sparsity-fine-tune/

## Facts

- arXiv 2509.23012 is the preprint; Apple's research page is dated December 2025. — source: `asserted`
- Table 6 of the paper positions PHDS as the first method giving global runtime sparsity control from one checkpoint, against token-level methods such as DynaMoE, AdaMoE, TC-MoE and top-k(P). — source: `asserted`
- Training and evaluation are restricted to k at most k_pre, because expert parameters were pretrained only up to k_pre. — source: `asserted`
- Anchoring to a k above the lowest trained value hurts the lowest evaluated k. — source: `asserted`
- Below k_pre/2, degradation grows and the authors call larger reductions best-effort. — source: `asserted`
- All reported models are 1B to 14.3B total parameters. — source: `asserted`
- PHDS Multi-k Training samples ktrain uniformly from a candidate set K_train on every forward pass, with every value no larger than k_pre, and stores and updates auxiliary components (load-balancing loss, layer norms) per k. — [source](https://arxiv.org/pdf/2509.23012)
- PHDS Curriculum Anchoring anneals training from multi-k toward a fixed lower k such as 2, to stabilise expert dynamics when evaluation k is far below k_pre. — [source](https://arxiv.org/pdf/2509.23012)
- PHDS masks experts that are inside top-k_pre but outside the sampled top-k with a small epsilon weight instead of zero; the authors chose epsilon 1e-6 after ablations from 1e-1 to 1e-8 showed no material difference below 1e-4. — [source](https://arxiv.org/pdf/2509.23012)
- The authors use a soft mask partly for deployment ease, because JAX works best without changes to array sizes. — [source](https://arxiv.org/pdf/2509.23012)
- The paper's default sampling set is K_train = {k_pre/2, ..., k_pre-1, k_pre}. — [source](https://arxiv.org/pdf/2509.23012)
- PHDS picks one checkpoint after anchoring, by best validation at a target k or by average over a range of k, and treats k_pre as the default operating point. — [source](https://arxiv.org/pdf/2509.23012)
- The pretrained models in the paper are OLMoE-1B-7B-0125 (64 experts, k_pre 8, 1B active of 7B), Qwen1.5-MoE-A2.7B (60 experts, k_pre 4, 2.7B active of 14.3B), and internal 1.032B-parameter baselines with 16 experts and k_pre of 2, 4 and 6. — [source](https://arxiv.org/pdf/2509.23012)
- The OLMoE and Qwen runs fine-tune on Tulu-3 with 39B initial tokens and 7.8B ablation tokens on 8 A100-40GB GPUs, while the internal-baseline runs on CommonSense170k use 491M tokens on 2 A100-40GB GPUs. — [source](https://arxiv.org/pdf/2509.23012)
- In the OLMoE Tulu-3 table at k_ev=4 the base model scores 0.6646 average multiple choice for the k=8 oracle, 0.6668 for PHDS [4,5,6,7,8], 0.6590 for PHDS [4,5,6,7,8] anchored to 5, and 0.6587 for a naive 8-to-4 fine-tune. — [source](https://arxiv.org/pdf/2509.23012)
- The authors state that on OLMoE and Qwen all SFT methods are similar and robust to sparsity, and that less-tuned internal models show the consistent gains from PHDS. — [source](https://arxiv.org/pdf/2509.23012)
- For the chat-tuned Qwen1.5-MoE, relative accuracy and perplexity degradation from k=4 is 0.2-0.8% and 1.5-1.6% at k=3 and 2.3-2.6% and 6.1-7.2% at k=2; for the base model it is 0.2-0.4% and 1.2-1.3% at k=3 and 1.6-2.0% and 5.7-5.8% at k=2. — [source](https://arxiv.org/pdf/2509.23012)
- TriviaQA accuracy can rise at a reduced evaluation k. — [source](https://arxiv.org/pdf/2509.23012)
- On OLMoE fine-tuned on out-of-distribution CommonSense170k, attention refits carry more of the gain as evaluation k falls, while expert refits dominate at k_ev equal to k_pre. — [source](https://arxiv.org/pdf/2509.23012)
- In the internal-baseline curriculum ablation, anchoring [4,5] to 6 with k_pre 6 scored 0.5745 and 0.6228 in the two lowest populated evaluation columns, against 0.6808 and 0.7206 when anchored to 4. — [source](https://arxiv.org/pdf/2509.23012)
- The 7-22% cross-sparsity agreement gain compares one PHDS checkpoint against two separate well-specified oracle checkpoints, at the (2,4) and (4,6) pairs, at similar accuracy. — [source](https://arxiv.org/pdf/2509.23012)
- The paper's practical guidance is that operators can often reduce k by 20-30% with minimal loss and should treat larger reductions as best-effort. — [source](https://arxiv.org/pdf/2509.23012)
- The authors propose PHDS for service-level-aware serving (declare k by user tier, request type, latency budget or load) and energy-aware inference, and say it composes with token-adaptive MoE methods that spend a fixed global budget. — [source](https://arxiv.org/pdf/2509.23012)
- Apple's page lists the authors as Hannah, Zibakhsh, Nishu, Kundu, Samragh Razlighi, Farajtabar and Cho, and says the method needs no architectural change. — [source](https://machinelearning.apple.com/research/moe-phds)
- PHDS is a training-time method, so a Mac user running an already-quantized public MoE checkpoint cannot obtain its benefit by changing a runtime flag. — source: `asserted`
