MoE-PHDS post-hoc declared sparsity fine-tune
Parent: Mac local LLMs: MoE streaming and offload · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
arXiv 2509.23012 is the preprint; Apple's research page is dated December 2025.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- arXiv 2509.23012 is the preprint; Apple's research page is dated December 2025. [source]
- Table 6 of the paper positions PHDS as the first method giving global runtime sparsity control from one checkpoint, against token-level methods such as DynaMoE, AdaMoE, TC-MoE and top-k(P). [source]
- Training and evaluation are restricted to k at most k_pre, because expert parameters were pretrained only up to k_pre. [source]
- Anchoring to a k above the lowest trained value hurts the lowest evaluated k. [source]
- Below k_pre/2, degradation grows and the authors call larger reductions best-effort. [source]
- All reported models are 1B to 14.3B total parameters. [source]
- PHDS Multi-k Training samples ktrain uniformly from a candidate set K_train on every forward pass, with every value no larger than k_pre, and stores and updates auxiliary components (load-balancing loss, layer norms) per k. [source]
- PHDS Curriculum Anchoring anneals training from multi-k toward a fixed lower k such as 2, to stabilise expert dynamics when evaluation k is far below k_pre. [source]
- PHDS masks experts that are inside top-k_pre but outside the sampled top-k with a small epsilon weight instead of zero; the authors chose epsilon 1e-6 after ablations from 1e-1 to 1e-8 showed no material difference below 1e-4. [source]
- The authors use a soft mask partly for deployment ease, because JAX works best without changes to array sizes. [source]
- The paper's default sampling set is K_train = {k_pre/2, ..., k_pre-1, k_pre}. [source]
- PHDS picks one checkpoint after anchoring, by best validation at a target k or by average over a range of k, and treats k_pre as the default operating point. [source]
- The pretrained models in the paper are OLMoE-1B-7B-0125 (64 experts, k_pre 8, 1B active of 7B), Qwen1.5-MoE-A2.7B (60 experts, k_pre 4, 2.7B active of 14.3B), and internal 1.032B-parameter baselines with 16 experts and k_pre of 2, 4 and 6. [source]
- The OLMoE and Qwen runs fine-tune on Tulu-3 with 39B initial tokens and 7.8B ablation tokens on 8 A100-40GB GPUs, while the internal-baseline runs on CommonSense170k use 491M tokens on 2 A100-40GB GPUs. [source]
- In the OLMoE Tulu-3 table at k_ev=4 the base model scores 0.6646 average multiple choice for the k=8 oracle, 0.6668 for PHDS [4,5,6,7,8], 0.6590 for PHDS [4,5,6,7,8] anchored to 5, and 0.6587 for a naive 8-to-4 fine-tune. [source]
- The authors state that on OLMoE and Qwen all SFT methods are similar and robust to sparsity, and that less-tuned internal models show the consistent gains from PHDS. [source]
- For the chat-tuned Qwen1.5-MoE, relative accuracy and perplexity degradation from k=4 is 0.2-0.8% and 1.5-1.6% at k=3 and 2.3-2.6% and 6.1-7.2% at k=2; for the base model it is 0.2-0.4% and 1.2-1.3% at k=3 and 1.6-2.0% and 5.7-5.8% at k=2. [source]
- TriviaQA accuracy can rise at a reduced evaluation k. [source]
- On OLMoE fine-tuned on out-of-distribution CommonSense170k, attention refits carry more of the gain as evaluation k falls, while expert refits dominate at k_ev equal to k_pre. [source]
- In the internal-baseline curriculum ablation, anchoring [4,5] to 6 with k_pre 6 scored 0.5745 and 0.6228 in the two lowest populated evaluation columns, against 0.6808 and 0.7206 when anchored to 4. [source]
- The 7-22% cross-sparsity agreement gain compares one PHDS checkpoint against two separate well-specified oracle checkpoints, at the (2,4) and (4,6) pairs, at similar accuracy. [source]
- The paper's practical guidance is that operators can often reduce k by 20-30% with minimal loss and should treat larger reductions as best-effort. [source]
- The authors propose PHDS for service-level-aware serving (declare k by user tier, request type, latency budget or load) and energy-aware inference, and say it composes with token-adaptive MoE methods that spend a fixed global budget. [source]
- Apple's page lists the authors as Hannah, Zibakhsh, Nishu, Kundu, Samragh Razlighi, Farajtabar and Cho, and says the method needs no architectural change. [source]
- PHDS is a training-time method, so a Mac user running an already-quantized public MoE checkpoint cannot obtain its benefit by changing a runtime flag. [source]
Children
- No children recorded.