Flash-MoE paper flash_moe.pdf 90+ experiment log
Parent: Mac local LLMs: MoE streaming and offload · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
`paper/flash_moe.pdf` is the write-up titled "Flash-MoE: Streaming a 397B Parameter Mixture-of-Experts Model from NVMe at 5.7 Tokens/Second on Consumer Hardware". The byline lists "Claude Opus 4.6" as primary author and Daniel Woods, and the paper reports 90 experiments in about 24 hours.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- `paper/flash_moe.pdf` is the write-up titled "Flash-MoE: Streaming a 397B Parameter Mixture-of-Experts Model from NVMe at 5.7 Tokens/Second on Consumer Hardware". The byline lists "Claude Opus 4.6" as primary author and Daniel Woods, and the paper reports 90 experiments in about 24 hours. [source]
- Expert size is 7,077,888 bytes at 4-bit and 3,932,160 bytes at 2-bit, a 44.5% cut. The README's "6.75 MB" is 7,077,888 bytes expressed in MiB. [source]
- 2-bit requantization is a second stage over the 4-bit values: dequantize each group of 64, set `s2 = (max - min) / 3` and `b2 = min`, then `q = clamp(round((f - b2) / s2), 0, 3)`, keeping group size 64 and 16 values per uint32. [source]
- The reported 2-bit RMSE (0.0017 gate, 0.0018 up, 0.0023 down on average, rising from 0.0012-0.0018 in layers 0-14 to 0.0021-0.0028 in layers 45-59) is measured against the dequantized 4-bit weights, not against the original model. [source]
- The final per-layer budget (2-bit, no `F_NOCACHE`, no app cache) is 2.9 ms: expert I/O 1.37 ms (47%), CMD1 wait 0.87 ms (30%), CMD2 wait 0.45 ms (15%), CPU attention 0.15 ms (5%), other 0.06 ms (2%). [source]
- The I/O of 1.37 ms per layer is about 15.7 MB (4 experts x 3.93 MB) at an effective 11.5 GB/s, below the 17.5 GB/s peak because of four scattered `pread` calls. [source]
- The paper's bottleneck model at the `F_NOCACHE` configuration is 1.49 ms I/O + 1.32 ms GPU + 0.33 ms CPU = 3.14 ms per layer, 188 ms per token (5.3 tok/s), 60% of the 53.9 ms I/O floor, 29% of the I/O-theoretical maximum. [source]
- The paper's I/O floor is 943 MB per token (4 experts x 60 layers x 3.93 MB, 2-bit) over 17.5 GB/s, 53.9 ms or about 18.6 tok/s. [source]
- A 3.9 MB `mmap` read takes about 244 page faults at 16 KB pages, against one system call and one DMA for `pread`; the appendix table lists `mmap` experts as "5x slower" and says per-page faults are 100x slower than one `pread` for uncached data. [source]
- Session caching persists the KV cache and the linear-attention state between turns, so a second message prefills only about 14 new tokens; for a 10-turn, 2,000-token chat it avoids about 98% of reprocessing. [source]
- A prefill change (section heading lost in extraction) cut time to first token 53%, from 5.6 s to 2.6 s for a 16-token prompt, by eliminating expert I/O for N-1 of N prefill tokens. [source]
- The experiment trajectory has five phases: runs 1-20 on MLX (best 3.14 tok/s), 20-40 plateau broken by K=4 pruning, 40-55 two-pass overlap (6 tok/s but degraded output), 55-72 the C/Metal rewrite (GPU-only 14 tok/s, end-to-end 5.29), 72-84 2-bit and polish, ending at experiment 90 "Trust OS" 5.74 tok/s. [source]
- The method was Karpathy's "autoresearch" protocol adapted for systems work: the model proposed changes, ran benchmarks and kept or reverted, with a 5-minute limit per experiment and mandatory before-and-after runs. Three human prompts are named as decisive: treat experts like a database, trust the OS cache, and look at the hardware cache hierarchy. [source]
- The paper says 42% of experiments were discarded and that the commonest failure was broken output from optimizations that silently corrupted the computation (cache eviction races, routing-buffer mapping bugs, wrong attention). [source]
- Variable K: K=0 on the middle 40 layers gave 10 tok/s with blank output, and K=2 or 3 on those layers still collapsed, so every layer needs K of at least 4. [source]
- The K table (K=10: 1.20 tok/s and 2,359 MB per token; K=6: 1.91 and 1,415; K=4: 3.15 and 943; K=3 and 2: EOS only after one or two words) mixes units: its I/O column uses 2-bit expert sizes while its K=4 speed (3.15) matches the MLX-era 4-bit row (3.14) of Table 1. [source]
- A C extension for `pread` reached 18.8 GB/s standalone but only 2.6 GB/s inside MLX because of GIL contention. [source]
- An 18 GB Metal expert cache caused OS page-outs and was a net loss despite a 52% hit rate; keeping delta-net state on GPU (195 MB per-layer buffers) caused memory pressure. [source]
- Speculative routing predicted next-token experts with 53% accuracy and ran 38% slower from cache pollution. [source]
- Speculative decoding with Qwen3.5-35B-A3B (37 tok/s resident) as the draft got 33% acceptance at 4 draft tokens and a net 4.5x slowdown, because its routing diverges from the 397B model's. [source]
- Routing statistics over 200 tokens: consecutive-token overlap is 8-34% per layer at K=4 (0 to 1 of 4 experts reused), cross-layer correlation is near zero, and per-layer LRU hit rates are 59% (32 entries), 68% (64) and 71% (128), read from a garbled table. [source]
- Quality checks are three qualitative prompts (factoring 255, a Python `sorted()` with `key=len`, a probability explanation). The paper reports no perplexity or benchmark score, so "indistinguishable from 4-bit" rests on output coherence. [source]
- The 4-bit rows of Table 1 are not comparable with each other because the paper says page-cache state (cold or warm) changes them; for example "+BLAS delta-net" shows 3.10 tok/s after 5.29. [source]
- Hit rate source. CLAUDE.md says the OS page cache achieves about 71% hit rate. The paper's 71% is the 128-entry per-layer LRU hit rate from its routing analysis; the paper gives no measured OS page-cache hit rate. These may be one number used two ways. [source]
- Per-token read size. The introduction says about 600 MB of expert weights per token; section 3.6 and the budget say 943 MB (2-bit, K=4, no cache). The two are not reconciled in the text. [source]
- Resident memory. The abstract says only 5.5 GB resident; the conclusion and the appendix total say about 6.5 GB. The 5.5 GB is the non-expert weights alone. [source]
- Projection versus later data. The paper projects 10+ tok/s for 397B within two to three hardware generations at about 20% SSD gain per generation; the Anemll fork reports 14.5 tok/s (2-bit) on an M5 Max, two chip generations after the M3 Max, although the existing dossiers note those I/O times come from warm-cache runs. [source]
- The supplementary 90-entry log with git hashes is cited but is not in the PDF text; whether `results.tsv` (58 rows per the README) is a subset is not stated. [source]
- Table 8's projected tok/s values were lost in extraction. [source]
- The README's "~6.75MB" per 4-bit expert is MiB; the paper's 7.08 MB is the same 7,077,888 bytes. [source]
- The paper's headline "38% speedup" is 4.11 to 5.74 tok/s, which is +39.7%. [source]
- The paper's 2-bit F_NOCACHE gain is 5.55 to 5.68 tok/s (+2.3%) and Trust OS is 5.74 (+1.1% over F_NOCACHE). [source]
- The paper credits Claude Opus 4.6 as primary author, reports 90 experiments in about 24 hours and says 42% were discarded. [source]
- The paper's experiment loop used a 5-minute limit per experiment with immediate revert on regression. [source]
- The paper's final per-layer budget is 1.37 ms I/O, 0.87 ms CMD1 wait, 0.45 ms CMD2 wait, 0.15 ms CPU attention and 0.06 ms other (2.9 ms total). [source]
- The paper's effective expert read rate is 11.5 GB/s against 17.5 GB/s peak. [source]
- The paper's 2-bit requantization is a per-group affine second stage with `s2 = (max - min) / 3` and RMSE measured against 4-bit values. [source]
- The paper's variable-K test (K=0 on 40 middle layers) gave 10 tok/s and blank output. [source]
- The paper's C-extension `pread` fell from 18.8 GB/s to 2.6 GB/s under MLX GIL contention. [source]
- The paper's draft-model speculative decoding with a 35B-A3B model had 33% acceptance and a 4.5x net slowdown. [source]
- The paper measured 8-34% consecutive-token expert overlap per layer at K=4 and near-zero cross-layer correlation. [source]
- The paper's session cache avoids about 98% of reprocessing in a 10-turn, 2,000-token chat. [source]
- The paper's quality evidence is three qualitative prompts and no perplexity. [source]
- The paper's 71% hit rate is a 128-entry LRU figure, not a measured OS page-cache rate. [source]
- The paper states ANE co-processing is unused and estimates about 16 TFLOPS FP16 on M3 Max, noting MoE top-K routing does not fit the ANE's restricted programming model. [source]
Corrections and disagreements
- CONTRADICTS: moe-expert-offload-to-ssd-on-macos.md (4-bit "about 1.6 GB per token ... ceiling of about 18.6 tok/s at 17.5 GB/s"): 1.6 GB at 17.5 GB/s is about 93 ms or 10.8 tok/s; the 18.6 tok/s ceiling is the paper's 2-bit figure for 943 MB. [source]
Children
- No children recorded.