<!-- llms-explorer concept facts · https://llms-explorer.com/tree/flash-moe-paper-flash-moe-pdf-90-plus-experiment/ · pack 2026-10-05 · ~3157 tokens -->

# Flash-MoE paper flash_moe.pdf 90+ experiment log

> `paper/flash_moe.pdf` is the write-up titled "Flash-MoE: Streaming a 397B Parameter Mixture-of-Experts Model from NVMe at 5.7 Tokens/Second on Consumer Hardware". The byline lists "Claude Opus 4.6" as primary author and Daniel Woods, and the paper reports 90 experiments in about 24 hours.

Parent: [Mac local LLMs: MoE streaming and offload](https://llms-explorer.com/tree/mac-local-llms-moe-streaming-and-offload/) · 2 facets · 46 facts · page: https://llms-explorer.com/tree/flash-moe-paper-flash-moe-pdf-90-plus-experiment/

## Facts

- `paper/flash_moe.pdf` is the write-up titled "Flash-MoE: Streaming a 397B Parameter Mixture-of-Experts Model from NVMe at 5.7 Tokens/Second on Consumer Hardware". The byline lists "Claude Opus 4.6" as primary author and Daniel Woods, and the paper reports 90 experiments in about 24 hours. — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
- Expert size is 7,077,888 bytes at 4-bit and 3,932,160 bytes at 2-bit, a 44.5% cut. The README's "6.75 MB" is 7,077,888 bytes expressed in MiB. — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
- 2-bit requantization is a second stage over the 4-bit values: dequantize each group of 64, set `s2 = (max - min) / 3` and `b2 = min`, then `q = clamp(round((f - b2) / s2), 0, 3)`, keeping group size 64 and 16 values per uint32. — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
- The reported 2-bit RMSE (0.0017 gate, 0.0018 up, 0.0023 down on average, rising from 0.0012-0.0018 in layers 0-14 to 0.0021-0.0028 in layers 45-59) is measured against the dequantized 4-bit weights, not against the original model. — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
- The final per-layer budget (2-bit, no `F_NOCACHE`, no app cache) is 2.9 ms: expert I/O 1.37 ms (47%), CMD1 wait 0.87 ms (30%), CMD2 wait 0.45 ms (15%), CPU attention 0.15 ms (5%), other 0.06 ms (2%). — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
- The I/O of 1.37 ms per layer is about 15.7 MB (4 experts x 3.93 MB) at an effective 11.5 GB/s, below the 17.5 GB/s peak because of four scattered `pread` calls. — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
- The paper's bottleneck model at the `F_NOCACHE` configuration is 1.49 ms I/O + 1.32 ms GPU + 0.33 ms CPU = 3.14 ms per layer, 188 ms per token (5.3 tok/s), 60% of the 53.9 ms I/O floor, 29% of the I/O-theoretical maximum. — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
- The paper's I/O floor is 943 MB per token (4 experts x 60 layers x 3.93 MB, 2-bit) over 17.5 GB/s, 53.9 ms or about 18.6 tok/s. — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
- A 3.9 MB `mmap` read takes about 244 page faults at 16 KB pages, against one system call and one DMA for `pread`; the appendix table lists `mmap` experts as "5x slower" and says per-page faults are 100x slower than one `pread` for uncached data. — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
- Session caching persists the KV cache and the linear-attention state between turns, so a second message prefills only about 14 new tokens; for a 10-turn, 2,000-token chat it avoids about 98% of reprocessing. — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
- A prefill change (section heading lost in extraction) cut time to first token 53%, from 5.6 s to 2.6 s for a 16-token prompt, by eliminating expert I/O for N-1 of N prefill tokens. — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
- The experiment trajectory has five phases: runs 1-20 on MLX (best 3.14 tok/s), 20-40 plateau broken by K=4 pruning, 40-55 two-pass overlap (6 tok/s but degraded output), 55-72 the C/Metal rewrite (GPU-only 14 tok/s, end-to-end 5.29), 72-84 2-bit and polish, ending at experiment 90 "Trust OS" 5.74 tok/s. — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
- The method was Karpathy's "autoresearch" protocol adapted for systems work: the model proposed changes, ran benchmarks and kept or reverted, with a 5-minute limit per experiment and mandatory before-and-after runs. Three human prompts are named as decisive: treat experts like a database, trust the OS cache, and look at the hardware cache hierarchy. — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
- The paper says 42% of experiments were discarded and that the commonest failure was broken output from optimizations that silently corrupted the computation (cache eviction races, routing-buffer mapping bugs, wrong attention). — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
- Variable K: K=0 on the middle 40 layers gave 10 tok/s with blank output, and K=2 or 3 on those layers still collapsed, so every layer needs K of at least 4. — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
- The K table (K=10: 1.20 tok/s and 2,359 MB per token; K=6: 1.91 and 1,415; K=4: 3.15 and 943; K=3 and 2: EOS only after one or two words) mixes units: its I/O column uses 2-bit expert sizes while its K=4 speed (3.15) matches the MLX-era 4-bit row (3.14) of Table 1. — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
- A C extension for `pread` reached 18.8 GB/s standalone but only 2.6 GB/s inside MLX because of GIL contention. — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
- An 18 GB Metal expert cache caused OS page-outs and was a net loss despite a 52% hit rate; keeping delta-net state on GPU (195 MB per-layer buffers) caused memory pressure. — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
- Speculative routing predicted next-token experts with 53% accuracy and ran 38% slower from cache pollution. — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
- Speculative decoding with Qwen3.5-35B-A3B (37 tok/s resident) as the draft got 33% acceptance at 4 draft tokens and a net 4.5x slowdown, because its routing diverges from the 397B model's. — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
- Routing statistics over 200 tokens: consecutive-token overlap is 8-34% per layer at K=4 (0 to 1 of 4 experts reused), cross-layer correlation is near zero, and per-layer LRU hit rates are 59% (32 entries), 68% (64) and 71% (128), read from a garbled table. — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
- Quality checks are three qualitative prompts (factoring 255, a Python `sorted()` with `key=len`, a probability explanation). The paper reports no perplexity or benchmark score, so "indistinguishable from 4-bit" rests on output coherence. — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
- The 4-bit rows of Table 1 are not comparable with each other because the paper says page-cache state (cold or warm) changes them; for example "+BLAS delta-net" shows 3.10 tok/s after 5.29. — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
- Hit rate source. CLAUDE.md says the OS page cache achieves about 71% hit rate. The paper's 71% is the 128-entry per-layer LRU hit rate from its routing analysis; the paper gives no measured OS page-cache hit rate. These may be one number used two ways. — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/CLAUDE.md)
- Per-token read size. The introduction says about 600 MB of expert weights per token; section 3.6 and the budget say 943 MB (2-bit, K=4, no cache). The two are not reconciled in the text. — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
- Resident memory. The abstract says only 5.5 GB resident; the conclusion and the appendix total say about 6.5 GB. The 5.5 GB is the non-expert weights alone. — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
- Projection versus later data. The paper projects 10+ tok/s for 397B within two to three hardware generations at about 20% SSD gain per generation; the Anemll fork reports 14.5 tok/s (2-bit) on an M5 Max, two chip generations after the M3 Max, although the existing dossiers note those I/O times come from warm-cache runs. — [source](https://github.com/Anemll/flash-moe)
- The supplementary 90-entry log with git hashes is cited but is not in the PDF text; whether `results.tsv` (58 rows per the README) is a subset is not stated. — source: `asserted`
- Table 8's projected tok/s values were lost in extraction. — source: `asserted`
- The README's "~6.75MB" per 4-bit expert is MiB; the paper's 7.08 MB is the same 7,077,888 bytes. — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
- The paper's headline "38% speedup" is 4.11 to 5.74 tok/s, which is +39.7%. — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
- The paper's 2-bit F_NOCACHE gain is 5.55 to 5.68 tok/s (+2.3%) and Trust OS is 5.74 (+1.1% over F_NOCACHE). — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
- The paper credits Claude Opus 4.6 as primary author, reports 90 experiments in about 24 hours and says 42% were discarded. — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
- The paper's experiment loop used a 5-minute limit per experiment with immediate revert on regression. — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
- The paper's final per-layer budget is 1.37 ms I/O, 0.87 ms CMD1 wait, 0.45 ms CMD2 wait, 0.15 ms CPU attention and 0.06 ms other (2.9 ms total). — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
- The paper's effective expert read rate is 11.5 GB/s against 17.5 GB/s peak. — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
- The paper's 2-bit requantization is a per-group affine second stage with `s2 = (max - min) / 3` and RMSE measured against 4-bit values. — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
- The paper's variable-K test (K=0 on 40 middle layers) gave 10 tok/s and blank output. — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
- The paper's C-extension `pread` fell from 18.8 GB/s to 2.6 GB/s under MLX GIL contention. — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
- The paper's draft-model speculative decoding with a 35B-A3B model had 33% acceptance and a 4.5x net slowdown. — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
- The paper measured 8-34% consecutive-token expert overlap per layer at K=4 and near-zero cross-layer correlation. — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
- The paper's session cache avoids about 98% of reprocessing in a 10-turn, 2,000-token chat. — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
- The paper's quality evidence is three qualitative prompts and no perplexity. — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
- The paper's 71% hit rate is a 128-entry LRU figure, not a measured OS page-cache rate. — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
- The paper states ANE co-processing is unused and estimates about 16 TFLOPS FP16 on M3 Max, noting MoE top-K routing does not fit the ANE's restricted programming model. — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)

## Corrections and disagreements

- CONTRADICTS: moe-expert-offload-to-ssd-on-macos.md (4-bit "about 1.6 GB per token ... ceiling of about 18.6 tok/s at 17.5 GB/s"): 1.6 GB at 17.5 GB/s is about 93 ms or 10.8 tok/s; the 18.6 tok/s ceiling is the paper's 2-bit figure for 943 MB. — [source](https://raw.githubusercontent.com/danveloper/flash-moe/main/paper/flash_moe.pdf)
