<!-- llms-explorer concept facts · https://llms-explorer.com/tree/rustane-warm-page-cache-pread-fanout-ncdrone-rus/ · pack 2026-10-05 · ~759 tokens -->

# rustane warm-page-cache pread fanout (ncdrone/rustane)

> Anemll's docs credit rustane as the inspiration for cached-read fanout, describe the idea in prose and publish no rustane figures.

Parent: [Mac local LLMs: MoE streaming and offload](https://llms-explorer.com/tree/mac-local-llms-moe-streaming-and-offload/) · 1 facets · 11 facts · page: https://llms-explorer.com/tree/rustane-warm-page-cache-pread-fanout-ncdrone-rus/

## Facts

- Anemll's docs credit rustane as the inspiration for cached-read fanout, describe the idea in prose and publish no rustane figures. — [source](https://raw.githubusercontent.com/Anemll/flash-moe/m5-nax/docs/cache-pread-microbench.md)
- ncdrone/rustane describes itself as a Rust training and inference engine for the Apple Neural Engine and Metal GPU, validated for training from 48M to 5B parameters on an M4 Max 128 GB using reverse-engineered private ANE APIs. — [source](https://github.com/ncdrone/rustane)
- The rustane README sections are Benchmarking, Checkpoint Inference (a generate CLI with Metal KV-cache decode) and HTTP Serving (OpenAI-style completion routes); none concerns expert streaming or page-cache reads. — [source](https://github.com/ncdrone/rustane)
- The rustane results directory holds ANE and GPU TFLOPS, IOSurface staging, dual-load, f16 conversion, training probe and Rust-versus-Objective-C comparison notes, and no pread or page-cache note. — [source](https://github.com/ncdrone/rustane/tree/master/results)
- The rustane bench directory holds three Python scripts (dual_load.py, gpu_metal_matmul.py, gpu_tflops.py). — [source](https://github.com/ncdrone/rustane/tree/master/bench)
- rustane's CREDITS file lists Anemll for ANE inference tricks and does not mention Flash-MoE or page-cache reads. — [source](https://github.com/ncdrone/rustane/blob/master/CREDITS.md)
- Anemll says its cache I/O split experiment "is not a code import from rustane" and is built on Flash-MoE's existing async routed-expert pread path. — [source](https://raw.githubusercontent.com/Anemll/flash-moe/m5-nax/docs/cache-io-split-experiment.md)
- Anemll's cachebench note says its scaling curve is "close to the pattern reported in rustane", naming only the shape, one worker near raw-storage speed then a sharp multi-worker climb. — [source](https://raw.githubusercontent.com/Anemll/flash-moe/m5-nax/docs/cache-pread-microbench.md)
- The Anemll cachebench with split greater than 1 schedules the same total bytes as split 1 (for 128 experts and split 4, 512 tasks instead of 128), so the gain comes from more concurrent cached reads, not extra data. — [source](https://raw.githubusercontent.com/Anemll/flash-moe/m5-nax/docs/cache-pread-microbench.md)
- ncdrone/rustane master shows its last commit on 2026-04-03 and 181 stars at the 2026-10-04 fetch. — [source](https://github.com/ncdrone/rustane)
- Nothing in rustane's public default branch supports or refutes the warm-cache pread scaling curve; the Anemll cachebench is the only measured evidence for that curve. — source: `asserted`
