<!-- llms-explorer concept facts · https://llms-explorer.com/tree/ds4-streaming-expert-cache-sizing-under-tp/ · pack 2026-10-05 · ~611 tokens -->

# ds4 streaming expert cache sizing under TP

> Because ownership skips the peer's experts before they are fetched, a rank's hot set is about half of what a single Mac sees; a cache count sized for single-Mac use wastes about half its slots on experts this rank will never request.

Parent: [Mac local LLMs: Clusters, RDMA, exo and ds4](https://llms-explorer.com/tree/mac-local-llms-clusters-rdma-exo-ds4/) · 1 facets · 10 facts · page: https://llms-explorer.com/tree/ds4-streaming-expert-cache-sizing-under-tp/

## Facts

- Because ownership skips the peer's experts before they are fetched, a rank's hot set is about half of what a single Mac sees; a cache count sized for single-Mac use wastes about half its slots on experts this rank will never request. — source: `asserted`
- `ds4_streaming_cacheable_expert_count` returns `layers * DS4_N_EXPERT` with no TP divisor — [source](https://raw.githubusercontent.com/antirez/ds4/main/ds4.c)
- The "cacheable by this layer slice" cap keys on the pipeline `distributed.role` and layer range, not TP rank — [source](https://raw.githubusercontent.com/antirez/ds4/main/ds4.c)
- A manual `--ssd-streaming-cache-experts` byte target is capped at 7/8 of the GPU recommended working-set size minus context and graph memory, rounded down to GiB with a 1 GiB floor — [source](https://raw.githubusercontent.com/antirez/ds4/main/ds4.c)
- If the prefill headroom is at least the cache target, ds4 reduces the headroom to the target minus one expert and prints "prefill reserve reduced" — [source](https://raw.githubusercontent.com/antirez/ds4/main/ds4.c)
- If no complete expert slot fits, ds4 prints "leaves no complete expert slot; using direct per-layer reads" and uses a cache budget of 0 — [source](https://raw.githubusercontent.com/antirez/ds4/main/ds4.c)
- Each TP rank sizes its cache from its own working-set size, and no budget is exchanged between ranks — source: `asserted`
- Metal streaming fast paths such as the GLM tiny pair-SwiGLU path are gated on `tp_world < 2` — [source](https://raw.githubusercontent.com/antirez/ds4/main/ds4.c)
- On ROCm the GLM streaming cache can grow by the prefill headroom after prefill unless `DS4_ROCM_GLM_STREAMING_GROW_CACHE_AFTER_PREFILL=0` — [source](https://raw.githubusercontent.com/antirez/ds4/main/ds4.c)
- A cache count sized for single-Mac use wastes about half its slots under TP, since the peer's experts are skipped before fetch — source: `asserted`
