ds4 streaming expert cache sizing under TP
Parent: Mac local LLMs: Clusters, RDMA, exo and ds4 · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Because ownership skips the peer's experts before they are fetched, a rank's hot set is about half of what a single Mac sees; a cache count sized for single-Mac use wastes about half its slots on experts this rank will never request.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Because ownership skips the peer's experts before they are fetched, a rank's hot set is about half of what a single Mac sees; a cache count sized for single-Mac use wastes about half its slots on experts this rank will never request. [source]
- `ds4_streaming_cacheable_expert_count` returns `layers * DS4_N_EXPERT` with no TP divisor [source]
- The "cacheable by this layer slice" cap keys on the pipeline `distributed.role` and layer range, not TP rank [source]
- A manual `--ssd-streaming-cache-experts` byte target is capped at 7/8 of the GPU recommended working-set size minus context and graph memory, rounded down to GiB with a 1 GiB floor [source]
- If the prefill headroom is at least the cache target, ds4 reduces the headroom to the target minus one expert and prints "prefill reserve reduced" [source]
- If no complete expert slot fits, ds4 prints "leaves no complete expert slot; using direct per-layer reads" and uses a cache budget of 0 [source]
- Each TP rank sizes its cache from its own working-set size, and no budget is exchanged between ranks [source]
- Metal streaming fast paths such as the GLM tiny pair-SwiGLU path are gated on `tp_world < 2` [source]
- On ROCm the GLM streaming cache can grow by the prefill headroom after prefill unless `DS4_ROCM_GLM_STREAMING_GROW_CACHE_AFTER_PREFILL=0` [source]
- A cache count sized for single-Mac use wastes about half its slots under TP, since the peer's experts are skipped before fetch [source]
Children
- No children recorded.