exo prefill/decode disaggregation
Parent: Mac local LLMs: Clusters, RDMA, exo and ds4 · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Existing coverage: exo-cluster-software.md already records the DGX Spark plus M3 Ultra blog demo (4.34 s vs 6.42 s) and 'MLX P/D landed in placement_utils (PR #1993, Apr 28 2026)'; only absent claims follow
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Existing coverage: exo-cluster-software.md already records the DGX Spark plus M3 Ultra blog demo (4.34 s vs 6.42 s) and 'MLX P/D landed in placement_utils (PR #1993, Apr 28 2026)'; only absent claims follow [source]
- exo PR #1993 'MLX P/D' was opened by rltakashige on Apr 27 2026 with the one-line motivation 'MLX only prefill server for Apple Silicon' [source]
- The PR title changed from 'Prefill/Decode disaggregation' to 'P/D' to 'MLX P/D' on Apr 27 2026 and the first commit was 'Prefill/Decode disaggregation (with UI)' [source]
- The branch leo/prefill was force-pushed more than 30 times on Apr 27 2026 and a commit 'Add prefill decode benchmarks' was part of it [source]
- The PR was approved by Evanev7 with the comment 'eh its alright i guess', auto-merged by squash (merge commit f0d1371) with 6 checks passing, and the branch was deleted on Apr 28 2026 [source]
- An earlier attempt, PR #1776 'Leo/prefill decode really', was closed without merge on May 2 2026 after #1993 landed [source]
- Code touched by P/D, as seen in a downstream conflict resolution: src/exo/worker/engines/mlx/disaggregated/adapter.py, remote_prefill scaffolding in generator/generate.py and batch_generate.py, and a prefill_only filter in src/exo/master/main.py (P/D-aware routing) [source]
- The disaggregated adapter raises NotImplementedError for cache types it cannot serialise; after mlx-lm PR 1192 removed DeepseekV4Cache, DeepSeek V4's new PoolingCache is not supported on the P/D wire protocol, so DeepSeek V4 cannot use MLX P/D [source]
- Engine abstraction PR #2000 (Builder plus Engine ABC, MlxBuilder, MfluxBuilder) merged right after #1993 and moved runner construction out of runner.py [source]
- A downstream fork keeps its own DSV4_PREFILL_STEP_SIZE=256 on top of upstream remote_prefill, so prefill chunk size is not governed by the P/D path [source]
- The merged P/D is Apple-Silicon-to-Apple-Silicon MLX only; the NVIDIA Spark prefill demo in the exo blog is a separate earlier path and no fetched source shows it merged [source]
Corrections and disagreements
- CONTRADICTS exo-cluster-software.md only in location: the fetched evidence places P/D routing in master/main.py and a worker disaggregated adapter, and no fetched source mentions placement_utils; the existing claim is unconfirmed rather than disproved [source]
Children
- No children recorded.