<!-- llms-explorer concept facts · https://llms-explorer.com/tree/exo-prefill-decode-disaggregation/ · pack 2026-10-05 · ~821 tokens -->

# exo prefill/decode disaggregation

> Existing coverage: exo-cluster-software.md already records the DGX Spark plus M3 Ultra blog demo (4.34 s vs 6.42 s) and 'MLX P/D landed in placement_utils (PR #1993, Apr 28 2026)'; only absent claims follow

Parent: [Mac local LLMs: Clusters, RDMA, exo and ds4](https://llms-explorer.com/tree/mac-local-llms-clusters-rdma-exo-ds4/) · 2 facets · 12 facts · page: https://llms-explorer.com/tree/exo-prefill-decode-disaggregation/

## Facts

- Existing coverage: exo-cluster-software.md already records the DGX Spark plus M3 Ultra blog demo (4.34 s vs 6.42 s) and 'MLX P/D landed in placement_utils (PR #1993, Apr 28 2026)'; only absent claims follow — source: `asserted`
- exo PR #1993 'MLX P/D' was opened by rltakashige on Apr 27 2026 with the one-line motivation 'MLX only prefill server for Apple Silicon' — [source](https://github.com/exo-explore/exo/pull/1993)
- The PR title changed from 'Prefill/Decode disaggregation' to 'P/D' to 'MLX P/D' on Apr 27 2026 and the first commit was 'Prefill/Decode disaggregation (with UI)' — [source](https://github.com/exo-explore/exo/pull/1993)
- The branch leo/prefill was force-pushed more than 30 times on Apr 27 2026 and a commit 'Add prefill decode benchmarks' was part of it — [source](https://github.com/exo-explore/exo/pull/1993)
- The PR was approved by Evanev7 with the comment 'eh its alright i guess', auto-merged by squash (merge commit f0d1371) with 6 checks passing, and the branch was deleted on Apr 28 2026 — [source](https://github.com/exo-explore/exo/pull/1993)
- An earlier attempt, PR #1776 'Leo/prefill decode really', was closed without merge on May 2 2026 after #1993 landed — [source](https://github.com/exo-explore/exo/pull/1993)
- Code touched by P/D, as seen in a downstream conflict resolution: src/exo/worker/engines/mlx/disaggregated/adapter.py, remote_prefill scaffolding in generator/generate.py and batch_generate.py, and a prefill_only filter in src/exo/master/main.py (P/D-aware routing) — [source](https://github.com/exo-explore/exo/pull/1993)
- The disaggregated adapter raises NotImplementedError for cache types it cannot serialise; after mlx-lm PR 1192 removed DeepseekV4Cache, DeepSeek V4's new PoolingCache is not supported on the P/D wire protocol, so DeepSeek V4 cannot use MLX P/D — [source](https://github.com/exo-explore/exo/pull/1993)
- Engine abstraction PR #2000 (Builder plus Engine ABC, MlxBuilder, MfluxBuilder) merged right after #1993 and moved runner construction out of runner.py — [source](https://github.com/exo-explore/exo/pull/1993)
- A downstream fork keeps its own DSV4_PREFILL_STEP_SIZE=256 on top of upstream remote_prefill, so prefill chunk size is not governed by the P/D path — [source](https://github.com/exo-explore/exo/pull/1993)
- The merged P/D is Apple-Silicon-to-Apple-Silicon MLX only; the NVIDIA Spark prefill demo in the exo blog is a separate earlier path and no fetched source shows it merged — source: `asserted`

## Corrections and disagreements

- CONTRADICTS exo-cluster-software.md only in location: the fetched evidence places P/D routing in master/main.py and a worker disaggregated adapter, and no fetched source mentions placement_utils; the existing claim is unconfirmed rather than disproved — source: `asserted`
