ds4 two-Mac tensor parallelism over Thunderbolt RDMA setup
Parent: Mac local LLMs: Clusters, RDMA, exo and ds4 · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
ds4 tensor parallelism runs the same graph on two ranks. Rank 0 (leader) runs the normal frontend session and mirrors every session sync and eval to rank 1 (worker) over a TCP control socket. Inside each decoded token the two ranks swap partial block outputs through a registered memory slab, by t...
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- ds4 tensor parallelism runs the same graph on two ranks. Rank 0 (leader) runs the normal frontend session and mirrors every session sync and eval to rank 1 (worker) over a TCP control socket. Inside each decoded token the two ranks swap partial block outputs through a registered memory slab, by two-sided RDMA SEND/RECV when Thunderbolt RDMA works, otherwise by a full-duplex TCP exchange. [source]
- Launch order: start the worker first (it retries while the coordinator loads), then the coordinator. The worker command is `./ds4 --tensor-parallel --role worker --coordinator 10.99.0.2 9911 --transport rdma --ctx 8192` and the coordinator command is the same with `--role coordinator --listen 10.99.0.2 9911`. Do not pass `--layers`. [source] — privacy-ok (example address quoted from the cited source)
- The coordinator role can be `ds4`, `ds4-agent`, `ds4-server` or `ds4-bench`; the worker always runs `ds4`. A vision model needs the same `--vision FILE` on both. GLM MTP needs `--mtp` on both. V4.1 supports vision but not speculative decoding. [source]
- The leader listens and accepts one worker; the worker dials with retry; both then exchange an identity (GGUF byte size, model id, layer count, embedding width, vocabulary size, quant bits, context size and the gate schedule) and a mismatched pair aborts before any inference. [source]
- Each token exchanges two partials per layer, one after attention and one after the FFN (`DS4_TP_GATES_PER_LAYER = 2`), so a model with L layers makes 2L exchanges per decoded token. [source]
- Each partial is a float32 vector of `n_embd * 4` bytes and is "never quantized on the wire"; the activation-bit option belongs to pipeline mode only. [source]
- The slab holds `S = n_layer * 2` out vectors, `S` in vectors, `S` eight-byte sequence flags (written after each vector, so a flag proves its vector landed) and a 16-byte token slot sent leader to worker. [source]
- ds4 loads Apple's verbs library itself with `dlopen("/usr/lib/librdma.dylib")` and uses an unreliable-connection (UC) queue pair on macOS (reliable connection on Linux), MTU 1024, and global routing through the IPv4-mapped GID. [source]
- ds4 passes the index of the GID it selected as `sgid_index` in the RTR transition. [source]
- Device choice: ds4 looks at one verbs device per Thunderbolt port (`rdma_enN`) and picks the active one; `--rdma-device` and `--rdma-gid-index` override it. [source]
- Work split on two ranks: routed experts as one contiguous half per rank, attention heads split by `tp_head0 = rank * tp_heads`, the shared expert split into halves, prefill rows split into halves, dense weights replicated. [source]
- Speculative verify blocks (at most 5 rows, buffers up to 8) move all rows of a layer in one bulk RDMA transfer; prefill uses a bulk exchange with 2 MB TCP rounds as the fallback. [source]
- RDMA is skipped (TCP is used) when `n_embd * 4` bytes exceeds twice the RDMA maximum message size. [source]
- The TP wire protocol is at `DS4_TP_PROTOCOL_VERSION` 14 with magics `DS4T` and `DS4B`, and the comment on the 14th version says V4.1 CUDA workers "now return half-logit frames after successful work". [source]
- The librdma is loaded at runtime on purpose so builds and machines without RDMA still run. [source]
- Error text when no GID matches: "tp rdma: no IPv4-mapped GID on the active port; give the Thunderbolt interface its own IPv4 (e.g. sudo ifconfig en1 inet 10.99.0.2/30 alias) on both machines". [source] — privacy-ok (example address quoted from the cited source)
- Error text when no port is active: "no device with an active port ... is the peer up and rdma_ctl enabled on both machines?". [source]
- The source comment says the driver "only connects through the IPv4-mapped GID", which exists only when the member interface has its own IPv4 and "the bridge's address does not count". [source]
- A new UC queue pair can drop every send in one direction right after a working run or a role swap; ds4 detects this in its warm-up, recreates the pair and retries. [source]
- TP is exclusive with ds4's pipeline roles and layer slices: the validator rejects `--role` pipeline modes and `load_slice` with TP. [source]
- A TP option such as `--transport`, `--rdma-device` or `--rdma-gid-index` without `--tensor-parallel` and `--role` is an error. [source]
- Tunable environment variables: `DS4_TP_TIMEOUT_SEC` (default 300), `DS4_TP_GATE_TIMEOUT_MS` (default 750), `DS4_TP_DISABLE_VERIFY_WINDOW` and `DS4_TP_BIG_GATE_DEBUG`. [source]
- JACCL's `queue_pair_rtr` hard-codes `sgid_index = 1` while its `info()` accepts an IPv4-mapped GID at any index, so a GID table whose `::ffff:` entry is not at index 1 would be advertised from one index and routed from another. This could explain the errno 96 seen after IPv4 renumbering. [source]
- ds4 versus JACCL on GID handling. ds4 scans, uses the chosen index in RTR and fails early with a fix hint. JACCL scanned without failing until mlx PR 4191 added a similar hint, and still hard-codes index 1 at RTR. [source]
- Header comment versus code on what TP rejects: see aggregate-ssd-bandwidth-when-each-cluster-node-s.md. [source]
- Measured TP decode tok/s and per-gate latency for two-Mac RDMA are not published in the files read; the docs say only that gain depends on model, link and comparison setup. [source]
- Whether macOS 27 changes UC delivery, which would make ds4's one-direction-drop workaround unnecessary. [source]
- ds4 TP launch is worker first, then coordinator, with `--tensor-parallel --role ... --transport rdma` and the coordinator listening on a port such as 9911; `--layers` must not be passed. [source]
- ds4 TP exchanges two float32 partials per layer through a registered slab and never quantizes them on the wire. [source]
- ds4 TP's leader mirrors every session sync and eval to the worker over a TCP control socket so both run the identical graph sequence. [source]
- ds4 TP exchanges an engine identity at connect and aborts a mismatched pair before inference. [source]
- ds4 TP uses its own verbs calls through `/usr/lib/librdma.dylib`, a UC queue pair on macOS, MTU 1024 and the IPv4-mapped GID, not JACCL. [source]
- ds4 TP's sequence flag is written after its vector so a flag implies the vector landed. [source]
- ds4 TP's speculative verify blocks exchange up to 8 rows per layer in one bulk transfer. [source]
- ds4 TP's source says the bridge's IPv4 address does not count for the IPv4-mapped GID. [source]
- ds4 TP recreates a UC queue pair when a fresh pair drops all sends in one direction. [source]
- ds4 TP splits routed experts into contiguous halves, attention heads and shared-expert rows by rank, and replicates dense weights. [source]
- ds4's validator rejects TP combined with pipeline roles or layer slices. [source]
- ds4 TP is a 50/50 split with exactly one worker. [source]
- ds4's TP protocol version is 14. [source]
- A model with L layers costs 2L cross-machine exchanges per decoded token in ds4 TP, against about 2L all-sums in mlx-lm TP, with different transports. [source]
- JACCL's hard-coded `sgid_index = 1` at RTR may disagree with the GID index `info()` chose. [source]
Children
- No children recorded.