JACCL peer-liveness side channel and progress timeout PR 4530
Parent: Mac local LLMs: Clusters, RDMA, exo and ds4 · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
zcbenz's stated reason: it is "an AI wrote solution" that tries to solve the problem in a very hacky way that makes future maintenance harder, and a decent fix is not their focus at the time
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- zcbenz's stated reason: it is "an AI wrote solution" that tries to solve the problem in a very hacky way that makes future maintenance harder, and a decent fix is not their focus at the time [source]
- #4278 stays open after the PR closed, labels bug, distributed and low priority, no assignee, no linked PR [source]
- The check runs in 10 wait loops: 6 in mesh_impl.h and 4 in ring_impl.h, each owning a ProgressGuard [source]
- The guard polls the side-channel fds with zero timeout plus a one-byte MSG_PEEK, after 256 empty polls and then at most every 250 ms [source]
- Rank 0 watches all peers and other ranks watch only rank 0, so a non-zero rank death reaches everyone in two hops [source]
- With a user-supplied all-gather there is no TCP coordinator, the fd list is empty and only the progress timeout protects the group [source]
- The timeout env vars are JACCL_PROGRESS_TIMEOUT_S and MLX_JACCL_PROGRESS_TIMEOUT_S, default 600 s, 0 disables [source]
- Clean-path latency was unchanged at 1.9 ms for a 16 MB all_sum on a 2-node mesh [source]
- Tested scenarios: mesh rank 1 _exit, mesh rank 0 coordinator _exit, ring rank 1 _exit, mesh SIGKILL; each raised on the survivor after 0.25 s and a new group formed afterwards [source]
- On a 5x M3 Ultra full TB5 mesh (20 RDMA edges) serving GLM-5.3 744B 8-bit pipeline-parallel, a crashed rank was reported by the other four within the same second and the orchestrator re-formed the group twice without operator action [source]
- The same group then ran a reasoning benchmark for hours with zero JACCL events and stock per-token latency [source]
- Probe sources used for the QP/MR limit measurements are in the Odyssai-eu/OdyssAI-X repo under scripts/jaccl (rdma_probe.c, rdma_pair.c, smoke_jaccl.py) [source]
- #4278 was found with 4x M4 Pro minis: a SIGSTOPped rank left the other three frozen at 99-100% CPU in state Rs+, and SIGCONT let all four finish with no corruption [source]
- `ifconfig enX down` does not stop a running RDMA collective, because queue pairs bypass the IP stack once established; the /30 addresses only bootstrap the side channel [source]
- Ring peer loss on a 2-rank group: v0.32.0 hangs, current main aborts with exit 134, and mlx PR #3742 turns it into a catchable RuntimeError; JACCL had none of these [source]
- mlx PR #3742 moves error handling from metal::EventImpl to the public Event class and lets the CPU scheduler poison pending events so the exception is thrown where the event is waited on [source]
Corrections and disagreements
- CONTRADICTS apple-libthunderboltrdma-tbt-post-recv-sigsegv.md: PR #4530 is closed, not open; zcbenz closed it on 2026-09-29 without merging [source]
Children
- No children recorded.