Apple libthunderboltrdma tbt_post_recv SIGSEGV
Parent: Mac local LLMs: Clusters, RDMA, exo and ds4 · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
#4192 reproduces the SIGSEGV by unplugging a Thunderbolt cable mid-collective: both ranks die, 4/4 pulls, on 2x Mac mini M4 Pro, macOS 26.5.1, mlx at 596dc79f4
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- #4192 reproduces the SIGSEGV by unplugging a Thunderbolt cable mid-collective: both ranks die, 4/4 pulls, on 2x Mac mini M4 Pro, macOS 26.5.1, mlx at 596dc79f4 [source]
- The #4192 repro is a bare C++ loop of jaccl all_sum over 256 floats, with no model and no Python [source]
- #4192 trace goes through `jaccl::MeshImpl::all_reduce<float>` while #4319 goes through `RingImpl::all_reduce<2,unsigned char>` [source]
- The last iteration before the #4192 crash takes about 200 ms versus about 3 us steady state [source]
- Affected builds reported: macOS 26.5.1 (#4192), macOS 26.6.1 25G76 with mlx 0.32.0 and 0.32.2 (#4319 and production) [source]
- On a 5x M3 Ultra TB5 mesh, links dropped on their own on 2026-09-18 and 09-20 and each time the rank at one end died with the tbt_post_recv trace via MeshImpl::all_gather [source]
- That 5-node group had run 39 h with 1.25 M tokens and zero JACCL events before the first drop, so it is a link-loss path, not steady-state instability [source]
- After the crash the node keeps about 140 GB wired and the port stays PORT_DOWN until reboot [source]
- Status: #4192 and #4319 are open with labels bug and distributed, no assignee, no linked PR; #4319 has no linked fix [source]
- Apple FB24371487 was filed for the no-unplug #4319 incident [source]
- A 1 MiB unique all_gather keepalive does not certify the 256 MiB all_reduce path [source]
- PR #4530 (open, from Sofille65) does not fix the crash; it adds side-channel peer-liveness so survivors raise `[jaccl] peer is gone` within about 0.25 s instead of spinning, plus JACCL_PROGRESS_TIMEOUT_S (default 600 s) [source]
- Apple's TB RDMA UC queue pairs give no error completion for a dead peer and RC queue pairs fail with errno 102 [source]
- Per-device limits measured: 10 queue pairs (errno 16 on the 11th), at least 512 MRs, no QP or MR leak across clean exit, _exit or SIGKILL [source]
- Workaround: no known prevention; mitigate with the #4530 liveness patch on survivors, an orchestrator that purges and re-forms the group, checking the exit code of each rank (mlx.launch prints 0), and rebooting the crashed node [source]
- The macOS 27 dylib rework claim is second-hand from dyld-cache diffs and unverified on hardware [source]
Children
- No children recorded.