<!-- llms-explorer concept facts · https://llms-explorer.com/tree/apple-libthunderboltrdma-tbt-post-recv-sigsegv/ · pack 2026-10-05 · ~866 tokens -->

# Apple libthunderboltrdma tbt_post_recv SIGSEGV

> #4192 reproduces the SIGSEGV by unplugging a Thunderbolt cable mid-collective: both ranks die, 4/4 pulls, on 2x Mac mini M4 Pro, macOS 26.5.1, mlx at 596dc79f4

Parent: [Mac local LLMs: Clusters, RDMA, exo and ds4](https://llms-explorer.com/tree/mac-local-llms-clusters-rdma-exo-ds4/) · 1 facets · 16 facts · page: https://llms-explorer.com/tree/apple-libthunderboltrdma-tbt-post-recv-sigsegv/

## Facts

- #4192 reproduces the SIGSEGV by unplugging a Thunderbolt cable mid-collective: both ranks die, 4/4 pulls, on 2x Mac mini M4 Pro, macOS 26.5.1, mlx at 596dc79f4 — [source](https://github.com/ml-explore/mlx/issues/4192)
- The #4192 repro is a bare C++ loop of jaccl all_sum over 256 floats, with no model and no Python — [source](https://github.com/ml-explore/mlx/issues/4192)
- #4192 trace goes through `jaccl::MeshImpl::all_reduce<float>` while #4319 goes through `RingImpl::all_reduce<2,unsigned char>` — [source](https://github.com/ml-explore/mlx/issues/4319)
- The last iteration before the #4192 crash takes about 200 ms versus about 3 us steady state — [source](https://github.com/ml-explore/mlx/issues/4192)
- Affected builds reported: macOS 26.5.1 (#4192), macOS 26.6.1 25G76 with mlx 0.32.0 and 0.32.2 (#4319 and production) — [source](https://github.com/ml-explore/mlx/issues/4192)
- On a 5x M3 Ultra TB5 mesh, links dropped on their own on 2026-09-18 and 09-20 and each time the rank at one end died with the tbt_post_recv trace via MeshImpl::all_gather — [source](https://github.com/ml-explore/mlx/issues/4192)
- That 5-node group had run 39 h with 1.25 M tokens and zero JACCL events before the first drop, so it is a link-loss path, not steady-state instability — [source](https://github.com/ml-explore/mlx/issues/4192)
- After the crash the node keeps about 140 GB wired and the port stays PORT_DOWN until reboot — [source](https://github.com/ml-explore/mlx/issues/4192)
- Status: #4192 and #4319 are open with labels bug and distributed, no assignee, no linked PR; #4319 has no linked fix — [source](https://github.com/ml-explore/mlx/issues/4319)
- Apple FB24371487 was filed for the no-unplug #4319 incident — [source](https://github.com/ml-explore/mlx/issues/4319)
- A 1 MiB unique all_gather keepalive does not certify the 256 MiB all_reduce path — [source](https://github.com/ml-explore/mlx/issues/4319)
- PR #4530 (open, from Sofille65) does not fix the crash; it adds side-channel peer-liveness so survivors raise `[jaccl] peer is gone` within about 0.25 s instead of spinning, plus JACCL_PROGRESS_TIMEOUT_S (default 600 s) — [source](https://github.com/ml-explore/mlx/pull/4530)
- Apple's TB RDMA UC queue pairs give no error completion for a dead peer and RC queue pairs fail with errno 102 — [source](https://github.com/ml-explore/mlx/pull/4530)
- Per-device limits measured: 10 queue pairs (errno 16 on the 11th), at least 512 MRs, no QP or MR leak across clean exit, _exit or SIGKILL — [source](https://github.com/ml-explore/mlx/pull/4530)
- Workaround: no known prevention; mitigate with the #4530 liveness patch on survivors, an orchestrator that purges and re-forms the group, checking the exit code of each rank (mlx.launch prints 0), and rebooting the crashed node — [source](https://github.com/ml-explore/mlx/pull/4530)
- The macOS 27 dylib rework claim is second-hand from dyld-cache diffs and unverified on hardware — [source](https://github.com/ml-explore/mlx/issues/4192)
