JACCL silent hang on peer loss
Parent: Mac local LLMs: Clusters, RDMA, exo and ds4 · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
#3910 captured a live `sample` showing 1156 of 1304 StreamThread samples inside MeshImpl::recv, calling tbt_poll_cq and tbt_poll_qp_recv, with ring_indicies_err on the error path polled in a loop
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- #3910 captured a live `sample` showing 1156 of 1304 StreamThread samples inside MeshImpl::recv, calling tbt_poll_cq and tbt_poll_qp_recv, with ring_indicies_err on the error path polled in a loop [source]
- The main thread of a hung rank sits in Event::wait on IOSurfaceSharedEvent waitUntilSignaledValue, so SIGTERM is never delivered and SIGKILL leaves wired pages held until reboot [source]
- #3910 reproduced on 4x Mac Studio M3 Ultra, macOS 26.4.1 and 26.5 mixed, mlx 0.32.0, TB5 full mesh [source]
- Re-plugging the Thunderbolt cable about a minute after the drop did not unblock the in-flight operation [source]
- The same hang signature appeared once after about 75 h of continuous serving with no physical action and no captured trigger [source]
- On 2x M4 Pro mini with macOS 26.5.1, SIGKILL of rank 1 left the survivor at 100% CPU in all_reduce with 1927 of 2009 samples in tbt_poll_cq [source]
- On the same M4 Pro hardware a physical unplug crashed both ranks instead of hanging [source]
- #4278: a SIGSTOPped rank froze the other three at the same iteration at 99-100% CPU, and SIGCONT let all four resume with no corruption [source]
- Setting en2, en3 and en4 down with ping at 100% loss did not stop a four-rank RDMA collective, since established queue pairs bypass the IP stack [source]
- mlx.launch kills the surviving ranks when any rank exits, which hides what survivors would do [source]
- Ring backend on peer loss: v0.32.0 hangs, main aborts with exit 134, #3742 gives a catchable RuntimeError [source]
- Ring issue #3862 is closed; its fix PR #4060 merged as commit 8c28c38 [source]
- Ring SocketThread::worker returned after 10 failed attempts without resolving pending promises, so every rank blocked in Event::wait; the ring also lacks SO_KEEPALIVE [source]
- PR #4060 chose to exit the process on peer loss because resolving the promises makes survivors return wrong collective results silently [source]
- Throwing from the stream thread reaches std::terminate because StreamThread::thread_fn calls task() with no try/catch [source]
- TN3205 states RDMA over Thunderbolt supports send and receive only, at most 10 UC queue pairs, messages up to 16,773,120 bytes and 4095 outstanding work requests [source]
- The only integration pattern reported before any fix is an application no-progress watchdog that SIGKILLs ranks and reboots leaked nodes [source]
Children
- No children recorded.