<!-- llms-explorer concept facts · https://llms-explorer.com/tree/jaccl-silent-hang-on-peer-loss/ · pack 2026-10-05 · ~934 tokens -->

# JACCL silent hang on peer loss

> #3910 captured a live `sample` showing 1156 of 1304 StreamThread samples inside MeshImpl::recv, calling tbt_poll_cq and tbt_poll_qp_recv, with ring_indicies_err on the error path polled in a loop

Parent: [Mac local LLMs: Clusters, RDMA, exo and ds4](https://llms-explorer.com/tree/mac-local-llms-clusters-rdma-exo-ds4/) · 1 facets · 17 facts · page: https://llms-explorer.com/tree/jaccl-silent-hang-on-peer-loss/

## Facts

- #3910 captured a live `sample` showing 1156 of 1304 StreamThread samples inside MeshImpl::recv, calling tbt_poll_cq and tbt_poll_qp_recv, with ring_indicies_err on the error path polled in a loop — [source](https://github.com/ml-explore/mlx/issues/3910)
- The main thread of a hung rank sits in Event::wait on IOSurfaceSharedEvent waitUntilSignaledValue, so SIGTERM is never delivered and SIGKILL leaves wired pages held until reboot — [source](https://github.com/ml-explore/mlx/issues/3910)
- #3910 reproduced on 4x Mac Studio M3 Ultra, macOS 26.4.1 and 26.5 mixed, mlx 0.32.0, TB5 full mesh — [source](https://github.com/ml-explore/mlx/issues/3910)
- Re-plugging the Thunderbolt cable about a minute after the drop did not unblock the in-flight operation — [source](https://github.com/ml-explore/mlx/issues/3910)
- The same hang signature appeared once after about 75 h of continuous serving with no physical action and no captured trigger — [source](https://github.com/ml-explore/mlx/issues/3910)
- On 2x M4 Pro mini with macOS 26.5.1, SIGKILL of rank 1 left the survivor at 100% CPU in all_reduce with 1927 of 2009 samples in tbt_poll_cq — [source](https://github.com/ml-explore/mlx/issues/3910)
- On the same M4 Pro hardware a physical unplug crashed both ranks instead of hanging — [source](https://github.com/ml-explore/mlx/issues/3910)
- #4278: a SIGSTOPped rank froze the other three at the same iteration at 99-100% CPU, and SIGCONT let all four resume with no corruption — [source](https://github.com/ml-explore/mlx/issues/4278)
- Setting en2, en3 and en4 down with ping at 100% loss did not stop a four-rank RDMA collective, since established queue pairs bypass the IP stack — [source](https://github.com/ml-explore/mlx/issues/4278)
- mlx.launch kills the surviving ranks when any rank exits, which hides what survivors would do — [source](https://github.com/ml-explore/mlx/issues/4278)
- Ring backend on peer loss: v0.32.0 hangs, main aborts with exit 134, #3742 gives a catchable RuntimeError — [source](https://github.com/ml-explore/mlx/issues/4278)
- Ring issue #3862 is closed; its fix PR #4060 merged as commit 8c28c38 — [source](https://github.com/ml-explore/mlx/pull/4060)
- Ring SocketThread::worker returned after 10 failed attempts without resolving pending promises, so every rank blocked in Event::wait; the ring also lacks SO_KEEPALIVE — [source](https://github.com/ml-explore/mlx/issues/3862)
- PR #4060 chose to exit the process on peer loss because resolving the promises makes survivors return wrong collective results silently — [source](https://github.com/ml-explore/mlx/pull/4060)
- Throwing from the stream thread reaches std::terminate because StreamThread::thread_fn calls task() with no try/catch — [source](https://github.com/ml-explore/mlx/pull/4060)
- TN3205 states RDMA over Thunderbolt supports send and receive only, at most 10 UC queue pairs, messages up to 16,773,120 bytes and 4095 outstanding work requests — [source](https://developer.apple.com/documentation/technotes/tn3205-low-latency-communication-with-rdma-over-thunderbolt)
- The only integration pattern reported before any fix is an application no-progress watchdog that SIGKILLs ranks and reboots leaked nodes — [source](https://github.com/ml-explore/mlx/issues/3910)
