<!-- llms-explorer concept facts · https://llms-explorer.com/tree/jaccl-rdma-completion-status-checking-and-silent/ · pack 2026-10-05 · ~706 tokens -->

# JACCL RDMA completion status checking and silent corruption

> zcbenz's closing reason: it is easier for maintainers to write and verify such changes themselves than to review contributor PRs over several rounds

Parent: [Mac local LLMs: Clusters, RDMA, exo and ds4](https://llms-explorer.com/tree/mac-local-llms-clusters-rdma-exo-ds4/) · 2 facets · 11 facts · page: https://llms-explorer.com/tree/jaccl-rdma-completion-status-checking-and-silent/

## Facts

- zcbenz's closing reason: it is easier for maintainers to write and verify such changes themselves than to review contributor PRs over several rounds — [source](https://github.com/ml-explore/mlx/pull/3152)
- angeloskath measured about 10% slowdown (about 1 us) on 16 KB 4-way all-reduce and held the PR for a JACCL refactor and optimisation pass, noting the team treats Thunderbolt as a reliable channel — [source](https://github.com/ml-explore/mlx/pull/3152)
- The rebased PR changed 3 files (+61 lines) and added 4 check sites in mesh_impl.h and 5 in ring_impl.h — [source](https://github.com/ml-explore/mlx/pull/3152)
- The first commit message cites IBV_WC_RETRY_EXC_ERR as an example of an RDMA transport error that was silently ignored — [source](https://github.com/ml-explore/mlx/pull/3152)
- The original PR also made Connection::poll() and the free poll() throw on a negative ibv_poll_cq return — [source](https://github.com/ml-explore/mlx/pull/3152)
- The PR author tested 256, 512, 5478+433 and 8490+512 token generations on the 612 GB model with zero RDMA errors after the fix, and saw no regression on models under 10 GB per rank — [source](https://github.com/ml-explore/mlx/pull/3152)
- The PR author states small models are unaffected because their RDMA operations always succeed, so the bug is latent — source: `asserted`
- After a peer is SIGKILLed, ibv_poll_cq returned 0 about 2.64 billion times, never negative, never a bad wc.status, so status checking cannot detect peer loss — [source](https://github.com/ml-explore/mlx/issues/3910)
- Related exo report: macOS 26.3.1 4x M3 Ultra clusters hit "[jaccl] Recv failed with errno=2", errno=60 and "Changing queue pair to RTR failed with errno 22"; the reporter says the #3451 race fix removed the post-kill RTR failures — [source](https://github.com/exo-explore/exo/issues/1847)
- Issue #3467 root cause: after the #3412 refactor Connection::info() only accepts IPv4-mapped GIDs (::ffff:x.x.x.x) but Apple Thunderbolt RDMA can expose only link-local fe80:: GIDs, leaving the GID uninitialised and giving RTR errno 22 — [source](https://github.com/ml-explore/mlx/issues/3467)

## Corrections and disagreements

- CONTRADICTS distributed-inference-across-macs.md: PR #3152 is closed, not open; zcbenz closed it without merging on 2026-10-02 — [source](https://github.com/ml-explore/mlx/pull/3152)
