<!-- llms-explorer concept facts · https://llms-explorer.com/tree/ibv-devinfo-hang-as-an-rdma-health-probe/ · pack 2026-10-05 · ~1236 tokens -->

# ibv_devinfo hang as an RDMA health probe

> The probe idea: run `ibv_devinfo` and treat a hang as a sign that Thunderbolt RDMA is wedged, then show a dashboard warning. It comes from one sentence in exo issue 1712.

Parent: [Mac local LLMs: Clusters, RDMA, exo and ds4](https://llms-explorer.com/tree/mac-local-llms-clusters-rdma-exo-ds4/) · 1 facets · 20 facts · page: https://llms-explorer.com/tree/ibv-devinfo-hang-as-an-rdma-health-probe/

## Facts

- The probe idea: run `ibv_devinfo` and treat a hang as a sign that Thunderbolt RDMA is wedged, then show a dashboard warning. It comes from one sentence in exo issue 1712. — [source](https://github.com/exo-explore/exo/issues/1712)
- The issue says jaccl errors such as exo issue 1711 recover only by unplugging and replugging the Thunderbolt 5 cable or rebooting, and that "usually when these errors happen, the `ibv_devinfo` command hangs". It asks for a dashboard warning with advice to reconnect cables, then reboot. — [source](https://github.com/exo-explore/exo/issues/1712)
- The issue has one participant, no comments, no assignee, no linked branch or pull request, and the label `enhancement`. It gives no logs, no macOS version and no reproduction. — [source](https://github.com/exo-explore/exo/issues/1712)
- ds4's docs use `ibv_devinfo -v` the other way round: as a static check that the port is active and has an IPv4-mapped GID, and as the place to read `--rdma-device` and `--rdma-gid-index` values. They do not mention hangs. — [source](https://raw.githubusercontent.com/antirez/ds4/main/docs/DISTRIBUTED.md)
- Exo's reliability PRs 2367 and 2370 took other routes to the same goal. PR 2367 stops a runner when its supervisor sees MLX's ring diagnostic "Too many send/recv errors. Aborting..." on stderr. PR 2370 stops the first-rank runner after 90 seconds without a token or prefill progress report while a generation is in flight. Neither uses `ibv_devinfo`. — [source](https://github.com/exo-explore/exo/pull/2370)
- Both PRs target a stuck runner process. A wedged driver that hangs `ibv_devinfo` is a different layer, and a runner blocked inside JACCL would already be caught only by the 90-second rule. — source: `asserted`
- A hung probe may not be killable: processes stuck in the Thunderbolt RDMA stack have been seen in uninterruptible sleep (`Us+`) where `kill -9` does nothing. A probe run without a timeout, or with one that only sends a signal, can leave stuck probe processes behind. — [source](https://valensas.com/blog/how-we-fixed-a-silent-killer-in-our-mac-studio-llm-cluster)
- `ibv_devinfo` reports `PORT_ACTIVE` only when all connected Macs have RDMA enabled (per distributed-inference-across-macs.md), so a clean exit with `PORT_DOWN` and a hang are different signals. — source: `asserted`
- Probe versus progress watchdog. Issue 1712 proposes a link-layer probe; the merged-or-open reliability PRs use runner-level signals (stderr diagnostic, token silence) and recreate the runner. The sources do not compare them. — [source](https://github.com/exo-explore/exo/pull/2370)
- Whether `ibv_devinfo` hangs on macOS 26.5 and later, or on macOS 27's reworked library. No fetched source tests it after March 2026. — source: `asserted`
- A safe probe form: a short-deadline child process that is expected to be abandoned if it hangs, plus a count of stuck probes. — source: `asserted`
- exo issue 1712 is still open with no assignee, comment, branch or pull request. — [source](https://github.com/exo-explore/exo/issues/1712)
- Exo issue 1712's only evidence is one sentence from the author that `ibv_devinfo` usually hangs when jaccl errors occur; it has no log or repro. — [source](https://github.com/exo-explore/exo/issues/1712)
- The issue asks for a dashboard warning that suggests reconnecting cables, then rebooting. — [source](https://github.com/exo-explore/exo/issues/1712)
- ds4 documents `ibv_devinfo -v` as a static link check, not as a hang probe. — [source](https://raw.githubusercontent.com/antirez/ds4/main/docs/DISTRIBUTED.md)
- exo PR 2367 detects a failed MLX ring by the stderr text "Too many send/recv errors. Aborting..." and stops that runner. — [source](https://github.com/exo-explore/exo/pull/2367)
- exo PR 2370 stops a first-rank runner after 90 seconds of silence during a generation request, and checks every 5 seconds. — [source](https://github.com/exo-explore/exo/pull/2370)
- Neither PR 2367 nor PR 2370 uses `ibv_devinfo`. — [source](https://github.com/exo-explore/exo/pull/2370)
- Processes stuck in the Thunderbolt RDMA stack have been reported in uninterruptible sleep where `kill -9` has no effect, so a hanging probe needs an abandon-on-timeout design. — [source](https://valensas.com/blog/how-we-fixed-a-silent-killer-in-our-mac-studio-llm-cluster)
- `ibv_devinfo` shows `PORT_ACTIVE` only when every connected Mac has RDMA enabled, per the existing dossier distributed-inference-across-macs.md. — source: `asserted`
