Uninterruptible Us+ rank processes in the Thunderbolt RDMA driver
Parent: Mac local LLMs: Clusters, RDMA, exo and ds4 · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
`Us+` ranks drift into the state gradually after their launcher dies, so a sweep at the next restart can still kill the newest orphans
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- `Us+` ranks drift into the state gradually after their launcher dies, so a sweep at the next restart can still kill the newest orphans [source]
- Orphaned ranks and their parent `sshd` both hang, so the orphan's ppid is the stuck `sshd`, not 1 [source]
- `launchd` KeepAlive restarts re-ran `mlx.launch` within seconds and accumulated peer orphans each cycle [source]
- The fix kills all processes matching the model's `--tmpdir` pattern with `kill -9` before launching, with no ppid filter [source]
- A PID-file lock released just before `exec` prevents a manual run and a launchd restart from killing each other's ranks, since an EXIT trap does not fire after `exec` [source]
- Killed `mlx_lm.share` ranks left temp directories of 817 GB and 868 GB on two nodes; kill processes before deleting them [source]
- Ranks already in `Us+` survive `kill -9` and only a node reboot clears them [source]
Children
- No children recorded.