Metal GPU watchdog 5 s timeout in distributed MLX generation
Parent: Mac local LLMs: GPU stability and kernel panics · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
#3830 was closed as completed by zcbenz on 2026-08-04 with the comment that MLX_METAL_FAST_SYNCH is not reliable and has no known fix
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- #3830 was closed as completed by zcbenz on 2026-08-04 with the comment that MLX_METAL_FAST_SYNCH is not reliable and has no known fix [source]
- The measured kill points were 7,325, 7,391 and 7,280 tokens, all on mlx 0.31.2 [source]
- mlx 0.32.0 added command-buffer error capture and event poisoning, so a watchdog kill may surface as a runtime error with the process alive; the 7.3k figures were not re-baselined on 0.32.0 [source]
- A 2-rank TCP-ring pipeline serving GLM-4.7-4bit on M2 Ultra, macOS 26.6, mlx 0.32.0, flag unset, completed 1,584 chat requests over about 21 h with no wedge or watchdog kill; generation lengths were not instrumented [source]
- Killing both wedged ranks released wired memory, but on the M2 Ultra every later mx.eval from a fresh process hung until reboot, and once the state escalated to a userspace freeze needing a power cycle [source]
- The reporters' diagnosis is that the ~5 s watchdog counts waiting buffers as running, so the GPU is idle when the buffer dies [source]
- A kidroca report found MLX_METAL_FAST_SYNCH was unnecessary for their workload, and removing it made a later stuck oMLX process recoverable by service restart without reboot [source]
- A CPU-hot hang and a CPU-idle GPU-pegged hang are the two ends of the same fence lock, so a CPU-side wait timeout alone would not rescue the GPU [source]
- Raw send/recv hit the Metal GPU Timeout on the receiver unless run on stream=mx.cpu, while mlx_lm.share, which does not set stream=mx.cpu, worked using all_sum as an init barrier [source]
Corrections and disagreements
- CONTRADICTS mlx-metal-fast-synch-and-jaccl-fence-wait-deadlock.md on the docs commit: the #3830 thread lists commit e031f07 ("Document MLX_METAL_FAST_SYNCH as unreliable instead of warning at run...") and 1c33dfd ("Warn once when a fast fence is created") as two commits added by zcbenz referencing the issue, so it is likely upstream, not only a fork commit; unresolved because the fetched upstream docs page still cites only #3142 [source]
Children
- No children recorded.