<!-- llms-explorer concept facts · https://llms-explorer.com/tree/metal-gpu-watchdog-5-s-timeout-in-distributed-ml/ · pack 2026-10-05 · ~743 tokens -->

# Metal GPU watchdog 5 s timeout in distributed MLX generation

> #3830 was closed as completed by zcbenz on 2026-08-04 with the comment that MLX_METAL_FAST_SYNCH is not reliable and has no known fix

Parent: [Mac local LLMs: GPU stability and kernel panics](https://llms-explorer.com/tree/mac-local-llms-gpu-stability-and-kernel-panics/) · 2 facets · 10 facts · page: https://llms-explorer.com/tree/metal-gpu-watchdog-5-s-timeout-in-distributed-ml/

## Facts

- #3830 was closed as completed by zcbenz on 2026-08-04 with the comment that MLX_METAL_FAST_SYNCH is not reliable and has no known fix — [source](https://github.com/ml-explore/mlx/issues/3830)
- The measured kill points were 7,325, 7,391 and 7,280 tokens, all on mlx 0.31.2 — [source](https://github.com/ml-explore/mlx/issues/3830)
- mlx 0.32.0 added command-buffer error capture and event poisoning, so a watchdog kill may surface as a runtime error with the process alive; the 7.3k figures were not re-baselined on 0.32.0 — [source](https://github.com/ml-explore/mlx/issues/3830)
- A 2-rank TCP-ring pipeline serving GLM-4.7-4bit on M2 Ultra, macOS 26.6, mlx 0.32.0, flag unset, completed 1,584 chat requests over about 21 h with no wedge or watchdog kill; generation lengths were not instrumented — [source](https://github.com/ml-explore/mlx/issues/3830)
- Killing both wedged ranks released wired memory, but on the M2 Ultra every later mx.eval from a fresh process hung until reboot, and once the state escalated to a userspace freeze needing a power cycle — [source](https://github.com/ml-explore/mlx/issues/3830)
- The reporters' diagnosis is that the ~5 s watchdog counts waiting buffers as running, so the GPU is idle when the buffer dies — [source](https://github.com/ml-explore/mlx/issues/3830)
- A kidroca report found MLX_METAL_FAST_SYNCH was unnecessary for their workload, and removing it made a later stuck oMLX process recoverable by service restart without reboot — [source](https://github.com/ml-explore/mlx/issues/3830)
- A CPU-hot hang and a CPU-idle GPU-pegged hang are the two ends of the same fence lock, so a CPU-side wait timeout alone would not rescue the GPU — [source](https://github.com/ml-explore/mlx/issues/3830)
- Raw send/recv hit the Metal GPU Timeout on the receiver unless run on stream=mx.cpu, while mlx_lm.share, which does not set stream=mx.cpu, worked using all_sum as an init barrier — [source](https://github.com/ml-explore/mlx/issues/3207)

## Corrections and disagreements

- CONTRADICTS mlx-metal-fast-synch-and-jaccl-fence-wait-deadlock.md on the docs commit: the #3830 thread lists commit e031f07 ("Document MLX_METAL_FAST_SYNCH as unreliable instead of warning at run...") and 1c33dfd ("Warn once when a fast fence is created") as two commits added by zcbenz referencing the issue, so it is likely upstream, not only a fork commit; unresolved because the fetched upstream docs page still cites only #3142 — [source](https://github.com/ml-explore/mlx/issues/3830)
