MLX PR 4552 Metal fence deadlock fix in 0.32.3 (MLX_METAL_FAST_SYNCH)
Parent: Mac local LLMs: GPU stability and kernel panics · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Before the PR, `eval_impl` kept one fence per stream with a counter of `fence::update()` calls scheduled by the main thread; the consumer stream waited for a value equal to that scheduled total, regardless of how many updates had finished. With independent reductions in one tape (matmul, matmul, ...
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Before the PR, `eval_impl` kept one fence per stream with a counter of `fence::update()` calls scheduled by the main thread; the consumer stream waited for a value equal to that scheduled total, regardless of how many updates had finished. With independent reductions in one tape (matmul, matmul, update, update, wait, wait on the GPU; spin until counter 2 then all_reduce, all_reduce on the CPU) the CPU would not start the first all_reduce before the GPU finished the inputs of every all_reduce in the tape. [source]
- The patch makes `Fence::update` return the new counter value and `Fence::wait` take that value (`wait(stream, x, value)`); `transforms.cpp` replaces its `needs_fence` pair with a `FenceInfo{stream_index, cross_device, value}` record, stores the value returned by `update` and passes it to the consumer's `wait`. [source]
- The header comment of `fence.h` changed from "wait waits until all previous calls to update are complete" to "`update` returns a value that marks when its array is computed and visible. `wait` orders work in the consumer stream after that value is signaled. Later updates do not extend an earlier array's wait." [source]
- The change is not limited to the fast path: the default `MTLSharedEvent` path now does `event.set_value(value); event.wait(stream)` instead of waiting on the latest count, and the CUDA backend keeps one GPU event per update (`gpu_events`) so a consumer can wait for an earlier update. [source]
- Main's `fence.cpp` after the change still spins in the same two places (`while (f.cpu_value()[0] < value)` on a CPU stream, the `fence_wait` kernel on the GPU), so the fix lives in which value is awaited, not in how the spin reads memory. [source]
- The PR touches five files (CUDA, Metal and no-GPU `fence.cpp`, `fence.h`, `transforms.cpp`), +51 lines and -41 lines. [source]
- 2026-09-22: first commit "first attempt: Fence::update() marks the input"; 2026-09-23 merge of main; 2026-09-24 suggestions from zcbenz applied to the Metal and CUDA fence files and the title changed from "[BUG] Deadlock in fence" to "[BUG][Metal] Deadlock in fence". [source]
- 2026-09-24: zcbenz approved ("Very good job locating the deadlock!") and nastya236 merged it as commit 77e1cfb23d3cc94cadf3e821e85aa8eed363ade2 with 29 checks passed. [source]
- The 0.32.3 release notes list it as "[BUG][Metal] Deadlock in fence" (PR 4552); the tag commit is 64ea011 dated 2026-09-28 and the release page is stamped 29 Sep 00:37. [source]
- 2026-09-30: Unsloth's MLX pins moved to mlx 0.32.3 (unsloth pull 12334, unsloth-zoo pull 1516); 2026-10-01: exo PR 2377 "use MLX fast synch only for RDMA instances" and exo PR 2384 "build MLX without null-this traps (needed for MLX >= 0.32.1)" were open. [source]
- Field signature of this deadlock, from the PR: the stalled node's `ps -Ao pid,etime,%cpu,command` shows 100% CPU (one stalled system thread, idle JACCL ring threads) while healthy nodes show about 300% (system thread plus two JACCL pool threads); `sample <pid> 3 -mayDie -f` on the stalled node shows the `Fence::wait(...)::$_1` lambda on the system CPU thread and idle JACCL threads, meaning the node never entered `all_reduce`, while healthy nodes sit in the ring pass. [source]
- Counter evidence via `lldb -p <pid>` on the stalled thread: the main thread had scheduled 12 `fence::update()` calls, the GPU counter was 8 and the CPU counter 0, so no all_reduce had started. [source]
- The author states the counters cannot tell whether GPU updates 9 to 12 never executed or whether the CPU could not see their writes, so the fix removes the scheduling error but does not settle the memory-visibility question. [source]
- The reproducer is a loop of 32 independent `x @ weight` chains (dim 4096, 8 rounds, float32) followed by 32 `mx.distributed.all_sum` calls and a second GPU chain per step, 100 steps, `mx.distributed.init(strict=True, backend="jaccl")`; the author ran it on four M5 Ultras joined in a ring with three cables per neighbour. [source]
- The PR lists no issue it closes ("Successfully merging this pull request may close these issues: None yet"), so issues 3142 and 3830 are cited in its text but not auto-closed. [source]
- Workloads with few independent reductions per tape (single sequential all_sum per layer) would not have hit this particular defect, which fits reports that FAST_SYNCH mostly worked for tensor-parallel decode; failure classes 1 to 5 in the sibling dossier remain possible. [source]
- Whether the flag is now safe: the PR makes no such claim, the environment-variables page still gives only "default is 0", and exo's open PR 2377 restricts the flag to RDMA instances rather than enabling it everywhere. [source]
- Whether the 7,300-token watchdog kill with the flag unset (issue 3830) changes now that the `MTLSharedEvent` path waits for a specific value; no source re-tests it. [source]
- Whether failure classes 1 to 4 reproduce on 0.32.3 with the flag on; the PR's own validation covers only its reproducer. [source]
- Whether upstream will update the distributed docs wording ("not reliable ... see #3142") after this fix. [source]
- PR 4552 was merged by nastya236 on 2026-09-24 as commit 77e1cfb23d3cc94cadf3e821e85aa8eed363ade2 with 29 checks passed, after zcbenz approved it. [source]
- zcbenz's review comment was "Very good job locating the deadlock!". [source]
- The 0.32.3 release notes list PR 4552, and the release tag commit is 64ea011 dated 2026-09-28. [source]
- Before PR 4552 each consumer wait targeted the total number of updates the main thread had scheduled on that fence. [source]
- `Fence::update` now returns the counter value and `Fence::wait` takes it, and `transforms.cpp` stores it per array in a `FenceInfo` record. [source]
- The `fence.h` contract now says later updates do not extend an earlier array's wait. [source]
- The MTLSharedEvent path of `Fence::wait` now sets the awaited value before waiting. [source]
- The CUDA fence keeps a vector of per-update GPU events after PR 4552. [source]
- PR 4552 changes five files for +51 and -41 lines. [source]
- Main's Metal fence still spins on `cpu_value()[0] < value` on a CPU stream and on a one-thread `fence_wait` kernel on the GPU. [source]
- A stalled rank showed 100% CPU against about 300% on healthy ranks, with `Fence::wait` on the system thread in `sample` output. [source]
- In the author's lldb capture 12 updates were scheduled, the GPU had completed 8 and the CPU counter was 0. [source]
- The author could not tell from the counters whether updates 9 to 12 never ran or were not visible to the CPU. [source]
- The PR's reproducer uses 32 independent GPU chains and 32 `all_sum` calls per step over JACCL, 100 steps, on four M5 Ultras in a ring. [source]
- The PR does not auto-close issues 3142 or 3830. [source]
- Unsloth Studio moved its MLX pin to 0.32.3 on 2026-09-30, and exo had open PRs 2377 (fast synch only for RDMA) and 2384 (null-this traps, needed for MLX 0.32.1 and later) on 2026-10-01. [source]
- The environment-variables page documents `MLX_METAL_FAST_SYNCH` as default 0 and requiring Metal 3.2 or later, with no reliability wording. [source]
- PR 4552 shows one cause of FAST_SYNCH deadlocks was scheduling, not memory coherence, which narrows but does not remove the "cannot be guaranteed" position. [source]
Corrections and disagreements
- CONTRADICTS: mlx-metal-fast-synch-and-jaccl-fence-wait-deadlock.md reads the hangs as memory-visibility and cross-queue ordering defects (classes 1 to 5) and zcbenz's April wontfix as "cannot be guaranteed"; PR 4552's author began from the coherence hypothesis ("The hypothesis was that stall happens because memory coherence is not guaranteed") and found a scheduling defect fixable without any memory-model change, and zcbenz approved the fix. [source]
Children
- No children recorded.