<!-- llms-explorer concept facts · https://llms-explorer.com/tree/mlx-metal-fast-synch-effect-on-event-synchroniza/ · pack 2026-10-05 · ~803 tokens -->

# MLX_METAL_FAST_SYNCH effect on event synchronization hangs

> Already covered: with the flag set MLX replaces `MTLSharedEvent` waits with a spin on a shared word; the existing dossier lists how that can hang.

Parent: [Mac local LLMs: GPU stability and kernel panics](https://llms-explorer.com/tree/mac-local-llms-gpu-stability-and-kernel-panics/) · 1 facets · 13 facts · page: https://llms-explorer.com/tree/mlx-metal-fast-synch-effect-on-event-synchroniza/

## Facts

- Already covered: with the flag set MLX replaces `MTLSharedEvent` waits with a spin on a shared word; the existing dossier lists how that can hang. — source: `asserted`
- Absent from the existing dossier: in issue 3830's stack samples of a wedged TCP-ring pipeline run, the main thread is parked in `Event::wait` on an `IOSurfaceSharedEvent`, both socket worker threads are idle, and a CPU stream thread burns 100% in the fast-fence spin, so the network layer is not stalled. — [source](https://github.com/ml-explore/mlx/issues/3830)
- The wedge in that harness depends on handoff shape and size: 10 KB messages ran 2,000,000 iterations clean with the flag on and 2.6 MB messages deadlocked within seconds. — [source](https://github.com/ml-explore/mlx/issues/3830)
- No new dated events beyond the existing dossier. — source: `asserted`
- The cross-stream event patch is cheap on a single process: 8 KB handoffs cost 1.9 ms per iteration with flag on and 2.3 ms with flag off, and 46.8 MB handoffs 9.8 ms (flag on) and 7.2 ms (flag off), on one M3 Ultra. — [source](https://github.com/rltakashige/mlx-jaccl-fix-small-recv/pull/4)
- Before the patch, the same stress with the flag on never completed at 46.8 MB (killed at 150 s) and ran 12 ms per iteration with a 361 ms worst case at 8 KB. — [source](https://github.com/rltakashige/mlx-jaccl-fix-small-recv/pull/4)
- Whether small messages avoid the hang: the 3830 author's 10 KB run was clean; another commenter's loop wedged at about 20 KB once a scalar `all_sum` was in it. The existing dossier does not record this split. — [source](https://github.com/ml-explore/mlx/issues/3830)
- Whether the 0.32.3 fence fix (PR 4552) changes the small-message behaviour reported in 3830. — source: `asserted`
- In 3830's wedge, socket worker threads were idle while a CPU stream thread spun, so the hang is in the fence, not the network. — [source](https://github.com/ml-explore/mlx/issues/3830)
- With the flag on, 10 KB messages were clean for 2,000,000 iterations and 2.6 MB messages deadlocked within seconds in the same harness. — [source](https://github.com/ml-explore/mlx/issues/3830)
- A second harness wedged at about 20 KB once a per-step scalar `all_sum` was added, so size alone does not predict safety. — [source](https://github.com/ml-explore/mlx/issues/3830)
- On one M3 Ultra the cross-stream event patch costs 9.8 ms versus 7.2 ms (flag on versus off) at 46.8 MB and 1.9 ms versus 2.3 ms at 8 KB. — [source](https://github.com/rltakashige/mlx-jaccl-fix-small-recv/pull/4)
- Without the patch and with the flag on, 46.8 MB cross-stream handoffs never completed in 150 s and 8 KB ones had a 361 ms worst case. — [source](https://github.com/rltakashige/mlx-jaccl-fix-small-recv/pull/4)
