MLX command-buffer error capture and event poisoning in 0.32.0
Parent: Mac local LLMs: GPU stability and kernel panics · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
The old failure signature was `libc++abi: terminating due to uncaught exception of type std::runtime_error: [METAL] Command buffer execution failed: ...` with the stack `mlx::core::gpu::check_error(MTL::CommandBuffer*)` called from a completion block in `-[_MTLCommandBuffer didCompleteWithStartTi...
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- The old failure signature was `libc++abi: terminating due to uncaught exception of type std::runtime_error: [METAL] Command buffer execution failed: ...` with the stack `mlx::core::gpu::check_error(MTL::CommandBuffer*)` called from a completion block in `-[_MTLCommandBuffer didCompleteWithStartTime:endTime:error:]`, ending in `__cxa_throw`, `_objc_terminate` and `abort`. [source]
- Field cases before the fix: issue 3317 (M2 Ultra, Qwen3.5-122B, about 3.5 hours of sustained inference) and an M4 Max running a four-server mlx-lm 0.31.3 pool where prompt-cache growth raised `kIOGPUCommandBufferCallbackErrorOutOfMemory` and aborted one pool member. [source]
- PR 3523 first handled the error "in same thread", then on 2026-05-20 was rewritten as "Catch error in CommandBuffer and poison the events": the error is set on every `Event` the failed buffer signals, and when a poisoned `Event` is waited on, the error passes to the waiting `CommandEncoder` and to the events that encoder signals, so an error in one stream reaches other streams through events. [source]
- At the time of 3523 the CPU stream was excluded: its errors "currently just throw and crash"; that gap is what PR 3742 closed in 0.32.1. [source]
- In current main `CommandEncoder::get_command_encoder()` calls `error_.check()` before it opens a compute encoder, which is the "before encoding" check point named in the PR description. [source]
- The throw is delayed to the next check point, so the exception does not point at the failing op; the author likened it to CUDA. [source]
- PR 3318 (defer the error to the next user-thread `eval()`) ran unattended for 5 days on Qwen3.5-122B and overnight on an M3 Ultra, then was rejected and closed by its author. zcbenz on issue 3317: "We can not simply transfer the error: MLX code is not exception-safe so when the exception is thrown the process would be in a stale state that will likely fail later in a much weirder way." awni on issue 2670: no guarantee the state is sane after an exception during eval. [source]
- PR 3519 (2026-05-10, closed 2026-05-11, +148/-6 over two files) answered the state-safety objection with per-stream poisoning: the handler never throws, it saves the message in a `StreamThread` slot; the next `eval`, `finalize` or `synchronize` on that stream re-throws once and then refuses every call with "[METAL] Stream is in error state from a prior failure. Call mx.clear_streams() ..."; `mx.clear_streams()` resets the slots. Maintainers did not take it; zcbenz opened 3523 the next day with event poisoning instead. [source]
- PR 3523 reviews: angeloskath on 2026-05-19 noted `check_error` in `Event` and `Fence` wait was ineffective because the waiting thread is practically never the scheduling thread; 2026-05-21 he called the event-based rewrite "an awesome cleanup", especially the fence cleanup. [source]
- 2026-06-12: approved by angeloskath and merged by zcbenz as commit a025496 with 16 checks passed; the issue it closes is 2670 ("when it crashed in the background, can not catch the exception", opened 2025-09-29). [source]
- 2026-06-04: tseylerd ported 3523 to mlx v0.31.1 so mlx-swift and mlx-swift-lm could use it (tseylerd/mlx pull 1). [source]
- 2026-06-13 (opened; listed in the 0.32.1 notes): PR 3675 "Fix state corruption when a primitive throws during eval" repaired the other half of exception safety: an exception out of `eval_cpu` or `eval_gpu` unwound `eval_impl` before its per-stream epilogue, so earlier arrays in the same batch were marked evaluated while their kernels sat in a command buffer never committed (later reads returned unwritten memory) and events attached to the failed eval were never signaled (later reads hung in `array::wait`). The fix signals every event created for the eval, synchronizes every touched stream and re-throws. [source]
- The 0.32.1 release notes list "Propagate CPU errors to events" (PR 3742) and "Fix state corruption when a primitive throws during eval" (PR 3675); this confirms that 3742 shipped in 0.32.1. [source]
- PR 3675's trigger is extension code: `mx.fast.metal_kernel` whose source fails to JIT-compile (it compiles lazily at eval) or a custom primitive that throws in `eval_gpu`; built-in ops validate at graph construction, so pure built-in code rarely reaches it; before the fix a caught exception left a bystander array unwritten in about half of runs. [source]
- A caught Metal error says nothing about whether arrays and KV state produced by the discarded command buffers are valid; PR 3523 and PR 3675 only keep the process alive and the streams consistent. [source]
- The 3523 author could not unit-test it ("no good way to test it") and used `mlx.launch --verbose -n 32 python python/tests/ring_test_distributed.py`, which reliably timed out before the change. [source]
- PR 3519's author had the same testing limit: the pre-flight allocation check in `metal::malloc` intercepts most synthetic out-of-memory repros, so a real command-buffer failure is needed. [source]
- Design: PR 3519 (per-stream "poisoned" state, explicit reset with `mx.clear_streams()`) against PR 3523 (error rides on events, no explicit reset call); the maintainers merged the event design. [source]
- Overhead: the 3523 author argued event waits add overhead while fewer completion handlers remove some; his benchmark could not separate the two (see Claims). [source]
- Whether PR 3675's recovery covers errors that arrive asynchronously from the command-buffer handler (it handles exceptions thrown inside eval, a different path); no source tests the combination. [source]
- Which release first carried a025496: the 0.32.0 notes were not re-read in this batch. [source]
- PR 3523 was approved by angeloskath on 2026-06-12 and merged as commit a025496 with 16 checks passed. [source]
- Before PR 3523 a failed Metal command buffer aborted the process through an exception thrown in a completion handler (`check_error` under `didCompleteWithStartTime:endTime:error:`). [source]
- Issue 3317 reported that abort on an M2 Ultra after about 3.5 hours of Qwen3.5-122B inference. [source]
- PR 3318 (defer the error to the next eval) was rejected by maintainers on state-safety grounds ("MLX code is not exception-safe"). [source]
- PR 3519 proposed per-stream poisoning with `mx.clear_streams()` as the reset and was closed on 2026-05-11. [source]
- PR 3523 was retitled on 2026-05-20 to "Catch error in CommandBuffer and poison the events" and propagates the error from the command buffer to its signaled events and from a waited event to the encoder. [source]
- At 3523's merge the CPU stream still threw and crashed on errors. [source]
- PR 3523's author tested with `mlx.launch -n 32 python python/tests/ring_test_distributed.py`, which reliably timed out before the change. [source]
- A Llama-3.1-8B-Instruct benchmark (`mlx_lm.benchmark -n 10 -p 64 -g 512`) gave main 363.322 prompt and 24.238 generation tokens per second and the 3523 branch 360.728 and 24.239, with peak memory 16.176 GB on both. [source]
- tseylerd ported 3523 to mlx v0.31.1 for mlx-swift users on 2026-06-04. [source]
- `CommandEncoder::get_command_encoder()` calls `error_.check()` before opening a compute encoder. [source]
- PR 3675 fixes unwritten-buffer reads and hung `array::wait` calls after an exception escapes `eval_gpu` or `eval_cpu`, by signaling events, synchronizing streams and re-throwing. [source]
- The 0.32.1 release notes list PR 3742 and PR 3675. [source]
- PR 2983 "Fix scheduler exception propagation to Python" was an earlier attempt, closed on 2026-01-19. [source]
- PR 4356 stopped a failed CUDA graph commit from poisoning the CUDA encoder (merged 2026-08-23), the CUDA counterpart of the poison design. [source]
- A caught Metal error does not certify the contents of arrays from discarded command buffers. [source]
Children
- No children recorded.