M5 Max watchdogd 90 s kernel panic under sustained GPU load
Parent: Mac local LLMs: GPU stability and kernel panics · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
The top backtrace frames are the watchdog itself (`AppleARMWatchdogTimer`, `AppleInterruptControllerV3`), so the panic report names the detector, not the originating fault.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- The top backtrace frames are the watchdog itself (`AppleARMWatchdogTimer`, `AppleInterruptControllerV3`), so the panic report names the detector, not the originating fault. [source]
- The reporter's reading: active-display compositing (WindowServer, the busiest thread in the panicked task) contends with long training command buffers until the whole system stalls. Memory (compressor 13%) and heat (thermal pressure Nominal) were ruled out. [source]
- Process isolation cannot help, because the stall is below the process boundary. [source]
- Only one public reporter exists for this signature on M5 Max; no corroboration from a second machine was found. [source]
- Runs 1-3 (before any mitigation): three hard reboots on mlx and mlx-metal 0.30.6 during sustained FLUX.2 Klein LoRA training, each at a similar depth into the run. [source]
- The reporter then pinned mlx and mlx-metal to 0.31.2, set `AGX_RELAX_CDM_CTXSTORE_TIMEOUT=1` before Metal init, and used `caffeinate -s`. A 9B bf16 run with segmentation off still rebooted about 11 min in, as the first training step began (93 s variant). [source]
- Telemetry on that run: `powermetrics` at 5 s cadence went silent at 10:12:19 and the panic fired at 10:15:18 (2 min 59 s of silence), GPU at about 1592 MHz, 97% active, 24.4 W. [source]
- A/B with the same config and only the display changed: display active reboots at step 0; `pmset displaysleepnow` (lid open) reached step 302 of 600 in a box up 5 h 20 min. All on 2026-06-27. [source]
- 2026-07-19: ltx-2-mlx added a "GPU watchdog" explanation to its error path and decided never to set driver env vars itself. [source]
- The failing configuration is the larger one: 9B bf16 base, `quantize: null`, `low_ram: true`, 768 px, batch size 1, 25-image dataset, 600 steps (mflux 0.17.5, mlx-lm 0.29.1, Python 3.14). Smaller models were not reported to panic, but no smaller-model run was reported either. [source]
- The reporter's untested alternative: chunk the work so no sustained-GPU stretch is long enough to starve the scheduler. [source]
- Other M5 Max (applegpu_g17s) GPU failures, none of them a watchdogd panic: mlx-lm 1206, LoRA on Qwen3.5-9B-4bit crashes at the first backward pass with `kIOGPUCommandBufferCallbackErrorOutOfMemory` at low system memory use, while Qwen3-8B-4bit trains; and mlx-vlm 1064, Qwen3-VL-2B gives GPU Hang and PageFault on an M5 Max. [source]
- MetalGuard hypothesises, from mlx-lm 1206, that `applegpu_g17s` has command-buffer limits independent of RAM, and filters its registry by GPU family; this is unverified. [source]
- Hybrid-architecture memory (a commenter on mlx-lm 1206 blames Qwen3.5 recurrent-state backprop) versus an M5-Max-specific command-buffer limit (another commenter, MetalGuard). Neither tested. [source]
- Does macOS 26.6 or 27 change the M5 Max behaviour? No report found. [source]
- Is the panic specific to M5-class silicon, or only to the heaviest config? Only one machine reported. [source]
- Does a headless virtual display (no panel) count as "display off" for this stall? Untested. [source]
- The M5 Max in the report is `Mac17,7`, SoC `T6050`, 128 GB, macOS 26.5.1 build `25F80`, kernel `Darwin 25.5.0 xnu-12377.121.6~2/RELEASE_ARM64_T6050`. [source]
- Three hard reboots occurred on mlx and mlx-metal 0.30.6 with mflux 0.17.5 and mlx-lm 0.29.1 during sustained FLUX.2 Klein LoRA training, each at a fairly consistent depth into the run. [source]
- The failing workload was a 25-image dataset at 768 px, batch size 1, `quantize: null` (bf16 base), `low_ram: true`, 600 steps. [source]
- After pinning mlx to 0.31.2 and setting `AGX_RELAX_CDM_CTXSTORE_TIMEOUT=1`, a 9B bf16 segmentation-off run still rebooted about 11 minutes in with the 93 s watchdogd panic. [source]
- That run's `powermetrics` log went silent at 10:12:19 and the panic fired at 10:15:18. [source]
- At the panic, compressor use was 13%, swap was fine and WindowServer was the busiest thread in the panicked task. [source]
- The display-asleep run (`pmset displaysleepnow`, lid open) used the same config and the same env var and reached step 302 of 600; all three runs are dated 2026-06-27. [source]
- The reporter calls `AGX_RELAX_CDM_CTXSTORE_TIMEOUT=1` plus display asleep "a reliable combination" on this machine. [source]
- The reporter proposes, without testing it, chunking work so no sustained-GPU stretch is long enough to starve the scheduler. [source]
- ltx-2-mlx added a GPU-watchdog explanation to its error path on 2026-07-19 and decided never to set driver env vars itself. [source]
- mlx-lm issue 1206 (opened 2026-04-26) reports LoRA on `mlx-community/Qwen3.5-9B-4bit` crashing at the first backward pass on an M5 Max (applegpu_g17s, 36 GB, mlx 0.31.2) with `kIOGPUCommandBufferCallbackErrorOutOfMemory` at low memory use. [source]
- On the same machine `mlx-community/Qwen3-8B-4bit` trains with identical settings, and batch size, sequence length 2048-8192, LoRA layers 4-16 and `--grad-checkpoint` do not change the crash. [source]
- A commenter on mlx-lm 1206 attributes the crash to hybrid linear-attention state storage during backprop and another suggests an M5 Max command-buffer limit independent of RAM; neither is tested. [source]
- mlx-vlm issue 1064 (opened 2026-04-24, M5 Max, mlx 0.32.0.dev20260424) reports `Qwen/Qwen3-VL-2B-Instruct` ending in `kIOGPUCommandBufferCallbackErrorHang` and `Qwen3-VL-2B-Thinking-bf16` in a PageFault; a maintainer closed it as completed on 2026-06-16 saying it "should be fixed on main". [source]
- MetalGuard maps `applegpu_g13`, `g14`, `g15`, `g16` and `g17` prefixes to M1 through M5 and registers Qwen3-VL-2B (M5 Max hang) and Qwen3.5-9B-4bit (M5 Max LoRA first-backward OOM) as known-bad on M5 Max. [source]
- No second public report of the `watchdogd` 90 s panic on any chip was found. [source]
- A `watchdogd` timeout means the kernel could not schedule the check-in thread for about 90 s, so the whole system stalled rather than one process. [source]
Children
- No children recorded.