<!-- llms-explorer concept facts · https://llms-explorer.com/tree/gpu-firmware-lockup-reports-gpuevent-ips/ · pack 2026-10-05 · ~1869 tokens -->

# GPU firmware lockup reports gpuEvent ips

> The report is named after the process, for example `gpuEvent-python3.13-*.ips`, and lands in `/Library/Logs/DiagnosticReports`.

Parent: [Mac local LLMs: GPU stability and kernel panics](https://llms-explorer.com/tree/mac-local-llms-gpu-stability-and-kernel-panics/) · 1 facets · 34 facts · page: https://llms-explorer.com/tree/gpu-firmware-lockup-reports-gpuevent-ips/

## Facts

- The report is named after the process, for example `gpuEvent-python3.13-*.ips`, and lands in `/Library/Logs/DiagnosticReports`. — source: `asserted`
- The `analysis` object holds the firmware's view: restart reason, signature, guilty data master, and per-queue firmware state slots. In eight reports from one host every field except `command_buffer_trace_id` was identical. — source: `asserted`
- Firmware slot state in all eight reports: compute list slots 0-2 each 1, 3D slots 0-2 each 0, tile-accelerator slots 0-1 each 0. Read with the reporter's inference, a compute buffer was stuck and nothing graphical was running. The meaning of the `fw_*_state` keys is not documented in any source read. — source: `asserted`
- Apple's AGX driver separates "firmware-detected lockup" from its timeout restart reasons ("timestamp timeout", "progress timeout", "CDM Kill timeout"). The Metal error code for these events is Hang (0x03), not Timeout (0x02). — source: `asserted`
- Background from Asahi Linux reverse engineering: AGX is driven through an ASC coprocessor running Apple firmware; work is split into queue types TA (vertex), 3D (pixel) and CP (compute); a context can own several work queues; submissions go through shared-memory channel rings. This is the structure the `fw_cl/3d/ta` slots most plausibly mirror (inference). — source: `asserted`
- 2026-08-27: first two oMLX events on macOS 26.5.1 (25F80). No report survives, so their firmware signature is unknown. — source: `asserted`
- 2026-09-28 and 2026-10-02: eight reports (one and seven) on macOS 27.0 (26A428), all with the same `analysis` fields. — source: `asserted`
- 2026-10-03: a second reporter on a different M3 Ultra 512 GB confirms that on macOS 27 each Hang leaves a gpuEvent report with `restart_reason 4` and `signature 579`. — source: `asserted`
- Reports undercount recoveries: AGX `recoveryCount` read 10 against 8 reports since the 09-15 boot, and the two missing recoveries left no oMLX log line either. Count both sources. — source: `asserted`
- Two of ten events fall on macOS 26.5.1 with no surviving report, so "every gpuEvent report has signature 579" holds only for macOS 27 reports. — source: `asserted`
- The Python frame in the traceback is the first encode or wait on the engine's stream after the reset. It names the engine whose stream saw the error, not the guilty kernel; at 11:06 on 10-02 the hung queue's owner logged nothing. — source: `asserted`
- A survivable restart is not the same as safe state: other processes' buffers on the same queues are also discarded as `InnocentVictim` (0x05), and in the oMLX 3706 variant repeated errors tip the process into `SubmissionsIgnored` (0x04). — source: `asserted`
- Error-code vocabulary seen across sources: OutOfMemory 0x08 "Insufficient Memory"; PageFault 0x0b "Caused GPU Address Fault Error"; SubmissionsIgnored 0x04 "Ignored (for causing prior/excessive GPU errors)". — source: `asserted`
- Whether the lockup is a firmware scheduling failure between two busy queues (oMLX 4224 reporter, inferred, firmware criterion undocumented) or a single long dispatch exceeding a firmware progress check (reporter says not excluded). Both stand. — source: `asserted`
- What `signature 579`, `guilty_dm 3` and the `fw_*_state` slots encode; Apple publishes no schema in anything found. — source: `asserted`
- Whether `restart_reason 4` appears for non-MLX Metal workloads on macOS 27 (no source found). — source: `asserted`
- Whether any `gpuEvent` report accompanies the M5 Max `watchdogd` panic (the reporter attached none). — source: `asserted`
- macOS names GPU restart reports `gpuEvent-<process>-<timestamp>.ips`, for example `gpuEvent-python3.13-*.ips`, under `/Library/Logs/DiagnosticReports`. — [source](https://github.com/jundot/omlx/issues/4224)
- Eight gpuEvent reports from one M3 Ultra host (one on 09-28, seven on 10-02) carry identical `analysis` fields apart from `command_buffer_trace_id`. — [source](https://github.com/jundot/omlx/issues/4224)
- The `analysis` object includes `fw_cl_state` slots 0-2 each equal to 1, `fw_3d_state` slots 0-2 each 0, `fw_ta_state` slots 0-1 each 0, `fw_power_state` 0, and `fw_perf_state_lo` and `fw_perf_state_hi` both 8. — [source](https://github.com/jundot/omlx/issues/4224)
- Apple's AGX driver lists "firmware-detected lockup" apart from its timeout restart reasons "timestamp timeout", "progress timeout" and "CDM Kill timeout". — [source](https://github.com/jundot/omlx/issues/4224)
- The command-buffer error for the firmware-detected lockup events is Hang (0x03), not Timeout (0x02). — [source](https://github.com/jundot/omlx/issues/4224)
- AGX `recoveryCount` read 10 against 8 gpuEvent reports since the 2026-09-15 boot, and the two unreported recoveries left no oMLX log line. — [source](https://github.com/jundot/omlx/issues/4224)
- The two 2026-08-27 events ran on macOS 26.5.1 (25F80) and have no surviving report, so their firmware signature is unknown. — [source](https://github.com/jundot/omlx/issues/4224)
- The Python frame where `Caused GPU Hang` surfaces is the first encode or wait on that stream after the reset and names the engine, not the kernel. — [source](https://github.com/jundot/omlx/issues/4224)
- At 11:06 on 2026-10-02 the hung queue held two buffers and its owner logged nothing, and at 14:22 and 16:00 queues with 10 and 11 discarded buffers logged nothing. — [source](https://github.com/jundot/omlx/issues/4224)
- A second M3 Ultra 512 GB reporter confirms that on macOS 27 each GPU Hang leaves a gpuEvent report with `restart_reason 4` ("firmware-detected lockup") and `signature 579`. — [source](https://github.com/jundot/omlx/issues/3706)
- Metal error 0x04 `kIOGPUCommandBufferCallbackErrorSubmissionsIgnored` reads "Ignored (for causing prior/excessive GPU errors)" and means the driver rejects every later command buffer from the process. — [source](https://github.com/jundot/omlx/issues/3706)
- Metal error 0x08 `kIOGPUCommandBufferCallbackErrorOutOfMemory` reads "Insufficient Memory". — [source](https://github.com/ml-explore/mlx-lm/issues/1206)
- Metal error 0x0b `kIOGPUCommandBufferCallbackErrorPageFault` reads "Caused GPU Address Fault Error". — [source](https://github.com/Blaizzy/mlx-vlm/issues/1064)
- Asahi Linux documents AGX as driven through an ASC coprocessor running Apple firmware, with work queue types TA, 3D and CP (compute) submitted over shared-memory channel rings. — [source](https://asahilinux.org/docs/hw/soc/agx/)
- Asahi Linux notes that GPU faults appear to be handled stop-the-world, with macOS dumping GPU MMIO registers directly. — [source](https://asahilinux.org/docs/hw/soc/agx/)
- No Apple documentation of the gpuEvent `analysis` schema, `signature 579` or `guilty_dm` was found in the macOS 27 release notes or elsewhere. — [source](https://developer.apple.com/documentation/macos-release-notes/macos-27-release-notes)
- The Metal-side meaning of the `fw_*_state` slot keys is unconfirmed; mapping them to the TA, 3D and compute queue types is an inference. — source: `asserted`
