<!-- llms-explorer concept facts · https://llms-explorer.com/tree/metalguard-user-space-gpu-crash-mitigations/ · pack 2026-10-05 · ~2710 tokens -->

# MetalGuard user-space GPU-crash mitigations

> Registry-driven gating: `KNOWN_PANIC_MODELS` (model by hardware by workload entries with a tier of `panic`, `abort` or `degradation`, `error_classes`, and a `verified_safe_alternative`), `MLX_VERSION_BLOCKLIST`, `MLX_LM_VERSION_BLOCKLIST` and `WORKLOAD_ADVISORIES` feed `authorize_mlx_spawn()`.

Parent: [Mac local LLMs: GPU stability and kernel panics](https://llms-explorer.com/tree/mac-local-llms-gpu-stability-and-kernel-panics/) · 1 facets · 42 facts · page: https://llms-explorer.com/tree/metalguard-user-space-gpu-crash-mitigations/

## Facts

- Registry-driven gating: `KNOWN_PANIC_MODELS` (model by hardware by workload entries with a tier of `panic`, `abort` or `degradation`, `error_classes`, and a `verified_safe_alternative`), `MLX_VERSION_BLOCKLIST`, `MLX_LM_VERSION_BLOCKLIST` and `WORKLOAD_ADVISORIES` feed `authorize_mlx_spawn()`. — source: `asserted`
- `authorize_mlx_spawn()` returns a `SpawnVerdict` and aggregates all blockers in one call: G1 panic cooldown (`evaluate_panic_cooldown`), G2 known-panic-model denylist filtered by GPU family, G3 MLX version blocklist at severity critical or high, G4 workload advisories at tier critical; non-critical workload advisories surface as a soft warning. — source: `asserted`
- Overrides need `override_with_reason` of at least 20 characters and write a JSONL record to `~/.cache/metal-guard/spawn_audit.jsonl`; audit failures never block the spawn. — source: `asserted`
- Prefill exposure: a per-worker cumulative prefill-token tracker (`prefill_exposure_tracker`) feeds an advisory worker-rotation gate (`worker_rotation.should_rotate_worker`) aimed at delayed-trigger panics; `oom_precheck` wraps the prefill allocation math with model dimensions and GPU memory auto-fetched. — source: `asserted`
- Panic scanning (`parse_panic_reports`) reads `/Library/Logs/DiagnosticReports`, `/var/db/PanicReporter` and `~/Library/Logs/DiagnosticReports`, in both `.panic` and modern `.ips` formats (since 1.1.0). — source: `asserted`
- GPU family detection (`apple_gpu_family`) reads `mx.device_info()` and maps `applegpu_g13` to `g17` prefixes to M1 through M5, so registry entries can be filtered per chip. — source: `asserted`
- It sets `AGX_RELAX_CDM_CTXSTORE_TIMEOUT=1` at import if unset, and records a `ctxstore_timeout` panic class. — source: `asserted`
- A `ResourceTracker` (default cold restart after 4000 inferences) targets the mlx-lm 1185 descriptor leak. — source: `asserted`
- 2026-04-10 v0.1.0 to v0.2.x (OOM recovery, size estimator); 04-12 v0.3.0 Metal health probe; 04-13 v0.4.0 hardware auto-config (a 2026-04-12 comment on mlx 3186 already cites v0.4.0, so CHANGELOG dates can lag tags); 04-14 v0.6.0 hardened lock reclaim; 04-16 v0.7.0 R-series defences and nine advisories (R2 `audit_wired_limit` flags overrides above 85% of unified memory, citing mlx-lm 1047); 04-17 v0.8.0 L9 cadence guard; 04-25 v0.9.0 panics 7-11 and a "when metal-guard is not enough" section; 04-27 v0.10.0 L10-L13 promoted and registry made community-curated; 04-28 v0.11.0-v0.11.7 (eight tags in a day, CI hotfixes, `WORKLOAD_ADVISORIES`); 05-02 v0.12.0 spawn gate; 05-18 v0.13.0 package split and drift-audit waves W3-W10; 05-19 v0.21-v0.24 and v1.0.0 stable; v1.1.0 shipped 2026-05-19 14:33 as the GitHub release. — source: `asserted`
- No release after 2026-05-19 appears on the releases page as of 2026-10-04, over four months and spanning macOS 27's release. — source: `asserted`
- The 1.1.0 CHANGELOG heading is dated 2026-05-15, four days before 1.0.0; the GitHub release date (2026-05-19, after v1.0.0) shows the heading date is a typo. — source: `asserted`
- v0.13.0 split a 7,297-line single file into a 17-module package; waves W3-W14 ported features from a private fork and ended with a "de-Harper cleanup sweep"; the registry's panic list is therefore partly vendor-internal production data. — source: `asserted`
- Repository metrics seen: 14 stars, 0 forks. — source: `asserted`
- Advisory contents (v0.19.0, 16 advisories): mlx-lm 1090 (before 0.31.3, global generation stream makes observer-mode thread pools race on Metal), mlx-lm 1256 (0.31.3 and later, thread-local stream fix incomplete), transformers 1011 (5.5.x removed `ReasoningEffort`, breaking Gemma 4 loads via mlx-vlm). — source: `asserted`
- Registry entries beyond the panic model: mlx-lm 0.31.3 flagged high for three concurrent server bugs (1208, 1215, 1206); Qwen3.5-122B-A10B VLM MTP 5-bit abort (M2 Ultra 64K-context prefill command-buffer timeout); Qwen3-Coder-Next-4bit server crash loop (about 420 restarts in 2.5 h); Qwen3.5-9B-4bit M5 Max LoRA first-backward OOM; Kimi K2.5 KV OOM on M3 Ultra; a `silent_corruption` class for VLM checkpoints loaded text-only. — source: `asserted`
- Abort-tier events are counted in `CooldownVerdict.abort_count_24h` for display only; only kernel panics that rebooted the machine drive the cooldown staircase. — source: `asserted`
- The `WORKLOAD_ADVISORIES` entry `lora_with_display_active` (v0.11.6) records the display-contention kill and its workaround (display sleep plus `caffeinate -s`); L7 subprocess isolation is documented as not helping, because the kill is above the process boundary. — source: `asserted`
- The vendor's own notes (v0.9.0) say it "narrows the race windows but does not eliminate panic" for Gemma 4 31B 8-bit, that `--prompt-cache-bytes` and adding RAM to 96 GB do not prevent the panic, and that 26.4.x and a 26.5 beta had not fixed it. — source: `asserted`
- The install path changed: the 2026-04-27 comment on mlx 3267 installs from a git tag ("no PyPI yet"); the README later uses `pip install metal-guard`. — source: `asserted`
- Nothing in the source supports use on macOS 27: no changelog entry, advisory or test names macOS 27 or IOGPUFamily 162. — source: `asserted`
- Value: vendor reports 9 panics per week to zero on an M1 Ultra pipeline (self-reported, no independent test); the shim-based controlled test in mlx 3186 reached zero panics without MetalGuard by removing wiring traffic. The two approaches are not compared head to head. — source: `asserted`
- Does any MetalGuard layer remain useful on macOS 27 and IOGPUFamily 162.x, and is the project maintained after May 2026? Unknown. — source: `asserted`
- Does `authorize_mlx_spawn` block the oMLX 4224 two-engine pattern? It has no concurrency gate; its cross-process lock (L8) covers processes, not threads. — source: `asserted`
- MetalGuard v1.1.0 is the latest GitHub release, published 2026-05-19 14:33, after v1.0.0 "first stable release". — [source](https://github.com/Harperbot/metal-guard/releases)
- The repository showed 14 stars and 0 forks when fetched on 2026-10-04. — [source](https://github.com/Harperbot/metal-guard/releases)
- The 1.1.0 CHANGELOG heading is dated 2026-05-15 while 1.0.0 is dated 2026-05-19. — [source](https://raw.githubusercontent.com/Harperbot/metal-guard/main/CHANGELOG.md)
- v0.12.0 (2026-05-02) added `authorize_mlx_spawn()` returning a `SpawnVerdict` with four hard gates: G1 panic cooldown, G2 known-panic-model denylist plus GPU family filter, G3 MLX version blocklist at critical or high, G4 workload advisories at critical. — [source](https://raw.githubusercontent.com/Harperbot/metal-guard/main/CHANGELOG.md)
- `authorize_mlx_spawn` overrides need `override_with_reason` of at least 20 characters and write a JSONL record to `~/.cache/metal-guard/spawn_audit.jsonl`. — [source](https://raw.githubusercontent.com/Harperbot/metal-guard/main/CHANGELOG.md)
- v0.11.6 (2026-04-28) added `WORKLOAD_ADVISORIES` with `lora_with_display_active`, `check_workload_advisory()`, `MLX_LM_VERSION_BLOCKLIST` (mlx-lm 0.31.3 flagged high) and `check_mlx_lm_version_blocked()`. — [source](https://raw.githubusercontent.com/Harperbot/metal-guard/main/CHANGELOG.md)
- The v0.11.6 advisory says subprocess isolation does not help the display-contention kill and that the workaround is display sleep plus `caffeinate -s`. — [source](https://raw.githubusercontent.com/Harperbot/metal-guard/main/CHANGELOG.md)
- v0.18.0 (2026-05-18) added `oom_precheck`, `prefill_exposure_tracker` and `worker_rotation` (advisory prefill-budget worker rotation). — [source](https://raw.githubusercontent.com/Harperbot/metal-guard/main/CHANGELOG.md)
- v0.19.0 brought `check_version_advisories()` to 16 advisories, adding mlx-lm 1090, mlx-lm 1256 and transformers 1011. — [source](https://raw.githubusercontent.com/Harperbot/metal-guard/main/CHANGELOG.md)
- v1.1.0 made `parse_panic_reports` scan `/var/db/PanicReporter` and `~/Library/Logs/DiagnosticReports` as well, and read `.ips` kernel-panic reports besides `.panic`. — [source](https://github.com/Harperbot/metal-guard/releases)
- v0.13.0 split a 7,297-line single-file `metal_guard.py` into a `metal_guard/` package of 17 submodules, and v0.24.0 was a "de-Harper cleanup sweep". — [source](https://raw.githubusercontent.com/Harperbot/metal-guard/main/CHANGELOG.md)
- v0.7.0 added `audit_wired_limit()` flagging `iogpu.wired_limit_mb` overrides above 85% of unified memory, citing the mlx-lm maintainer on mlx-lm 1047. — [source](https://raw.githubusercontent.com/Harperbot/metal-guard/main/CHANGELOG.md)
- v0.11.0 added `apple_gpu_family()` mapping `applegpu_g13` to `g17` to M1 through M5, and `ResourceTracker(cold_restart_after=4000)` for the mlx-lm 1185 descriptor leak. — [source](https://raw.githubusercontent.com/Harperbot/metal-guard/main/CHANGELOG.md)
- `CooldownVerdict.abort_count_24h` is informational and does not influence the exit code; the staircase lockout stays reserved for kernel panics that rebooted the machine. — [source](https://raw.githubusercontent.com/Harperbot/metal-guard/main/CHANGELOG.md)
- The v0.9.0 notes say macOS 26.4.x and a 26.5 beta had not fixed the panic, `--prompt-cache-bytes` does not prevent it, adding RAM to 96 GB does not prevent it, and MetalGuard narrows race windows without eliminating panic on one Gemma 4 31B workload. — [source](https://raw.githubusercontent.com/Harperbot/metal-guard/main/CHANGELOG.md)
- The registry in v0.11.x lists Qwen3.5-122B-A10B VLM MTP 5-bit, Qwen3-Coder-Next-4bit (about 420 restarts in 2.5 h), Qwen3.5-9B-4bit, Qwen3.6-35B-A3B VLM MTP 8-bit (`silent_corruption`) and kimi-k2.5 as abort or degradation entries. — [source](https://raw.githubusercontent.com/Harperbot/metal-guard/main/CHANGELOG.md)
- On 2026-04-27 the vendor told mlx 3267 readers to install from a git tag ("no PyPI yet"). — [source](https://github.com/ml-explore/mlx/issues/3267)
- No MetalGuard changelog entry, advisory or test names macOS 27 or IOGPUFamily 162. — [source](https://raw.githubusercontent.com/Harperbot/metal-guard/main/CHANGELOG.md)
- MetalGuard's lock covers processes and thread registry, but no gate targets two engine threads submitting concurrently in one process. — source: `asserted`
