Health monitoring, alerting and safe automated recovery for a Thunderbolt eGPU

Parent: Thunderbolt eGPU on Linux for local LLM inference · Published reference · snapshot 2026-09-08 · skill devops-linux-internals/references/egpu-health-monitoring-and-automated-recovery-linux.md

Also known as: egpu auto recovery, egpu monitoring, egpu watchdog, egpu-reinit automation, liveness probe gpu, node_exporter textfile, nvidia xid alert, systemd OnFailure

↓ Facts as markdown↓ Download this reference fileall context files

Watching a Thunderbolt eGPU on a headless Linux box and recovering it safely — the kernel-log and sysfs signatures worth alerting on, a guarded liveness probe that tells a dead GPU from a bridge with

These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.

Health monitoring, alerting and safe automated recovery for a Thunderbolt eGPU

Thunderbolt eGPU health monitoring and automated recovery (Linux)

Core Concepts

Signals to Watch

Log Signatures

sysfs and NVML Health Probes

Liveness Probe

Alerting

Automated Recovery and Safeguards

Templates

Anti-patterns

  • Related references added later: egpu-unattended-remote-recovery-and-out-of-band-linux.md (host-first escalation ladder, remote power control and out-of-band access); egpu-reproducible-bringup-and-drift-detection-linux.md (capturing this wiring as a restorable manifest, drift verifier and restore order). [source]
  • Sources

  • Not fetched (two man-page hosts returned 403 during research): journald.conf(5), journalctl(1), NVML API reference, egpu-community watchdog scripts. Statements resting on those are tagged [INFERRED]/[UNVERIFIED]. [source]
  • Unverified items (summary)

    Children

    Frontier under this node: AER sysfs counters and rate alerting, BAR0 chip-ID liveness probe, DCGM vs NVML vs nvidia-smi on GeForce, NVIDIA Xid recovery-action taxonomy, PCI bridge COMMAND Mem/BusMaster diagnosis, auto-recovery rate limit, backoff and lockout state machine, break-glass manual recovery runbook, node_exporter textfile collector for hardware health, persistent journal kernel-evidence retention, systemd OnFailure oneshot watchdog pattern

    ← the whole tree · 3D view· how to read this page