<!-- llms-explorer concept facts · https://llms-explorer.com/tree/egpu-suspend-resume-sleep/ · pack 2026-09-08 · ~8240 tokens -->

# Suspend, resume and sleep states with a Thunderbolt NVIDIA eGPU on Linux

> What happens to a Thunderbolt NVIDIA eGPU across suspend on Linux and what to do about it on a headless box — s2idle versus S3 and modern standby on a NUC, the fate of the tunnel across sleep, NVIDIA'

Parent: [Thunderbolt eGPU on Linux for local LLM inference](https://llms-explorer.com/tree/thunderbolt-egpu-linux/) · 14 facets · 84 facts · page: https://llms-explorer.com/tree/egpu-suspend-resume-sleep/

## Suspend, Resume and Sleep States with a Thunderbolt NVIDIA eGPU on Linux

- What happens to a Thunderbolt NVIDIA eGPU across suspend on Linux and what to do about it on a headless box - s2idle versus S3 and modern standby on a NUC, the fate of the tunnel across sleep, NVIDIA's suspend machinery and the open RTX 50 suspend regressions, how to disable suspend safely with undo steps, and a pre-sleep hook with post-resume validation for when sleep is required. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#suspend-resume-and-sleep-states-with-a-thunderbolt-nvidia-egpu-on-linux)
- verified-as-of: 2026-09-25 tag legend: [SOURCED url] = read in this research pass; [INFERRED] = reasoned from sourced facts, not directly stated; [UNVERIFIED] = plausible, not confirmed. No commands here were run on the target box. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#suspend-resume-and-sleep-states-with-a-thunderbolt-nvidia-egpu-on-linux)
- Safety summary. Every state-changing command below lists its UNDO, whether it needs sudo, and the file path involved. Nothing in this file should be run until you have read "Test Procedure" (hang recovery). Default recommendation: do not suspend a headless LLM box (Decision Guide, option A). [INFERRED] — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#suspend-resume-and-sleep-states-with-a-thunderbolt-nvidia-egpu-on-linux)

## Core Concepts

- Sleep state vs suspend variant. The kernel offers up to four states: suspend-to-idle (S2Idle, sysfs string freeze), standby, suspend-to-RAM (mem, ACPI S3) and hibernation (disk). mem is an alias whose meaning is picked by /sys/power/mem_sleep (s2idle, shallow, deep). [SOURCED https://docs.kernel.org/admin-guide/pm/sleep-states.html] — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#core-concepts)
- s2idle is software-only. It freezes user space and puts devices in low power so CPUs reach deepest idle; wake relies on in-band interrupts, so resume is fast but savings are smaller. S3 takes non-boot CPUs offline and hands control to firmware. [SOURCED kernel sleep-states] — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#core-concepts)
- Hibernate writes RAM to storage and powers nearly everything off, so it is a different failure surface (needs swap/resume device; the NVIDIA hibernate unit saves VRAM too). [SOURCED kernel sleep-states; NVIDIA README] — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#core-concepts)
- A Thunderbolt eGPU is a PCIe endpoint behind a tunnel. The tunnel exists only while the TB link is up and the device is authorized (security level user/secure requires writing authorized). Kernel Thunderbolt docs do not describe suspend/resume behaviour at all. [SOURCED https://docs.kernel.org/admin-guide/thunderbolt.html] Everything about tunnel, D3cold, pciehp and bolt behaviour across sleep below is therefore reasoned, not sourced, and is tagged [INFERRED] wherever it appears; confirm it with the Test Procedure. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#core-concepts)
- NVIDIA driver has two suspend paths: the default kernel callback (saves only "essential" video memory) and the /proc/driver/nvidia/suspend interface driven by systemd units (needed for CUDA/UVM-class apps; can preserve all VRAM). [SOURCED https://download.nvidia.com/XFree86/Linux-x86_64/580.95.05/README/powermanagement.html] — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#core-concepts)
- Failure mode that matters on a headless LLM box: a failed resume of a GPU behind TB usually means a dead link or GSP fault needing a reboot or cold power cycle, and Ollama/CUDA state is lost either way. Sleep saves little (idle GPU already low power) and risks a lot. [INFERRED] — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#core-concepts)

## Sleep States and How to See Them

- Read-only inspection (safe; no root needed except where noted; none of these change state): — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#sleep-states-and-how-to-see-them)
  - /sys/power/state lists supported strings; freeze = s2idle, mem = suspend-to-RAM, disk = hibernate. [SOURCED kernel sleep-states] — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#sleep-states-and-how-to-see-them)
  - mem_sleep_default= (kernel cmdline) overrides which variant mem uses; default is deep where S2RAM is supported, else s2idle. [SOURCED kernel sleep-states] systemd can also set MemorySleepMode= (default empty = kernel default / mem_sleep_default= respected). [SOURCED https://man7.org/linux/man-pages/man5/systemd-sleep.conf.5.html] — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#sleep-states-and-how-to-see-them)
  - If deep is not listed in mem_sleep, the firmware is not exposing S3; forcing mem_sleep_default=deep cannot create it. [INFERRED] — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#sleep-states-and-how-to-see-them)
  - systemctl suspend uses the mem string (variant per above); systemctl hibernate uses disk. logind SleepOperation= (v256+) default tries suspend-then-hibernate suspend hibernate in order for the fallback chain of the lid/idle actions. [SOURCED https://man7.org/linux/man-pages/man5/logind.conf.5.html] — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#sleep-states-and-how-to-see-them)
  - Wake sources: s2idle can in theory wake on any interrupt-capable device; standby/S3 wake sources are more limited and platform-configured. [SOURCED kernel sleep-states] Inspect with cat /proc/acpi/wakeup and /sys/class/net/*/device/power/wakeup [UNVERIFIED path names, standard but not re-read here]. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#sleep-states-and-how-to-see-them)
- To change the variant for a test (root needed for the persistent form). UNDO: remove the token, run sudo update-grub, reboot. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#sleep-states-and-how-to-see-them)
- Prefer the one-shot form: it cannot outlive a hang-and-power-cycle. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#sleep-states-and-how-to-see-them)

## Modern Standby on a NUC

- Windows "Modern Standby" is the OS-level S0ix idea; on Linux it maps to s2idle plus platform idle states (S0ix/package C10). Firmware that advertises only modern standby usually lacks S3. [INFERRED] — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#modern-standby-on-a-nuc)
- This box: ASUS NUC15CRKU5 (Arrow Lake). ASUS documents a "modern standby indicator" and "ErP Ready" only; S3 availability is undocumented. Decision rule: trust /sys/power/mem_sleep on this machine, not the spec sheet. If it shows only [s2idle], S3 is not on offer. [INFERRED from kernel docs] — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#modern-standby-on-a-nuc)
- S0ix depth depends on every PCIe device reaching D3/L1.2; a live TB controller with an eGPU behind it is a classic blocker for deep S0ix residency. [INFERRED] NVreg_EnableS0ixPowerManagement=1 is only meaningful with suspend state s2idle and a GPU that reports "Video Memory Self Refresh" in /proc/driver/nvidia/gpus/<BDF>/power; below a 256 MB VRAM-use threshold (default) the driver copies VRAM to system RAM and powers it off. [SOURCED NVIDIA README] It targets laptop dGPUs; for an eGPU keep it off (its effect over Thunderbolt is [UNVERIFIED]). — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#modern-standby-on-a-nuc)
- ErP Ready (EU standby power) in BIOS can cut standby power to USB/LAN and therefore wake sources. [INFERRED] See asus-nuc15-pro-firmware-for-thunderbolt-egpu-linux.md. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#modern-standby-on-a-nuc)

## What Happens to the Tunnel

- No primary source documents this end to end. The whole sequence below is reasoned from PCIe/Thunderbolt design, not sourced, and each step is tagged [INFERRED]; it must be verified by the Test Procedure. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#what-happens-to-the-tunnel)
  - Pre-sleep (reasoned): userspace frozen; NVIDIA units (if run) suspend the GPU via /proc/driver/nvidia/suspend; PCI core moves devices to D3hot; the TB controller may enter D3cold and tear the tunnel down to save power. [INFERRED] — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#what-happens-to-the-tunnel)
  - Resume (reasoned): the TB host controller re-enumerates the link. Whether the eGPU comes back authorized without user action depends on the security level: at none it should re-tunnel automatically (dponly creates no PCIe tunnel at all, so an eGPU cannot work there [UNVERIFIED; confirm with the kernel thunderbolt doc]); at user/secure it needs re-authorization (bolt policy auto/stored keys probably handle this [INFERRED]). [SOURCED kernel thunderbolt doc for the authorization rule; resume behaviour INFERRED] — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#what-happens-to-the-tunnel)
  - pciehp (reasoned) on the downstream ports may report presence-detect change and remove/re-add the whole subtree; the GPU then appears as a new device and BAR/bridge windows are re-assigned. A GPU that vanished during sleep and re-enumerated is a fresh device to the driver, which is exactly the case a saved VRAM image cannot survive. [INFERRED] See egpu-hot-unplug-pciehp-safety-linux.md and linux-egpu-hotplug-boot-orchestration.md. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#what-happens-to-the-tunnel)
  - Symptoms to grep after resume: pciehp, thunderbolt, NVRM, Xid, fallen off the bus, GSP. A GPU reading 0xFFFFFFFF in config space or BAR0 is unreachable (not by itself proof it was removed; see linux-nvidia-egpu-fallen-off-bus-diagnosis.md for the BAR0 probe and pcie-power-management-aer-dpc-egpu-linux.md for D3cold versus gone). [INFERRED] — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#what-happens-to-the-tunnel)
  - Because the box loads the driver late via a systemd oneshot, a re-enumerated GPU will not be re-bound unless that oneshot (or a udev rule) re-runs; see linux-egpu-hotplug-boot-orchestration.md. [INFERRED from context] — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#what-happens-to-the-tunnel)

## NVIDIA Suspend Machinery

- Units: nvidia-suspend.service, nvidia-hibernate.service (pre-sleep), nvidia-resume.service (post), plus the nvidia-sleep.sh helper; installed unless the installer used --no-systemd. [SOURCED NVIDIA README] (Ubuntu packages ship them; user reports they are enabled here.) — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#nvidia-suspend-machinery)
- /proc/driver/nvidia/suspend is the interface those units write to. The verbs (suspend, hibernate, resume) are [UNVERIFIED] wording from memory; confirm with the nvidia-sleep.sh helper the packages install (find it with dpkg -L on the driver package or systemctl cat nvidia-suspend.service). Do not write to this file by hand: it needs root, and freezing the GPU outside a real system sleep can wedge the driver. If you ever must, the matching undo is the resume verb [UNVERIFIED], and a wedged driver means a reboot. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#nvidia-suspend-machinery)
- NVreg_PreserveVideoMemoryAllocations=1 preserves all VRAM allocations (default preserves only essential ones). Backups go to NVreg_TemporaryFilePath (default /tmp); README recommends space of total VRAM plus about 5 percent and a non-tmpfs filesystem because /tmp and /run are often size-limited tmpfs. [SOURCED NVIDIA README] So for a 16 GB RTX 5080, budget about 17 GB free on a disk-backed path per suspend; a tmpfs /tmp that is smaller will fail the suspend. (16 GB figure is the card's nominal VRAM [INFERRED].) — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#nvidia-suspend-machinery)
- For a headless inference box, preserving VRAM is mostly pointless: model weights reload from disk, and the corrected =0 avoids writing ~16 GB per sleep. On this box a package modprobe file had overridden the value to 1; a later-sorting zz-nvidia-egpu-pm.conf modprobe file now sets it back to 0, effective only at the next driver load (a still-loaded driver keeps the old value; reload needs the safe detach in egpu-hot-unplug-pciehp-safety-linux.md, or a reboot). Check /proc/driver/nvidia/params after the reload; do not assume. [INFERRED] — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#nvidia-suspend-machinery)
- The enabled units do nothing unless a sleep is actually triggered; masking sleep targets makes them inert. [INFERRED] — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#nvidia-suspend-machinery)

## Known Issues

- Issue numbers below were returned by the GitHub search API for NVIDIA/open-gpu-kernel-modules on 2026-09-25 and details read from each page; all are open unless stated. Only #979 is titled as a TB-eGPU case (title only); none is confirmed as a TB-eGPU sleep failure. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#known-issues)
- [SOURCED https://github.com/NVIDIA/open-gpu-kernel-modules/issues/1291 , /1284 , /1389 , and search listing via api.github.com]. The other rows are from title/listing only; open the issue before citing symptoms. No maintainer response was visible on #1284/#1291/#1389 at read time. Whether a fixed driver exists after 610.57.04 is [UNVERIFIED]; recheck before relying on it. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#known-issues)

## Decision Guide

- Recommendation: A now; consider B after the box has been stable for weeks and a pin of kernel/driver exists. [INFERRED] — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#decision-guide)

## Disabling Suspend

- This section implements option A. It is safe to apply and every item has an UNDO. All system-level commands need root. This file has not applied any of them; check the current state first with systemctl is-enabled suspend.target (masked means already done). — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#disabling-suspend)
- Mask all sleep targets (blocks systemctl suspend/hibernate/hybrid-sleep; logind actions also fail). Masking creates /dev/null symlinks under /etc/systemd/system/, owned by root: — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#disabling-suspend)
- logind drop-in (never act on keys/idle). Path /etc/systemd/logind.conf.d/90-no-sleep.conf, owner root:root, mode 0644; create the directory with sudo install -d -m 0755 /etc/systemd/logind.conf.d and the file with sudoedit. UNDO: sudo rm /etc/systemd/logind.conf.d/90-no-sleep.conf, then reboot (restarting systemd-logind can end graphical sessions [INFERRED], so prefer a reboot). A drop-in only takes effect after logind restarts or the box reboots [INFERRED]. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#disabling-suspend)
- Defaults: HandleSuspendKey and HandleLidSwitch default to suspend; IdleAction defaults to ignore, acts only when all sessions are idle and no idle inhibitor is held. [SOURCED man7 logind.conf] systemd sleep.conf alternative: /etc/systemd/sleep.conf.d/90-no-sleep.conf (root:root, mode 0644, created with sudoedit) with [Sleep] AllowSuspend=no AllowHibernation=no AllowHybridSleep=no AllowSuspendThenHibernate=no. UNDO: sudo rm the file (takes effect at the next sleep request; no restart needed [UNVERIFIED]). AllowSuspend/AllowHibernation exist and default to advertising any supported mode [SOURCED man7 systemd-sleep.conf]; the Hybrid/SuspendThenHibernate option names are [UNVERIFIED] here. GNOME idle suspend (per user, no root; run as the desktop user inside their session, since over plain SSH gsettings has no session bus [UNVERIFIED]): — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#disabling-suspend)
- [UNVERIFIED key names; standard for GNOME but not re-read this pass; verify with gsettings list-keys.] Also [UNVERIFIED], from memory: the schema name and value strings; confirm with gsettings list-keys org.gnome.settings-daemon.plugins.power and gsettings range <schema> <key> before setting. The GDM login screen keeps its own settings under the gdm user [UNVERIFIED]; the masks above cover it regardless. Verify: systemctl status suspend.target shows masked; systemctl suspend returns an error and the box stays up. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#disabling-suspend)

## Pre-Sleep Hook and Post-Resume Validation

- Skeleton for option B, not run on the box. Do not install it until the Test Procedure preconditions and steps 3-4 (driver unloaded) have passed. Hooks live in /usr/lib/systemd/system-sleep/ with args pre|post and the sleep verb [UNVERIFIED path/args, standard systemd convention]. Also [UNVERIFIED], from memory: /usr/lib/... is package-owned and a local script belongs in /etc/systemd/system-sleep/, which is what the skeleton uses; confirm both with man systemd-sleep before installing. Install as root, root:root, mode 0755 (it must be executable): sudo install -m 0755 -o root -g root 50-egpu /etc/systemd/system-sleep/50-egpu. UNDO: sudo rm /etc/systemd/system-sleep/50-egpu, or disable it without editing via the kill switch sudo touch /etc/egpu-sleep.disable (remove the file to re-enable). — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#pre-sleep-hook-and-post-resume-validation)
- Design rules, so the hook cannot loop forever or strand the GPU half-detached [INFERRED]: no while/retry loops (only a fixed, bounded for list); every external command under timeout; it never kills processes (fuser -k on /dev/nvidia* can kill the desktop or persistenced); unload modules with modprobe -r, which refuses while a device file is open, never with a driver unbind (a remove with a non-zero usage count can spin in-kernel, see egpu-hot-unplug-pciehp-safety-linux.md); and a state file records progress so post can always recover. Whether a failing pre hook stops the sleep is [UNVERIFIED; confirm with man systemd-sleep]; assume it does not, which is why the hook is only ever tested after steps 3-4. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#pre-sleep-hook-and-post-resume-validation)
- Note the dmesg | tail check can match stale lines from before the sleep; treat a failure as "look at the log", not as proof. [INFERRED] — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#pre-sleep-hook-and-post-resume-validation)
- Post-resume validation, in order, abort on first failure and leave Ollama stopped: — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#pre-sleep-hook-and-post-resume-validation)
  - boltctl list (or /sys/bus/thunderbolt/devices/*/authorized) shows the eGPU authorized. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#pre-sleep-hook-and-post-resume-validation)
  - lspci -nn -d 10de: finds the GPU; config-space vendor ID reads 10de, not ffff. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#pre-sleep-hook-and-post-resume-validation)
  - BAR0 chip-ID probe (read the identification register as described in linux-nvidia-egpu-fallen-off-bus-diagnosis.md; do not hard-code an offset here [UNVERIFIED]). — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#pre-sleep-hook-and-post-resume-validation)
  - nvidia-smi -L then a short CUDA test; dmesg | grep -E 'Xid|NVRM|GSP' clean. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#pre-sleep-hook-and-post-resume-validation)
  - Start Ollama; run one small generation. Wake sources (all need root): Wake-on-LAN (sudo ethtool -s <if> wol g, must also be enabled in BIOS; [UNVERIFIED for this NUC], undo sudo ethtool -s <if> wol d; the setting probably does not survive a reboot [UNVERIFIED]), USB keyboard, or RTC alarm via sudo rtcwake -m mem -s 300 for tests (rtcwake exists in util-linux; flags [UNVERIFIED] here). Prove one wake source works before the first test; without one, the box stays asleep until someone presses the power button. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#pre-sleep-hook-and-post-resume-validation)

## Test Procedure

- Hang-recovery preflight (complete before any suspend; a suspend can hang the whole box and only a hard power cycle then recovers it): — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#test-procedure)
  - Be physically present, or have remote power (smart plug) you have already tested; a hung suspend cannot be fixed over SSH. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#test-procedure)
  - Second login path: a second machine that can SSH in after wake, plus a local console you can reach. SSH sessions die at suspend, so these are for diagnosis after wake, not during the sleep. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#test-procedure)
  - Save the journal first: confirm journalctl --list-boots shows earlier boots (persistent journal; if not, sudo mkdir -p /var/log/journal and reboot; UNDO by removing the directory [INFERRED]), then copy journalctl -k -b off-box before each cycle, because a hard power cycle can lose the last seconds of log. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#test-procedure)
  - Test only when no job is running: no inference, fine-tune, download or large write in progress; run sync right before each attempt. A hard power cycle can lose unsynced data. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#test-procedure)
  - Physical power-cycle warning: hold power 5 s only as the last resort; expect a filesystem journal replay on next boot. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#test-procedure)
  - Recent backup; Ollama stopped; suspend targets unmasked for the test (reverse the mask above) and re-masked afterwards. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#test-procedure)
  - Baseline: capture the read-only commands in "Sleep States" and journalctl -k -b to a file off-box. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#test-procedure)
  - Confirm /sys/power/mem_sleep and the NVreg_* values actually loaded. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#test-procedure)
  - Dry run without the eGPU driver loaded: stop Ollama and nvidia-persistenced, unload the modules with sudo modprobe -r nvidia_uvm nvidia_drm nvidia_modeset nvidia (it refuses while in use; do not force it), run sudo rtcwake -m freeze -s 120 (or systemctl suspend with an RTC wake). Success = returns in about 120 s, TB link and network back. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#test-procedure)
  - Repeat with the eGPU authorized but the driver unloaded. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#test-procedure)
  - Repeat with the driver loaded, idle, PreserveVideoMemoryAllocations as intended. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#test-procedure)
  - Repeat with the hook (option B), 20 cycles, then 20 more with a loaded model before suspend. Abort criteria (stop and revert to option A): any hang over 60 s past the wake time; any Xid 79/119/120/154 or GSP timeout; GPU missing from lspci after 30 s; nvidia-smi failing; a needed hard power cycle even once with the driver loaded; resume that needs a physical unplug/replug. Recovery if the box hangs: hold power 5 s, unplug the TB cable and the eGPU PSU for 30 s (drains GPU state), power on with the eGPU attached, then check journalctl -b -1 -k. From another machine: ping/SSH; if only the GPU is lost, try the re-init steps in linux-egpu-hotplug-boot-orchestration.md before rebooting. Afterwards revert everything the test changed: re-apply the mask from "Disabling Suspend", remove /etc/systemd/system-sleep/50-egpu (or touch /etc/egpu-sleep.disable), and revert cmdline/mem_sleep_default changes before normal use. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#test-procedure)

## Anti-patterns

- Enabling nvidia-hibernate paths or running systemctl hibernate without confirming swap/resume device and about VRAM+5 percent free disk on a non-tmpfs path. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#anti-patterns)
- Leaving NVreg_PreserveVideoMemoryAllocations=1 on a headless box by accident (a package modprobe file overrode the intended value here; check /proc/driver/nvidia/params, not the conf file). — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#anti-patterns)
- Forcing mem_sleep_default=deep when firmware does not list deep. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#anti-patterns)
- Suspending with Ollama or any CUDA process alive, or with the driver bound, over TB. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#anti-patterns)
- Trusting the spec sheet ("modern standby") or an unmasked GNOME idle timer; GNOME can suspend a headless-but-logged-in box. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#anti-patterns)
- Testing first with a loaded model and no remote recovery path, or installing the hook before the driver-unloaded dry runs pass. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#anti-patterns)
- Hand-writing to /proc/driver/nvidia/suspend, or using fuser -k / a driver unbind in a sleep hook (see the hook design rules). — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#anti-patterns)
- Leaving <placeholder> text in a shell hook: bash reads <name> as a redirect and can create stray files. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#anti-patterns)
- Assuming a 610.57.04 or 7.0 result on one card (#1284 is a 5090 on 7.1.6) transfers exactly to a 5080 over TB4; test, do not extrapolate. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#anti-patterns)

## Related references (not duplicated here)

  - pcie-power-management-aer-dpc-egpu-linux.md: shorter suspend/resume section, ASPM/D3cold, power/control=on and pcie_port_pm=off keep-port-awake options. This file is the fuller treatment of sleep; the two agree on NVreg_PreserveVideoMemoryAllocations=0 for an eGPU. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#related-references-not-duplicated-here)
  - egpu-hot-unplug-pciehp-safety-linux.md: the safe detach order the hook borrows. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#related-references-not-duplicated-here)
  - linux-egpu-hotplug-boot-orchestration.md: the load oneshot and udev re-init. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#related-references-not-duplicated-here)
  - asus-nuc15-pro-firmware-for-thunderbolt-egpu-linux.md: modern standby and ErP BIOS options. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#related-references-not-duplicated-here)
  - nvidia-open-kernel-modules-blackwell-linux.md: module parameters. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#related-references-not-duplicated-here)
- Related references added later: egpu-idle-power-and-energy-accounting-linux.md (idle watts, persistence mode, power limits and energy cost); egpu-unattended-remote-recovery-and-out-of-band-linux.md (host-first escalation ladder, remote power control and out-of-band access). — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#related-references-not-duplicated-here)

## Sources

- Linux kernel, System Sleep States: https://docs.kernel.org/admin-guide/pm/sleep-states.html (read) — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#sources)
- Linux kernel, Thunderbolt admin guide: https://docs.kernel.org/admin-guide/thunderbolt.html (read; silent on suspend) — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#sources)
- NVIDIA driver README ch. Power Management (580.95.05 copy): https://download.nvidia.com/XFree86/Linux-x86_64/580.95.05/README/powermanagement.html (read; check the copy matching driver 610 before relying on details) — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#sources)
- logind.conf(5): https://man7.org/linux/man-pages/man5/logind.conf.5.html (read) — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#sources)
- systemd-sleep.conf(5): https://man7.org/linux/man-pages/man5/systemd-sleep.conf.5.html (read) — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#sources)
- open-gpu-kernel-modules issues #1291, #1284, #1389 (read) and listing for #1325, #1125, #1142, #1209, #1372, #979, #1151, #1200 (titles only): https://github.com/NVIDIA/open-gpu-kernel-modules/issues — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#sources)
- Not obtained: Arch wiki suspend page (blocked by bot protection), freedesktop.org man pages (HTTP 403), a systemd-sleep hook man page, Ubuntu docs. — [source](https://llms-explorer.com/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux/#sources)

## Context files

- [Suspend, resume and sleep states with a Thunderbolt NVIDIA eGPU on Linux](https://llms-explorer.com/downloads/sources/global-ai-hub/egpu-suspend-resume-and-sleep-states-linux.md)
