GPU hot-unplug and surprise-removal safety for a Thunderbolt eGPU on Linux

GPU hot-unplug and surprise-removal safety for a Thunderbolt eGPU on Linux

Why unplugging a Thunderbolt eGPU is unsafe on Linux with NVIDIA and what the kernel does about it — pciehp surprise link-down handling, the PCI core’s disconnected flag and all-ones MMIO, DPC versus surprise removal, the NVIDIA driver’s missing hot-removal path against amdgpu and xe’s DRM hot-unplug contract, and a safe planned-detach runbook.


eGPU Hot-Unplug and Surprise-Removal Safety on Linux

verified-as-of 2026-09-24. Worked example: RTX 5080 in Razer Core X V2 over Thunderbolt 4 (TB4; “TB” below means Thunderbolt), Intel NUC 15 Pro, Ubuntu 26.04, kernel 7.0, nvidia 610.57.04-open. Web research only; nothing here was tested on the machine. Tags: [SOURCED url] = read on that page in this run; [INFERRED] = my synthesis; [UNVERIFIED] = not confirmed. A qualifier after the URL (“via search summary”, “page not opened”, “search snippet”) means only a snippet or summary was read, not the page itself: treat those tags as lower confidence.

Detaching now? Go to “Planned Detach Runbook” and read its “Before you start” block first.

Core Concepts

  1. Surprise vs planned removal. Planned (pciehp’s term is “safe removal”) = software tears down driver, then PCI device, then tunnel, then cable is pulled. Surprise = link drops first; software learns from pciehp Link Down / Presence Detect change. Over Thunderbolt the “slot” is a tunnelled PCIe downstream port; cable pull collapses the tunnel and pciehp sees the link go down. [INFERRED] from [SOURCED https://github.com/CachyOS/linux-cachyos/issues/794] (log: “pciehp: Slot(6): Link Down”, “Card not present”, “amdgpu: device lost from bus!” on Razer Core X).
  2. Reads of a gone device return all-ones. “Read requests to the device will time out after (typically) 17ms, and return a fabricated ‘all ones’ response”; naive removal of many devices took seconds and could raise machine checks. [SOURCED https://lwn.net/Articles/767885/]. PCI error-recovery doc: isolated slot “all reads return 0xffffffff, all writes are ignored” (powerpc description). [SOURCED https://docs.kernel.org/PCI/pci-error-recovery.html]
  3. Disconnected flag. A 4.12 kernel change marks surprise-removed devices so the PCI core skips accesses; removal dropped from seconds to microseconds; the PCI subsystem maintainer cautioned the flag is set asynchronously and may hide driver bugs. [SOURCED https://lwn.net/Articles/767885/]. Driver rule of thumb: an all-ones read is not proof of removal; pci_dev_is_disconnected() distinguishes real surprise disconnect from module unload (rmmod) [SOURCED https://lore.kernel.org/linux-pci/[email protected]/t/ via search summary].
  4. Quiesce differs by path. pciehp_unconfigure_device(): on surprise (presence false) it runs pci_walk_bus(parent, pci_dev_set_disconnected, NULL); on safe removal it clears Bus Master and SERR and sets INTx disable before unbinding. [SOURCED https://raw.githubusercontent.com/torvalds/linux/master/drivers/pci/hotplug/pciehp_pci.c]
  5. Driver-level hot-unplug contract (DRM). Userspace should keep working “more or less” until it closes the fd; it learns via uevent, ioctls returning ENODEV, open() returning ENXIO; GPU jobs must have fences force-signalled; SIGBUS on mmap “is not an option”. [SOURCED https://docs.kernel.org/gpu/drm-uapi.html]. NVIDIA’s proprietary/open Resource Manager (RM) stack is not documented to implement this contract (see Driver Comparison). [INFERRED]
  6. Safe removal can race surprise removal. If a user-initiated remove (e.g. driver unbind) blocks waiting for an interrupt from the device and the device is then surprise-removed, remove “hangs forever”; a 2026 RFC schedules work from pciehp_isr() to report surprise removal while the IRQ thread is blocked. [SOURCED https://ratatoskr.run/linux-pci/2026/09/17519182/t] (RFC; the thread shows no merge).
  7. Tunnel de-authorization = PCIe hot-remove. Writing 0 to a Thunderbolt device’s authorized attribute tears down the PCIe tunnel; needs connection-manager support (domain deauthorization attr); doc warns of data loss if storage isn’t shut down. [SOURCED https://docs.kernel.org/admin-guide/thunderbolt.html] Caveat: on a software connection manager, de-authorization can leave the PCI devices in the tree until pciehp reacts, so the runbook removes them first. [INFERRED from the thunderbolt-usb4-pcie-tunnel-bolt-iommu-linux sibling]

Surprise Removal in the PCI Core

DPC vs surprise removal

Kernel 6.19 regression signal

CachyOS 6.19.10-1 dropped a TB Razer Core X (RX 6800 XT) with pciehp Link Down; 6.19.6-2 worked; cause not identified in the issue. [SOURCED https://github.com/CachyOS/linux-cachyos/issues/794]. Lesson [INFERRED]: link-down on a tunnelled port can be a kernel/tunnel regression, not only a cable pull; check dmesg ordering before blaming the GPU.

Driver Comparison

Aspect nvidia (proprietary + open modules) amdgpu xe (i915 not researched)
Hot-unplug design No reliable support claimed. An NVIDIA maintainer (2023-01-31): “I suspect there is a lot of work still necessary to reliably support GPU hotplug/hotunplug.” [SOURCED https://github.com/NVIDIA/open-gpu-kernel-modules/discussions/451] README: stability on unplug “is not guaranteed”; X screens not configured on eGPUs by default (AllowExternalGpus). [SOURCED https://download.nvidia.com/XFree86/Linux-x86_64/396.51/README/egpu.html] (older 396.51 README; re-check current) drm_dev_unplug-based support built via multi-year RFC series (v2..v7), forcibly unmapping BO VMAs and failing faults after removal. [SOURCED https://lore.kernel.org/linux-pci/[email protected]/t/ via search summary] xe: gates PM/HW work on drm_dev_enter(); 2026 patches skip GuC CLEANUP for exec queues after unplug and tests via IGT core_hotunplug. [SOURCED https://ratatoskr.run/intel-xe/2026/07/17330261/t]
Removal with open fds nvidia_remove sees users and loops os_schedule() while holding NV_LINUX_DEVICES lock: hangs “quite deliberately”. [SOURCED https://lab.whitequark.org/notes/2018-10-28/patching-nvidia-gpu-driver-for-hot-unplug-on-linux/] (2018 driver; may have changed) Log: “NVRM: Attempting to remove device with non-zero usage count!”, then 100% CPU lockup on RTX 5070 Ti eGPU, driver 575.64 open. [SOURCED https://github.com/NVIDIA/open-gpu-kernel-modules/issues/894] Remove completes; fds stay valid but return errors (ENODEV) per DRM contract [INFERRED from DRM doc] Remove completes; exec queues destroyed locally [SOURCED xe patch above]
eGPU flag nv->is_external_gpu set by RmCheckForExternalGpu; issue #894 reporter found it misdetected an ADT-Link dock and forcing NV_TRUE avoided the lockup. [SOURCED #894] Whether it is set for Razer Core X V2 on this host: [UNVERIFIED] Runtime PM disabled if pci_is_thunderbolt_attached() or dev_is_removable(); the latter added because ASM4242 USB4 hosts lack a marked TB ancestor. [SOURCED https://ratatoskr.run/amd-gfx/2026/08/17400523/t] n/a
Upstream fix state PR #985 “Thunderbolt eGPU hot-unplug kernel support” (20 commits, RTX 3060 TB3): open, unmerged, no maintainer engagement as of the 2026-09-24 research. [SOURCED https://github.com/NVIDIA/open-gpu-kernel-modules/pull/985] in-tree [INFERRED; the cited thread is an RFC series] in-tree, actively patched
Recovery after unplug Reboot in the reported cases where the driver wedged [SOURCED #894, whitequark] replug works with patches [SOURCED lore thread] replug expected [INFERRED]

What happens to userspace at removal

Removal State Machine

PLANNED                                   SURPRISE
--------                                  --------
ATTACHED_ACTIVE                           ATTACHED_ACTIVE
  | stop clients (CUDA, compositor)         | cable pulled / enclosure power lost
  v                                         v
DRAINED  (drain -m 1; persistenced off)   LINK_DOWN (TB tunnel torn down; pciehp DLLSC/presence)
  | unbind nvidia (fds all closed)           | dpc_handler: surprise-down ignored / or DPC frozen state
  v                                         v
DRIVER_UNBOUND                            PCI_MARKED_DISCONNECTED (pci_walk_bus set_disconnected)
  | echo 1 > .../remove  (sysfs remove)      | driver ->remove() runs with device already gone
  v                                         v
PCI_REMOVED                               +-- driver tolerant (amdgpu/xe): fds get ENODEV -> CLEAN
  | deauthorize TB device (authorized=0)    +-- driver intolerant (nvidia): remove waits for usage=0
  v                                              -> WEDGED (D-state, lock held) -> reboot
TUNNEL_DOWN -> physical unplug -> DETACHED

[INFERRED] synthesis of the sourced pieces above.

Planned Detach Runbook

Untested on hardware: the sysfs steps use standard kernel interfaces; the nvidia-smi drain step is unconfirmed.

Before you start

  1. Stop consumers: sudo systemctl stop <inference-service> (for example ollama or a llama.cpp server), then kill remaining CUDA jobs and stop GPU containers.
  2. Stop nvidia-persistenced (and dcgm) since they pin fds [SOURCED gpu-operator #1263]; stop any other daemon that opens the device, and if a stopped daemon restarts on demand, mask it until step 8 (sudo systemctl mask, then unmask) [INFERRED]. Then confirm nothing holds the GPU: sudo fuser -v /dev/nvidia* and sudo lsof /dev/nvidia* must both print nothing (without sudo they skip other users’ processes and give a false all-clear). A clean result is necessary but not sufficient, since kernel-side references (for example nvidia_uvm) are invisible to both; also check the “Used by” counts in lsmod | grep nvidia. Also run sudo fuser -v on the eGPU’s DRM nodes (/dev/dri/card*, renderD*; the device symlinks under /sys/class/drm/ show which GPU owns each). [INFERRED]
  3. Turn persistence mode off, then drain: sudo nvidia-smi -i 0000:BB:00.0 -pm 0 && sudo nvidia-smi drain -p 0000:BB:00.0 -m 1 [SOURCED via search summary; verify with nvidia-smi drain -h]. If drain is rejected as unsupported, re-run the step 2 check and continue; if anything else fails, stop before step 4. Note that nvidia-smi may print bus IDs with an 8-digit domain (00000000:BB:00.0); use its form if -i or -p rejects yours. [INFERRED]
  4. Unbind only when usage is zero: echo 0000:BB:00.0 | sudo tee /sys/bus/pci/drivers/nvidia/unbind. If it logs “non-zero usage count”, STOP; do not retry blindly (lock-holding loop). [SOURCED #894] The write is synchronous: with a client still attached, the remove path loops in the kernel and the command may never return, so watch the second terminal instead of waiting on this one. [INFERRED] Unlike this write, sudo modprobe -r of the nvidia modules (the boot-orchestration sibling’s variant) refuses while a device file is open. [INFERRED from that sibling]
  5. Confirm the unbind took effect (ls -l /sys/bus/pci/devices/0000:BB:00.0/driver should fail). Then remove the PCI function(s): echo 1 | sudo tee /sys/bus/pci/devices/0000:BB:00.0/remove, then the same for 0000:BB:00.1 (the audio function). [INFERRED standard sysfs]
  6. Tear down the tunnel: find the enclosure’s entry under /sys/bus/thunderbolt/devices/ (match it with boltctl list), check that the domain’s deauthorization attribute (next to security) reads 1, then run echo 0 | sudo tee /sys/bus/thunderbolt/devices/<dev>/authorized; the kernel doc warns this is a PCIe hot-remove. [SOURCED thunderbolt doc] If deauthorization reads 0 or the write fails, skip this step: step 5 already removed the PCI devices, and on firmware-connection-manager hosts unplugging is the only teardown. [INFERRED from the thunderbolt sibling] No boltctl de-authorize command is verified [UNVERIFIED]; boltctl forget deletes the enrollment and is not a detach step. [INFERRED from the boot-orchestration sibling]
  7. Verify, then unplug. sudo dmesg | tail -50 (or sudo journalctl -k -b) should show the pciehp removal and no NVRM errors or asserts, and lspci -D | grep -i nvidia should print nothing. If either check fails, do not unplug: use the fallback below. Then unplug the cable, then power off the enclosure.
  8. Re-attach: power on the enclosure, plug in the cable, and authorize the device (bolt/udev). Once lspci -D lists the GPU, load the nvidia modules if you unloaded them, start nvidia-persistenced, then restart the services you stopped in step 1. To restore the GPU after step 5 without unplugging, run echo 1 | sudo tee /sys/bus/pci/rescan. [INFERRED] If the GPU still reports as drained, clear it with sudo nvidia-smi drain -p 0000:BB:00.0 -m 0 [UNVERIFIED]. [INFERRED]

Fallback. If any step from 3 to 7 fails or hangs: stop, check the second terminal, and never retry the unbind. If the unbind write has not returned, treat the driver as wedged: the safe path is a planned reboot with the cable connected, since pulling the cable on a wedged driver adds NVRM assertions and is unlikely to help. [INFERRED from #894] If drain is simply rejected as unsupported (untested, see Open questions), the step 2 check is your only guard against new clients, so repeat it immediately before step 4. [INFERRED] A wedged remove path can also stall a normal shutdown; give systemd its stop timeout before forcing a reboot, because a forced reboot skips clean filesystem shutdown. [INFERRED]

Userspace Tooling

Anti-patterns

Open questions / unverified

See also (sibling references in this hub)

Related references added later: egpu-unattended-remote-recovery-and-out-of-band-linux.md (host-first escalation ladder, remote power control and out-of-band access).

Sources