Thunderbolt 5, Barlow Ridge and OCuLink eGPU topologies on Linux

Thunderbolt 5, Barlow Ridge and OCuLink eGPU topologies on Linux

Choosing an eGPU connection topology beyond Thunderbolt 4 on Linux for local-LLM inference — what Thunderbolt 5 and the Barlow Ridge controllers actually change and their Linux status, how a TB5 enclosure behaves on a TB4 host, OCuLink and M.2 direct-PCIe as the lower-overhead alternative, and how link bandwidth affects resident inference, MoE offload and multi-GPU splits.


TB5 / Barlow Ridge and OCuLink eGPU Topologies (Linux, LLM inference)

Verified-as-of 2026-09-25 (source-reading date; a claim is verified only to the level its tag states). Tags: [SOURCED url] = read in a source; [INFERRED] = reasoning from sourced facts; [BOX] = this machine’s stated context (not probed); [UNVERIFIED] = could not confirm. Within SOURCED, “search summary” means only a search snippet was seen and “title only” means only the page title was readable; treat those, vendor listings and secondary blogs as claims, not measurements. Cross-refs (not re-covered; sibling files here): thunderbolt-usb4-pcie-tunnel-bolt-iommu-linux (tunnel sibling: TB4 bandwidth table), pcie-link-training-speed-width-thunderbolt-egpu-linux (link-training sibling: why 2.5 GT/s and “(downgraded)” are cosmetic), blackwell-sm120-llm-inference-stack-linux (in the ai-llm-model-layer hub) (sm_120 sibling: 16 GB residency, tunnel cost), egpu-hot-unplug-pciehp-safety-linux (hot-unplug sibling), egpu-power-enclosure-and-thermals-linux (power sibling: ATX PSU, TGP = total graphics power), linux-nvidia-egpu-fallen-off-bus-diagnosis (fallen-off-bus sibling: Xid 79), asus-nuc15-pro-firmware-for-thunderbolt-egpu-linux (NUC-firmware sibling).

Core Concepts

  1. Tunnel vs native PCIe. Thunderbolt/USB4 carries PCIe as a tunneled protocol inside the link; OCuLink/M.2/slot risers are plain PCIe lanes with no tunnel, no controller pair, no security-level authorization. [INFERRED from sources below]
  2. TB4 vs TB5 PCIe budget. TB4 allots ~32 Gb/s to PCIe (Gen3 x4-class); TB5 allots 64 Gb/s (Gen4 x4). [SOURCED https://egpu.io/forums/laptop-computing/thunderbolt-5-specs-confirmed80gpbs-bi-drectional-up-to-120gpbs-boost-speed-intel-demos-laptop-protoype-with-tb5-port/paged/6/ ; search summary of thunderboltlaptop.com] These are nominal ceilings; measured TB4 payload is lower (link-training sibling). The 80 Gb/s link and 120/40 Gb/s boost do NOT raise the PCIe tunnel above Gen4 x4 for an eGPU. [SOURCED https://videocardz.com/newz/razer-core-x-v2-now-available-349-99-egpu-enclosure-with-thunderbolt-5 (via search summary)]
  3. Three parties must all be TB5 for TB5 PCIe: host controller, cable (80 Gb/s-rated), enclosure/device controller. Any weaker party sets the link. [INFERRED]
  4. Barlow Ridge is a discrete controller family. JHL9580 = host/controller SKU, PCIe 4.0 x4 upstream, dual port, Q3’24, $19 RCP; JHL9480 = device/accessory controller (Q3’24); JHL9540 = TB4-class SKU of the family. [SOURCED https://www.intel.com/content/www/us/en/products/sku/225921/intel-jhl9580-thunderbolt-5-controller/specifications.html ; https://www.techpowerup.com/318236/details-of-intels-barlow-ridge-thunderbolt-5-controller-leaks] Which source backs each part-number role is not recorded (the JHL9480 page was not fetched), so JHL9540 = TB4-class is [UNVERIFIED]. JHL9586 and exact hub/host/device partitioning per part number: [UNVERIFIED] (not found in fetched sources). The Razer enclosure’s hub 8086:5786 is [BOX].
  5. Inference uses the link unevenly. Weights cross the link once (load); afterwards a fully-resident single GPU exchanges token IDs and logits, which one secondary blog puts at ~1.3% of link capacity. [SOURCED https://localaimaster.com/blog/egpu-local-ai-benchmarks - article’s own figures, treat as estimate] That is a worst-case estimate (fp32 logits at an assumed 100 tok/s versus ~4 GB/s theoretical TB4; sm_120 sibling), not a measured load [UNVERIFIED].
  6. What crosses the link depends on placement: resident = almost nothing; MoE CPU-expert offload = all CPU-side expert weights copied to GPU per prompt batch; KV spill = per-token reads; tensor-parallel = per-layer all-reduce. [SOURCED https://huggingface.co/blog/Doctor-Shotgun/llamacpp-moe-offload-guide ; https://github.com/ggml-org/llama.cpp/blob/master/docs/multi-gpu.md]
  7. Hot-plug is a topology property. Thunderbolt supports (managed) hot-plug (hot-unplug sibling); OCuLink docks are sold as not hot-pluggable. [SOURCED search summary of OCuLink dock listings, e.g. https://www.amazon.com/OwlTree-External-Graphics-SFF-8612-SFF-8611/dp/B0GGR8SYPJ - vendor claim; consistent with OCuLink being a raw PCIe cable]

Thunderbolt 5 and Barlow Ridge on Linux

Kernel timeline

Known issues (discrete Barlow Ridge HOST ports)

Implication for this box [BOX]: the NUC’s integrated TB4 host is a different controller from the discrete Barlow Ridge (BR) host chips behind the bugs above (Arrow Lake per the NUC-firmware sibling; an earlier note here said Meteor Lake). The enclosure’s BR hub is device-side and negotiates down. [INFERRED] Different is not risk-free: the MS-02 report above also shows Xid 79 on an integrated TB4 port.

TB5 Enclosure on a TB4 Host

Topology Comparison

Topology Link Nominal PCIe bandwidth (each way, best case) Hot-plug Cost class Linux status
TB3/TB4 eGPU USB4/TB 40G, PCIe tunnel ~32 Gb/s spec allotment; ~4 GB/s cited [SOURCED localaimaster] Yes (authorize; see hot-unplug sibling) $$ enclosure (+ ATX PSU if not bundled) Mature; Xid 79 / AER fragility reported [SOURCED MS-02 issue]
TB5 enclosure on TB4 host as TB4 same as TB4 Yes $$$ (+ ATX PSU for Core X V2) Works like TB4 [BOX/INFERRED]
TB5 host + TB5 enclosure 80G (120/40 boost for DP) PCIe Gen4 x4 = 64 Gb/s nominal Yes $$$$ (+ ATX PSU for Core X V2) Barlow Ridge host: broken/incomplete in reports (2026) [SOURCED ratatoskr, MS-02]
OCuLink (M.2/port/card) native PCIe 4.0 x4 ~64 Gb/s raw, ~7.88 GB/s (Gen4 x4 both ends) No (vendor claim) $ dock+cable+ATX Plain PCIe; standard driver [SOURCED localaimaster]; Xid 79 reported (NVIDIA #900)
PCIe x16 slot (desktop) native x16 Gen4/5 ~31.5 GB/s Gen4 x16 [SOURCED localaimaster] n/a desktop Best

Measured end-to-end tokens/s per topology on the same GPU: [UNVERIFIED] (only vendor-blog figures found; e.g. 60-80 tok/s OCuLink RTX 3090 Ollama claim is unattributed). The sm_120 sibling records one anecdotal TB4-vs-TB5 whole-system comparison (secondary blog, methodology unpublished).

LLM Workload Sensitivity

Speeds are nominal ceilings (times are floors, ratios best-case); the TB5-host column is theoretical while Linux support is unproven. The sm_120 sibling redoes the load arithmetic at an assumed ~3 GB/s TB4 throughput.

Workload Link traffic TB4 (~4 GB/s theoretical, ~3 realised per sm_120 sibling) TB5 host (~8 GB/s) OCuLink (~7.9 GB/s) Verdict
Single GPU, fully resident, decode token ID + logits only negligible negligible negligible Link irrelevant [SOURCED localaimaster]
Model load (e.g. 4.8 GB Q4 8B) weights once >=1.2 s floor ~0.6 s >=0.6 s floor [SOURCED localaimaster]; real loads disk/CPU-bound [INFERRED] Minor, once
MoE CPU-expert offload, prompt processing all CPU experts copied per batch (>=32*E/k tokens threshold; E, k undefined here, presumably total and active experts [INFERRED]; large -b/-ub amortize) slowest 2x faster copy 2x faster copy Link-bound; benefits from wider link [SOURCED HF MoE guide]
MoE offload, decode experts computed on CPU, activations only CPU/RAM-bound same same Link mostly irrelevant [INFERRED]
KV-cache spill to host per-token KV reads severe half the transfer time half the transfer time Avoid; shrink ctx/quantize KV [INFERRED]
Multi-GPU layer split activations at layer boundaries fine on PCIe 3.0 x8-class (source’s claim); TB4 is x4-class [INFERRED]; egpu-over-TB works but prefill hops via host fine fine Default; tolerates slow links [SOURCED llama.cpp multi-gpu.md; knightli search summary]
Multi-GPU tensor parallel (llama.cpp tensor, vLLM TP) per-layer all-reduce every token poor poor-moderate moderate Needs NVLink/P2P-class; experimental in llama.cpp, needs NCCL, no KV quant, no MoE [SOURCED multi-gpu.md]

Why layer split over Thunderbolt can still cost throughput: activations bounce GPU->host->GPU (P2P is generally workstation-only) and tunnel latency adds per-token. [SOURCED multi-gpu.md P2P note; INFERRED latency effect]. Claim “layer split over TB noticeably slower than single GPU” is [INFERRED]; egpu.io multi-eGPU thread was 404 - [UNVERIFIED].

Decision Guide

Rule 1: the model + KV must fit in one GPU’s VRAM => link barely matters; buy the cheapest reliable link (keep TB4; OCuLink only for the bandwidth-bound rows below, accepting no hot-plug). [BOX: RTX 5080 16 GB.] If the TB4 link itself is unstable (Xid 79), fix that first (fallen-off-bus sibling). [INFERRED] Rule 2: decide by where bytes go (MoE offload, KV spill, multi-GPU), not by “80 Gb/s”. [INFERRED]

Situation Recommendation
<=~14B dense, or a quantized model whose weights + KV cache fit in 16 GB (sm_120 sibling), fully resident Stay on TB4 [BOX current]; no upgrade justified
MoE too big for VRAM, CPU experts, long prompts OCuLink (up to 2x the nominal TB4 copy rate; expected, not measured [INFERRED]; Xid 79 reported on OCuLink, NVIDIA #900, no RTX 5080 success report found) and large -b/-ub; else accept TB4 prompt-processing cost
Want hot-plug/laptop portability TB4 now; TB5 only after a public Linux success report for your host
Two GPUs to fit a ~30B-class dense model Layer split; OCuLink x4 each or one internal (desktop only; this NUC has only the x1 Gen3 header) + one eGPU; avoid tensor parallel over any x4 link; two 16 GB cards (32 GB) cannot hold a 70B dense model at 4-bit (~35 GB) [INFERRED arithmetic]
Want tensor parallel Not with eGPU links; use desktop x16/NVLink-class or expect it to lose to layer split [INFERRED]
Buying TB5 enclosure for a TB4 host Only if future TB5 host is planned; no benefit today

For this NUC [BOX]: only if an upgrade is justified (a resident model does not), OCuLink via the 2280 Gen5 x4 M.2 slot is the best bandwidth upgrade path but forfeits hot-plug and one NVMe slot, carries the Xid 79 caveat above, and strands the Core X V2 and its ATX PSU already owned; TB5 requires replacing the host, and the Linux Barlow Ridge host stack is unproven. [INFERRED]

Anti-patterns

Sources

  1. Intel JHL9580 spec page - https://www.intel.com/content/www/us/en/products/sku/225921/intel-jhl9580-thunderbolt-5-controller/specifications.html
  2. Intel JHL9480 - https://www.intel.com/content/www/us/en/products/sku/225919/intel-jhl9480-thunderbolt-5-accessory-controller/specifications.html (listed, not fetched)
  3. TechPowerUp Barlow Ridge - https://www.techpowerup.com/318236/details-of-intels-barlow-ridge-thunderbolt-5-controller-leaks
  4. Phoronix Linux 6.5 USB4 v2 - https://www.phoronix.com/news/Linux-6.5-USB4-v2-Barlow-Ridge
  5. LWN asymmetric switching - https://lwn.net/Articles/947726/
  6. Kernel docs - https://docs.kernel.org/admin-guide/thunderbolt.html
  7. Barlow Ridge RFC patch thread - https://ratatoskr.run/linux-usb/2026/07/17208566/t
  8. MS-02-Ultra issue - https://github.com/minisforum-docs/MS-02-Ultra/issues/32
  9. egpu.io TB5 specs thread - https://egpu.io/forums/laptop-computing/thunderbolt-5-specs-confirmed80gpbs-bi-drectional-up-to-120gpbs-boost-speed-intel-demos-laptop-protoype-with-tb5-port/paged/6/
  10. Razer Core X V2 coverage - https://videocardz.com/newz/razer-core-x-v2-now-available-349-99-egpu-enclosure-with-thunderbolt-5 ; Tom’s Hardware URL above
  11. llama.cpp multi-GPU doc - https://github.com/ggml-org/llama.cpp/blob/master/docs/multi-gpu.md
  12. MoE offload guide - https://huggingface.co/blog/Doctor-Shotgun/llamacpp-moe-offload-guide
  13. eGPU for Local AI (secondary blog) - https://localaimaster.com/blog/egpu-local-ai-benchmarks
  14. ADT-Link OCuLink adapter - https://www.adt.link/product/F9GV4.html ; retail OCuLink dock listings (Amazon) as vendor claims
  15. LKML asym recovery fix - https://lkml.iu.edu/hypermail/linux/kernel/2502.0/00952.html

Unverified / gaps

JHL9586 part role; asymmetric-link release version; asym_threshold units/default; Barlow Ridge status on kernel 7.0.0-34; measured tok/s per topology; OCuLink dock retimer details; RTX 5080 power figure; Razer 140 W PD; TB4-port Gen1 vs Gen4 conflict (stated, unresolved); JHL9540 role; ~1.3% figure (estimate); Razer price; OCuLink cable length; “knightli” search summary (no URL recorded); E, k in the MoE threshold. The tunnel sibling records asym_threshold as Mb/s, default 45000.