Linux eGPU for LLM Inference (Thunderbolt 3/4/5, USB4, OCuLink)
Parent: Local LLM Inference on Consumer Hardware (RTX 5080 eGPU, Linux, Remote Access) · Topic entry · 8 branches · skill ai-llm-model-layer
A published reference is not available for this topic yet.
Children
- Hybrid MoE prefill vs link bandwidth benchmark (x16 vs OCuLink vs TB5 vs TB4) (frontier)
- llama.cpp op-offload threshold and ubatch tuning on slow links (frontier)
- NVIDIA open-module eGPU patches (TB4/TB5 detection, hot-unplug) (frontier)
- BAR1 / Resizable BAR sizing over Thunderbolt tunnels (frontier)
- TB5 controller (Barlow Ridge) vs ASM2464PD USB4 bridge behavior (frontier)
- OCuLink signal integrity and AER / replay counters (frontier)
- Blackwell power-cap curves: prefill vs decode per watt (frontier)
- Tensor-parallel over TB/OCuLink links (frontier)
Frontier under this node: BAR1 / Resizable BAR sizing over Thunderbolt tunnels, Blackwell power-cap curves: prefill vs decode per watt, Hybrid MoE prefill vs link bandwidth benchmark (x16 vs OCuLink vs TB5 vs TB4), NVIDIA open-module eGPU patches (TB4/TB5 detection, hot-unplug), OCuLink signal integrity and AER / replay counters, TB5 controller (Barlow Ridge) vs ASM2464PD USB4 bridge behavior, Tensor-parallel over TB/OCuLink links, llama.cpp op-offload threshold and ubatch tuning on slow links