Blackwell sm_120 local-LLM inference stack on Linux: llama.cpp, Ollama, PyTorch and vLLM over a Thunderbolt eGPU

Parent: Thunderbolt eGPU on Linux for local LLM inference · Published reference · snapshot 2026-09-08

Also known as: cuda 13 llama.cpp, egpu llm bandwidth, local llm 16gb vram, ollama blackwell, pytorch blackwell wheels, rtx 5080 llama.cpp, sm_120, vllm rtx 50

↓ Facts as markdown↓ Download this reference fileall context files

Building and choosing the local-LLM inference stack for a consumer Blackwell GPU (sm_120, RTX 5080 16 GB) on Linux — the llama.cpp CUDA build matrix, Ollama's bundled CUDA backends and GPU discovery,

These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.

Blackwell sm_120 local-LLM inference stack on Linux: llama.cpp, Ollama, PyTorch and vLLM over a Thunderbolt eGPU

Local-LLM inference stack for consumer Blackwell (sm_120) on Linux

1. `sm_120` vs `sm_120a` vs `sm_120f` — why "120" alone breaks a native llama.cpp build

2. Toolkit ↔ driver floors and CUDA minor-version compatibility

3. Real vs virtual architectures decide first-load latency and driver dependence

4. Where the Thunderbolt tunnel is on the critical path — and where it is not

5. 16 GB residency budget

6. MoE expert offload keeps the *hot* set on the GPU

Compatibility Matrix

Option A — prebuilt (fastest, no nvcc needed)

Option B — from source (needed for `GGML_CUDA_FA_QUANTS`, NCCL, custom flags)

Runtime sanity

How the bundled runners are chosen

Env knobs (from `envconfig/config.go`, default in parentheses)

Verification checklist (run as the user; nothing here changes state)

PyTorch wheel matrix (2.14.0)

vLLM 0.30.0 on a 16 GB SM120 card

TensorRT-LLM / ExLlamaV3 (short)

Thunderbolt Bandwidth Costs

16 GB Residency

Anti-patterns

  • Related references added later: local-llm-model-load-path-over-thunderbolt-linux.md (model load path and keep-resident policy, in the ai-llm-model-layer hub). [source]
  • Sources

  • PyTorch / vLLM / FlashInfer / bitsandbytes / TRT-LLM / ExLlamaV3 [source]
    • https://dev-discuss.pytorch.org/t/transitioning-pypi-cuda-wheels-to-cuda-13-0-as-the-stable-release-2-11/3325 (2026-03-11) [source]
    • https://dev-discuss.pytorch.org/t/introducing-cuda-13-2-and-deprecating-cuda-12-8-release-2-12/3337 - per-wheel arch lists, cu128 removal [source]
    • https://api.github.com/repos/pytorch/pytorch/releases/latest - v2.14.0 (2026-09-02) notes [source]
    • https://github.com/pytorch/pytorch/issues/159207 - sm_120 support history (2.7.0 first stable) [source]
    • https://discuss.pytorch.org/t/rtx-5070-ti-blackwell-pytorch-nightly-triton-still-getting-sm-120-is-not-defined-for-option-gpu-name-error/220460 [source]
    • https://discuss.huggingface.co/t/ptx-jit-broken-on-rtx-5080-blackwell-sm-120-missing-libnvptxcompiler-so-in-cuda-12-8-12-9/161827 (title only; page 403 at verification) [source]
    • https://api.github.com/repos/vllm-project/vllm/releases/latest - v0.30.0 (2026-09-22): CUDA 13.0 PyPI, cu129 wheels, B12X_ATTN, NVFP4 W4A4 default on SM120/121 [source]
    • https://docs.vllm.ai/en/latest/getting_started/installation/gpu.html - CC ≥ 7.5, wheel variants [source]
    • https://github.com/vllm-project/vllm/pull/50288, https://github.com/vllm-project/vllm/pull/46329, https://github.com/vllm-project/vllm/issues/47749, https://github.com/vllm-project/vllm/issues/35065, https://github.com/vllm-project/vllm/issues/31085 [source]
    • https://github.com/flashinfer-ai/flashinfer/pull/5517, https://github.com/flashinfer-ai/flashinfer/issues/5290, https://github.com/flashinfer-ai/flashinfer/issues/5209 [source]
    • https://api.github.com/repos/bitsandbytes-foundation/bitsandbytes/releases - 0.49.0, 0.50.0, 0.50.1 [source]
    • https://github.com/NVIDIA/TensorRT-LLM/issues/5018, https://github.com/NVIDIA/TensorRT-LLM/discussions/8334, https://github.com/JohnTDI-cpu/trtllm-nvfp4-blackwell-fix, https://nvidia.github.io/TensorRT-LLM/reference/support-matrix.html [source]
    • https://api.github.com/repos/turboderp-org/exllamav3/releases and https://github.com/turboderp-org/exllamav3 [source]
    • https://localaimaster.com/blog/egpu-local-ai-benchmarks - link bandwidth table, load-time formula, per-token traffic arithmetic (explicitly theoretical) [source]
    • https://botmonster.com/self-hosting/best-egpu-enclosures-linux-2026/ - RTX 5080 TB4 ≈ 85 %, TB5 ≈ 95 % of internal [source]
    • https://github.com/BFinn/5080-llm-configs - measured 16 GB RTX 5080 configs [source]
    • https://www.glukhov.org/llm-performance/benchmarks/best-llm-on-16gb-vram-gpu/ - 16 GB model/context/speed table [source]
    • https://www.localscore.ai/accelerator/489, https://modelfit.io/gpu/rtx-5080/ - 5080 tok/s reference points [source]
  • Children

    Frontier under this node: 16 GB VRAM residency budgeting (weights + KV cache), MoE expert CPU offload traffic over a Thunderbolt tunnel, Ollama cuda_v12/cuda_v13 backend selection and NVML dependency, PyTorch cu126/cu128/cu130 wheel matrix and driver floors, llama.cpp CMAKE_CUDA_ARCHITECTURES for sm_120 and MXFP4/NVFP4 kernels, vLLM consumer-Blackwell kernels (NVFP4, FlashInfer, bitsandbytes gaps)

    ← the whole tree · 3D view· how to read this page