Blackwell sm_120 local-LLM inference stack on Linux: llama.cpp, Ollama, PyTorch and vLLM over a Thunderbolt eGPU

Blackwell sm_120 local-LLM inference stack on Linux: llama.cpp, Ollama, PyTorch and vLLM over a Thunderbolt eGPU

Building and choosing the local-LLM inference stack for a consumer Blackwell GPU (sm_120, RTX 5080 16 GB) on Linux — the llama.cpp CUDA build matrix, Ollama’s bundled CUDA backends and GPU discovery, the PyTorch wheel and vLLM status, and what a ~3 GB/s Thunderbolt tunnel costs in load time and offload traffic and how to keep everything resident in 16 GB.


name: blackwell-sm120-llm-inference-stack-linux title: Local-LLM inference stack for consumer Blackwell (sm_120, RTX 5080 16 GB) on Linux description: Build/choose the CUDA inference stack for an RTX 5080 (sm_120) on Linux, incl. a ~3 GB/s Thunderbolt eGPU. TRIGGER — llama.cpp CUDA build for sm_120 (CMAKE_CUDA_ARCHITECTURES=120 vs 120a-real, GGML_NATIVE, MXFP4 ptxas errors, cudart bundles, 12.8 vs 13.4 tarballs); Ollama cuda_v12/cuda_v13 runners, GPU discovery, OLLAMA_FLASH_ATTENTION / OLLAMA_KV_CACHE_TYPE / num_gpu; PyTorch cu126-cu132 wheels and driver floors; vLLM / TensorRT-LLM / ExLlamaV3 / bitsandbytes on SM120; TB tunnel cost (load time, MoE expert offload, multi-GPU); keeping a model resident in 16 GB. SKIP — driver branches, DKMS, nvidia module params (sibling reference); Mac eGPU; KV-cache theory → kv-cache-optimization. verified-as-of: 2026-09-24

Local-LLM inference stack for consumer Blackwell (sm_120) on Linux

Verified-as-of 2026-09-24. Target box: RTX 5080 16 GB in a Thunderbolt-5-capable enclosure (Razer Core X V2) on a Thunderbolt 4 host, an Intel NUC 15 Pro, Ubuntu 26.04.1, kernel 7.0.0-34, NVIDIA driver 610.57.04-open (KMD 610 / CUDA UMD 13.3), Ollama already serving via its bundled llama-server, apt CUDA libs 12.4 (apt nvcc is too old for sm_120). Driver branches / DKMS / module parameters are covered by the sibling reference and are not repeated here.

Tagging: [SOURCED url] = read from the cited page; [INFERRED] = derived from sourced facts plus arithmetic or well-known engine behaviour; nothing else is asserted. No version or flag below is invented.


Core Concepts

1. sm_120 vs sm_120a vs sm_120f — why “120” alone breaks a native llama.cpp build

2. Toolkit ↔ driver floors and CUDA minor-version compatibility

CUDA toolkit Minimum Linux driver Source
12.4 GA ≥ 550.54.14 [SOURCED https://docs.nvidia.com/cuda/cuda-toolkit-release-notes/index.html]
12.8 GA ≥ 570.26 same
12.9 GA ≥ 575.51.03 same
13.0 R580 (≥ 580.65.06) same + [SOURCED https://docs.nvidia.com/cuda/archive/13.0.0/cuda-toolkit-release-notes/index.html]
13.1 R590 same
13.2 R595 same
13.3 R610 same
13.4 (current, 13.4 Update 1) R615 same

3. Real vs virtual architectures decide first-load latency and driver dependence

4. Where the Thunderbolt tunnel is on the critical path — and where it is not

5. 16 GB residency budget

6. MoE expert offload keeps the hot set on the GPU


Compatibility Matrix

Stack (version verified) Min CUDA to build / run sm_120 status Install path on this box Known issues
llama.cpp b11170 (2026-09-24 nightly tag; the tagged v0.5.0 release carries only nightly-tag.txt) Build: toolkit ≥ 12.8 (nvcc). Run: cuda-12.8 tarball needs driver ≥ 570.26; cuda-13.4 tarball ≥ 580 via minor-version compat (R615 for full 13.4 feature set) Native. 120a-real SASS; MXFP4 tensor-core kernels (PR #17906, +1.2–1.3× pp on MXFP4 MoE); FA (FlashAttention) on Prebuilt: llama-b11170-bin-ubuntu-cuda-12.8-x64.tar.gz (170 MB) or …-cuda-13.4-x64.tar.gz (150 MB) + matching cudart-llama-b11170-bin-ubuntu-cuda-{12.8,13.4}-x64.tar.gz (594 / 440 MB). Or build from source with a 12.8+/13.x toolkit (see recipe) #19662 (plain 120 → ptxas mxf4 errors; open), #18398 (compute_120a unsupported on an old nvcc), FA dynamic-smem ceiling on sm_120 (see Ollama #18276)
Ollama v0.33.x (0.33.0 2026-08-21) Driver ≥ 550 (docs); cuda_v13 runner implies ≥ 580; cuda_v12 runner (CUDA 12.8) ≥ 570 Supported. “Ollama is now compiled for NVIDIA Blackwell” (v0.5.13, 2025-02-27); “Support for CUDA 13” (v0.11.11, 2025-09-11); docs list CC 12.0 GeForce RTX 50xx Official installer ships /usr/local/lib/ollama/cuda_v12/ and cuda_v13/; discovery bootstraps each dir and picks one; OLLAMA_LLM_LIBRARY=cuda_v12 forces #13163 (0.12.11, RTX 5070 Ti, driver 580.95: cuda_v13 loaded then total vram="0 B" CPU fallback), #12136 (driver 580.65.06/CUDA 13.0: no nvidia devices detected by library …libcuda.so.580.65.06), #18276 (qwen3moe + auto-FA crash at warmup on sm_120; OLLAMA_FLASH_ATTENTION=0 works around), suspend/resume needs rmmod nvidia_uvm && modprobe nvidia_uvm
PyTorch 2.14.0 (2026-09-02) cu130 wheel → driver ≥ 580; cu132 → ≥ 595 (R595 = CUDA 13.2); cu126 → ≥ 560; (cu128 → ≥ 570, wheel discontinued) Supported in cu130/cu132 (arch list “Turing 7.5, Ampere 8.0/8.6, Hopper 9.0, Blackwell 10.0, 12.0”). Not in cu126 (Maxwell→Hopper only). First stable with sm_120 binaries: 2.7.0 (cu128) pip install torch (PyPI = cu130 since 2.11); --index-url https://download.pytorch.org/whl/cu132 for experimental 13.2; runtime comes from nvidia-* pip packages, apt CUDA 12.4 is irrelevant cu128 removed from the build matrix in 2.12 (week of 2026-04-06); torch.compile/Triton historically needed a Triton whose bundled ptxas knows sm_120 (Value 'sm_120' is not defined for option 'gpu-name'); a reported gap in libnvptxcompiler.so in CUDA 12.8/12.9 pip runtime packages broke PTX JIT on a 5080 (HF forum report)
vLLM v0.30.0 (2026-09-22) PyPI wheel = CUDA 13.0 (driver ≥ 580); +cu129 wheel also published (docs page still says 12.9 default — docs lag the release) Partial → improving. SM120-specific: B12X_ATTN causal paged-attention backend for SM120/121, W4A4 NVFP4 now default over weight-only kernels on SM120/121 (#55170), DeepGEMM fork with SM120 paged-MQA ports, SM12x blockwise FP8. Requires CC ≥ 7.5 pip install vllm (cu130) or uv pip install vllm --torch-backend=auto; Docker vllm/vllm-openai:v0.30.0 (-cu129 tag for 12.9) On 0.24 a ModelOpt NVFP4 checkpoint on a 5090 fell back to Marlin W4A16 with a misleading “no native FP4” warning (#47749); Nemotron NVFP4 MoE “No NvFp4 MoE backend” (#35065); SM120 uses SM80-style mma.sync, not SM100 tcgen05 — SM100 kernels crash on SM120; FlashInfer Triton GEMM configs exceed SM120’s 99 KiB (101 376 B) shared-memory limit (FlashInfer PR #5517)
FlashInfer 0.6.x cu128/cu130 wheels (see vLLM) Partial: FA2 tensor-core path and XQA NVFP4-KV decode are SM120-enabled; trtllm-gen FMHA sm120 still a feature request (#5290) pulled in by vLLM smem-limit launch failures; ongoing SM120 tracking issues (#5209, #5290)
bitsandbytes 0.50.1 (2026-08-13) CUDA 12.8+ wheels carry sm_100/sm_120 since 0.45.3; CUDA 13 CI since 0.49.0; CUDA 13.2 wheels since 0.50.0 Supported (4-bit fused dequant+GEMM kernels in 0.50.0) pip install bitsandbytes matching the torch CUDA major none Blackwell-specific found beyond the version floor
TensorRT-LLM (1.x line) Container-based; CUDA 13.x Works, thinly documented. Support matrix lists “NVIDIA Blackwell Architecture” without naming RTX 50 or FP4; FP4 GEMM on GeForce Blackwell was gated in 0.19.0 and enabled from 0.20.0rc3 (issue #5018); community RTX 5090 FP4 deployments (discussion #8334) and C++ runtime patches for 30B NVFP4 MoE on v1.2.0rc4 exist NGC container only; heavy (tens of GB) — on a 16 GB card only small NVFP4 models fit multimodal reported not working on 5090 (#8334); no official RTX 50 statement; TensorRT-for-RTX (a separate product) explicitly targets CC 12.0/12.1
ExLlamaV3 v1.5.1 (2026-09-20) torch ≥ 2.6, CUDA ≥ 12.4; wheels +cu128.torch2.8…2.11 and +cu132.torch2.11…2.13 Wheels are built with 12.8/13.2 toolkits (which support sm_120); attention/cache/recurrent kernels are Triton, so a Blackwell-aware Triton is required. README makes no explicit Blackwell statement [INFERRED] pip install exllamav3 --extra-index-url per release page, pick the wheel matching your torch none documented; torch 2.7 wheels retired in 1.4.9

Sources for this table: llama.cpp GitHub release API (b11166–b11171) and issues 19662/18398/20195; Ollama release API (v0.5.13, v0.11.11, v0.33.0), CMakePresets.json, Dockerfile, docs.ollama.com/gpu, issues 13163/12136/18276; PyTorch release v2.14.0 notes, dev-discuss posts 3325 and 3337; vLLM v0.30.0 release notes, docs installation/gpu, issues 47749/35065; FlashInfer issues 5290/5209 and PR 5517; bitsandbytes releases 0.49.0/0.50.0/0.50.1; TensorRT-LLM issue 5018, discussion 8334, support-matrix page; ExLlamaV3 release API. Full URLs in Sources.


llama.cpp Build

Option A — prebuilt (fastest, no nvcc needed)

# pick ONE CUDA line and keep binary + cudart from the SAME build number and CUDA version
B=b11170
curl -LO https://github.com/ggml-org/llama.cpp/releases/download/$B/llama-$B-bin-ubuntu-cuda-12.8-x64.tar.gz
curl -LO https://github.com/ggml-org/llama.cpp/releases/download/$B/cudart-llama-$B-bin-ubuntu-cuda-12.8-x64.tar.gz
mkdir -p ~/llama-cuda && tar -C ~/llama-cuda -xzf llama-$B-bin-ubuntu-cuda-12.8-x64.tar.gz \
                       && tar -C ~/llama-cuda -xzf cudart-llama-$B-bin-ubuntu-cuda-12.8-x64.tar.gz
# the cudart tarball supplies libcudart.so.12 / libcublas.so.12 / libcublasLt.so.12 next to the binaries,
# so the apt CUDA 12.4 libs are never used

Asset names as published on 2026-09-24: llama-b11170-bin-ubuntu-cuda-12.8-x64.tar.gz, llama-b11170-bin-ubuntu-cuda-13.4-x64.tar.gz, cudart-llama-b11170-bin-ubuntu-cuda-12.8-x64.tar.gz, cudart-llama-b11170-bin-ubuntu-cuda-13.4-x64.tar.gz (Windows: cuda-12.4 and cuda-13.4). [SOURCED https://api.github.com/repos/ggml-org/llama.cpp/releases]

Option B — from source (needed for GGML_CUDA_FA_QUANTS, NCCL, custom flags)

Prerequisite: a CUDA toolkit ≥ 12.8 (12.8 is the floor for sm_120; 13.x is fine on driver 610). Do not use the apt nvidia-cuda-toolkit 12.4 — install the toolkit from NVIDIA’s repo or the runfile with --toolkit only (driver handling is in the sibling reference).

git clone https://github.com/ggml-org/llama.cpp && cd llama.cpp
cmake -B build \
  -DGGML_CUDA=ON \
  -DCMAKE_CUDA_COMPILER=/usr/local/cuda/bin/nvcc \
  -DCMAKE_CUDA_ARCHITECTURES="120a-real" \
  -DGGML_NATIVE=ON \
  -DGGML_CUDA_FA_QUANTS=all \
  -DCMAKE_BUILD_TYPE=Release
cmake --build build --config Release -j"$(nproc)" --target llama-server llama-cli llama-bench llama-gguf-split

Runtime sanity

~/llama-cuda/llama-bench -m model.gguf -ngl 99 -fa 1 -p 512 -n 128
# expect ggml_cuda_init: found 1 CUDA devices: Device 0: NVIDIA GeForce RTX 5080, compute capability 12.0

Reference points from the community CUDA benchmark thread (7B Q4_0, pp512/tg128): RTX 5080 8297→9488 pp, 182→185 tg (pairs read as FA off → on [INFERRED]); RTX 5090 14073/290 → 14970/300; RTX 4090 11993/186 → 14771/189. [SOURCED https://github.com/ggml-org/llama.cpp/discussions/15013]


Ollama

How the bundled runners are chosen

Env knobs (from envconfig/config.go, default in parentheses)

Variable Default Effect
OLLAMA_FLASH_ATTENTION auto (FA is enabled “automatically when the selected backend and devices support it”); 1 forces on, 0 forces off attention scratch stays flat with context; prerequisite for KV quantization
OLLAMA_KV_CACHE_TYPE f16 q8_0 ≈ ½ KV memory, “very small loss”; q4_0 ≈ ¼, “small-medium loss … more noticeable at higher context”; global, applies to all models
OLLAMA_CONTEXT_LENGTH 0 = auto (“4k/32k/256k based on VRAM” per code; FAQ says 4096) default num_ctx
OLLAMA_KEEP_ALIVE 5m negative value keeps the model loaded forever (avoids a TB reload every idle period)
OLLAMA_GPU_OVERHEAD 0 bytes reserve VRAM per GPU for other processes
OLLAMA_MAX_LOADED_MODELS / OLLAMA_NUM_PARALLEL / OLLAMA_MAX_QUEUE 0 / 1 / 512 scheduler limits
OLLAMA_SCHED_SPREAD false force spreading across all GPUs
OLLAMA_LLM_LIBRARY unset force cuda_v12, cuda_v13, vulkan, cpu_avx2…
OLLAMA_VULKAN true disable to stop Vulkan probing the same GPU
OLLAMA_DEBUG 0 1 for discovery/load logs
CUDA_VISIBLE_DEVICES unset device visibility
[SOURCED https://raw.githubusercontent.com/ollama/ollama/main/envconfig/config.go; https://raw.githubusercontent.com/ollama/ollama/main/docs/faq.mdx]

Verification checklist (run as the user; nothing here changes state)

  1. ollama --version → ≥ 0.11.11 for a cuda_v13 runner; ≥ 0.5.13 for any Blackwell SASS. [SOURCED release notes]
  2. journalctl -u ollama --since -10m | grep -E 'load_backend|inference compute|total vram|library=' → expect load_backend: loaded CUDA backend from /usr/local/lib/ollama/cuda_v13/libggml-cuda.so (or cuda_v12) and an inference compute line with library=cuda and compute=12.0; failure signature is library=cpu plus "total vram"="0 B" / entering low vram mode. [SOURCED issue 13163 log lines]
  3. nvidia-smi --query-gpu=name,memory.used,driver_version --format=csv while a model is loaded → memory.used ≈ model size + KV.
  4. ollama ps → 100% GPU, not 48%/52% CPU/GPU. [SOURCED FAQ]
  5. If discovery fails after suspend/resume: sudo rmmod nvidia_uvm && sudo modprobe nvidia_uvm. [SOURCED docs.ollama.com/gpu]
  6. A/B the runner: OLLAMA_LLM_LIBRARY=cuda_v12 ollama serve vs cuda_v13; on a Blackwell CPU-fallback (#13163 pattern) the v12 runner with real sm_120 SASS is the first thing to try. [INFERRED]
  7. MoE model + FA crash at warmup (CUDA error: shared object initialization failed at fattn-mma-f16.cuh) → OLLAMA_FLASH_ATTENTION=0 (global) until fixed; the kernel asked for more dynamic shared memory than sm_120 allows. [SOURCED https://github.com/ollama/ollama/issues/18276]
  8. Confirm KV quantization took effect: the load log prints the cache type; if FA is off the setting is silently ignored. [SOURCED FAQ (“cache is only quantized when flash attention is enabled”)]
  9. Set OLLAMA_KEEP_ALIVE=-1 in the systemd override so the ~4 s TB reload (see Thunderbolt Bandwidth Costs) doesn’t recur after every 5-minute idle.

PyTorch & vLLM

PyTorch wheel matrix (2.14.0)

Wheel Arch list Driver floor Status
cu126 5.0, 6.0, 7.0, 7.5, 8.0, 8.6, 9.0 — no 12.0 ≥ 560 legacy
cu128 had 12.0 (first stable sm_120 binaries in 2.7.0) ≥ 570 removed from the matrix in 2.12 (April 2026)
cu130 7.5, 8.0, 8.6, 9.0, 10.0, 12.0 ≥ 580 PyPI default since 2.11
cu132 7.5, 8.0, 8.6, 9.0, 10.0, 12.0 R595+ (CUDA 13.2 branch) experimental since 2.12
[SOURCED https://dev-discuss.pytorch.org/t/transitioning-pypi-cuda-wheels-to-cuda-13-0-as-the-stable-release-2-11/3325; https://dev-discuss.pytorch.org/t/introducing-cuda-13-2-and-deprecating-cuda-12-8-release-2-12/3337; https://github.com/pytorch/pytorch/issues/159207; CUDA release-notes driver table]

vLLM 0.30.0 on a 16 GB SM120 card

TensorRT-LLM / ExLlamaV3 (short)


Thunderbolt Bandwidth Costs

Assumptions: TB4 tunnel assumed ≈ 3 GB/s on this box (figure supplied by the box owner, not measured here; theoretical PCIe 3.0 x4 ≈ 4 GB/s), vs ≈ 25 GB/s realised on internal PCIe 4.0 x16 (31.5 GB/s theoretical). Formulae from [SOURCED https://localaimaster.com/blog/egpu-local-ai-benchmarks]; numbers below are arithmetic on those formulae plus published model sizes — [INFERRED] unless noted.

Operation Bytes crossing the tunnel Time @ 3 GB/s (TB4) Time @ 25 GB/s (x16) Verdict
Load 8B Q4_K_M (4.8 GB) 4.8 GB, once ≥ 1.6 s ≥ 0.2 s noticeable only if reloaded often
Load gpt-oss-20b MXFP4 (~12.1 GB) ~12 GB, once ≥ 4 s ≥ 0.5 s set OLLAMA_KEEP_ALIVE=-1
Load Qwen3.5-35B-A3B UD-IQ3_S (13.6 GB) 13.6 GB, once ≥ 4.5 s ≥ 0.55 s same
Decode, all weights resident ≤ 513 KB/token (worst case, fp32 logits of a 128 K vocab) → 51 MB/s @ 100 tok/s ~1.7 % of link ~0.2 % not link-bound [SOURCED localaimaster]
Prefill, all weights resident token embeddings in, logits out; KV stays on GPU negligible negligible not link-bound; RTX 5080 pp512 ≈ 9.5 k tok/s (7B Q4_0)
Decode with --n-cpu-moe N per offloaded layer per token ≈ 2 × n_embd × 2 B of activations (e.g. Qwen3-30B-A3B n_embd 2048 → ~8 KB/layer, ~0.4 MB/token for all 48 layers) < 1 % of link — limiter is host DDR5 bandwidth and CPU FLOPs, not TB
Prefill with --n-cpu-moe N at large -ub llama.cpp may stream CPU-resident expert weights to the GPU per micro-batch for batched matmul; worst case = all offloaded expert bytes per ubatch (e.g. 8 GB of experts → ≥ 2.7 s per ubatch) seconds per ubatch 0.3 s link-bound; the only common path where TB really hurts. Mitigate with smaller -ub, fewer offloaded layers, or accept slower prefill [INFERRED] — the 5080-llm-configs repo notes “PCIe streaming efficiency” is the real question on a 16 GB 5080
Two eGPUs on one TB controller, -sm layer activations at each layer boundary (KB/token) small — fine for decode; both GPUs share 40 Gbps
Two eGPUs, -sm row / tensor-parallel all-reduce per layer (MB/token) link-bound — avoid over TB
KV cache never leaves the GPU for GPU layers 0 0 —
Model swap (two models alternating under OLLAMA_MAX_LOADED_MODELS=1) full reload each swap ≥ 4 s per swap for 12 GB — raise OLLAMA_MAX_LOADED_MODELS only if both fit

Whole-system reference points (not decomposed): RTX 5080 TB4 ≈ 85 % of internal throughput, TB5 ≈ 95 % on a 70B Q4 workload that is itself partly CPU-offloaded on a 16 GB card — consistent with the offload-prefill row above. [SOURCED botmonster] Disk is often the real load bottleneck: a cold GGUF read from NVMe (3–7 GB/s) is on the same order as the tunnel; a page-cache-warm second load is tunnel-bound. [INFERRED]


16 GB Residency

Goal: weights + KV + scratch ≤ ~15 GiB so nothing spills to CPU and nothing re-crosses the tunnel after load.

Model classes that stay resident on a 16 GB RTX 5080 (published measurements):

Model / quant Weights Context that fit Speed Source
gpt-oss-20b MXFP4 (ggml-org GGUF) ~12 GB --ctx-size 32768 -ub 4096 -b 4096 on 16 GB (llama.cpp guide) ~89 tok/s est. on 5080 [SOURCED https://github.com/ggml-org/llama.cpp/discussions/15396; modelfit.io]
Qwen3.5-35B-A3B UD-IQ3_S 13.6 GB ~100 K (RTX 4080 16 GB) 136–138 tok/s at 19–64 K [SOURCED https://www.glukhov.org/llm-performance/benchmarks/best-llm-on-16gb-vram-gpu/]
Qwen3.5-27B UD-IQ3_XXS (dense) 11.5 GB 128 K (9.6 tok/s there); 45 tok/s ≤ 32 K — same
Qwen3.8-27B UD-IQ3_S (dense, vision) 12 GB 98 304 95–96 tok/s decode w/ speculative, 1 756 tok/s prefill [SOURCED 5080-llm-configs repo, URL in Sources]
14B Q4_K_M ~8.4 GB 32 K f16 KV comfortably ~58 tok/s [SOURCED modelfit.io/gpu/rtx-5080]
8B Q4_K_M ~4.8 GB 64 K+ ~45 tok/s (Llama 3.1 8B, localscore) [SOURCED https://www.localscore.ai/accelerator/489]

llama.cpp flags that keep it resident

llama-server -m model.gguf -ngl 99 -fa on -c 32768 -ctk q8_0 -ctv q8_0 -ub 512 -b 2048 --jinja
# MoE that does not fit: keep attention/router/KV on GPU, experts partly on CPU
llama-server -m qwen3-30b-a3b-Q4_K_M.gguf -ngl 99 --n-cpu-moe 12 -fa on -c 32768 -ctk q8_0 -ctv q8_0

Ollama equivalents

# /etc/systemd/system/ollama.service.d/override.conf
[Service]
Environment="OLLAMA_FLASH_ATTENTION=1"
Environment="OLLAMA_KV_CACHE_TYPE=q8_0"
Environment="OLLAMA_KEEP_ALIVE=-1"
Environment="OLLAMA_CONTEXT_LENGTH=32768"
Environment="OLLAMA_MAX_LOADED_MODELS=1"

plus PARAMETER num_gpu 999 (Modelfile) or "options":{"num_gpu":999,"num_ctx":32768} (API) when Ollama’s automatic estimate under-offloads; verify with ollama ps → 100% GPU. If Ollama insists on splitting, lower num_ctx or the KV type before touching num_gpu. [SOURCED FAQ; api/types.go] Ollama’s own MoE offload is automatic (it has no --n-cpu-moe knob); for fine-grained expert placement use llama-server directly. [INFERRED] from the absence of such an option in envconfig and api/types.go.

Rules of thumb [INFERRED]


Anti-patterns

  1. -DCMAKE_CUDA_ARCHITECTURES=120 on a current tree → ptxas … '.kind::mxf4' not supported on .target 'sm_120'; use 120a-real (or let native rewrite it). -DGGML_CUDA_MXFP4=OFF does not prevent the template instantiation. [SOURCED issue 19662]
  2. apt nvcc 12.4 → Unsupported gpu architecture 'compute_120'; the floor is toolkit 12.8. [SOURCED issue 20195]
  3. GGML_NATIVE=ON with the eGPU unplugged/asleep at configure time → CMake’s native list is “garbage” and the build targets nothing useful; configure with the GPU attached or pass the arch list explicitly. [SOURCED CMakeLists comment]
  4. Mixing a cuda-13.4 binary with the cuda-12.8 cudart bundle (or with the apt 12.4 libs) → libcudart.so.13 not found or symbol errors; keep build number and CUDA major identical. [SOURCED release asset naming] [INFERRED] for the failure mode.
  5. Old-toolkit-built binary + new PTX: a PTX-only build (Ollama cuda_v13 is all -virtual) depends on the driver’s JIT; on a lower driver than the PTX ISA it fails. On driver 610 this is fine for CUDA 13.0 PTX, but not for a 13.4-era PTX-only artefact. [SOURCED minor-version compatibility doc; CMakePresets]
  6. OLLAMA_NUM_GPU=… — not an upstream variable; it does nothing (issue #11437 documents the resulting confusion). Use num_gpu as an option. [SOURCED envconfig; api/types.go]
  7. OLLAMA_KV_CACHE_TYPE=q8_0 with FA off → silently ignored, KV stays f16 and the model spills. [SOURCED FAQ]
  8. Leaving FlashAttention on (auto or forced =1) for a qwen3moe-class model on sm_120 → warmup crash (shared object initialization failed, dynamic shared memory ceiling). Set OLLAMA_FLASH_ATTENTION=0 (global, so it also disables KV quantization) or wait for the fix. [SOURCED issue 18276]
  9. pip install torch --index-url …/cu126 on Blackwell → sm_120 is not compatible with the current PyTorch installation; cu126 has no 12.0 SASS. Use cu130 (default) or cu132. [SOURCED dev-discuss 3337]
  10. Assuming SM100 kernels/flags apply to SM120 (tcgen05, datacenter NVFP4 MoE backends) → compile or runtime failure; consumer Blackwell is mma.sync + 99 KiB smem. [SOURCED vLLM PR 50288; FlashInfer PR 5517]
  11. Large -ub with heavy --n-cpu-moe over a TB link → prefill becomes a weight-streaming exercise at 3 GB/s; shrink -ub or offload fewer layers. [INFERRED]
  12. Default OLLAMA_KEEP_ALIVE=5m on an eGPU → a 12 GB model re-crosses the tunnel after every idle gap (≥ 4 s each time). Set -1. [INFERRED] from the load-time formula.
  13. Trusting 2025 search snippets (“no prebuilt SM120 binaries”, “stable torch stops at sm_90”, “vLLM lacks 5090 support”) — all superseded; check the dated sources above.
  14. Running vLLM beside a loaded Ollama model on 16 GB → allocator collision; vLLM reserves 90 % by default. [INFERRED]

Related references added later: local-llm-model-load-path-over-thunderbolt-linux.md (model load path and keep-resident policy, in the ai-llm-model-layer hub).

Sources

llama.cpp

Ollama

NVIDIA CUDA

PyTorch / vLLM / FlashInfer / bitsandbytes / TRT-LLM / ExLlamaV3

Thunderbolt / 16 GB