Local LLM model load path over a Thunderbolt eGPU

Local LLM Model Load Path Over a Thunderbolt eGPU (Linux) and Keep-Warm Policy

What happens between a model file on NVMe and weights resident in VRAM when the host link is slow: the stages and which one bounds cold and warm loads, llama.cpp load modes, Ollama and LM Studio keep-alive and eviction, page-cache prewarming, formats and loaders, and a cold-start budget with a keep-warm decision guide.

verified-as-of 2026-09-25 (documentation and source reading only; no rate here was measured on the box, and no flag or command was run against the real GPU or services). Box: RTX 5080 with 16 GB of VRAM (video memory), Razer Core X V2 enclosure (it has no built-in power supply; the owner installs an ATX power supply unit, PSU, see egpu-power-enclosure-and-thermals-linux.md), Thunderbolt 4 (TB4) port of an Intel NUC 15 Pro (vendor spec: PCIe (PCI Express) tunnelling 32 Gbps, “PCIe 3.0 x4 compliant” [UNVERIFIED: vendor page not re-read here; the 32 Gb/s figure is discussed in thunderbolt-usb4-pcie-tunnel-bolt-iommu-linux.md]), Ubuntu 26.04.1, driver 610.57.04-open, Ollama and LM Studio headless. The “~3 GB/s” host-link figure is an UNMEASURED estimate from secondary sources [UNVERIFIED] (the sibling files call it owner-supplied); replace it with your own number from measuring-a-thunderbolt-egpu-bandwidth-and-inference-linux.md.

Read this first. Every rate and time here is an assumption or arithmetic, not a measurement from this box: the ~3 GB/s Thunderbolt link, the 25 GB/s internal-PCIe comparison, the 3 GB/s NVMe (PCIe-attached SSD) and 0.5 GB/s SATA (older disk interface) disk rates, and every second-count built on them [UNVERIFIED]. Every flag name and default was read from documentation or source and never run against the real services: confirm each with the installed binary’s --help. A claim that could not be confirmed says so in its own sentence and carries an [UNVERIFIED] or [INFERRED] tag. Replace every figure with your own measurement (procedure: measuring-a-thunderbolt-egpu-bandwidth-and-inference-linux.md, devops-linux-internals hub).

Scope: only the LOAD path and residency policy. Siblings are cross-referenced by exact filename and not duplicated here:

Core Concepts

  1. Load time is a pipeline, so the slowest stage bounds it. Weights travel disk -> page cache -> staging buffers -> PCIe / Thunderbolt tunnel -> VRAM. Stages overlap only partly; the tunnel (~3 GB/s, unmeasured) is normally about as fast as a good NVMe read or slower, so the warm-cache load is link-bound and the cold load is bounded by the slower of disk and link plus non-overlap. [INFERRED] In this file “link” and “tunnel” both mean the PCIe tunnel carried over Thunderbolt; thunderbolt-usb4-pcie-tunnel-bolt-iommu-linux.md separates the 40 Gb/s physical link from the ~32 Gb/s PCIe tunnel inside it.
  2. Cold vs warm is a page-cache property, not a GPU property. A second load is fast when the file is still in RAM; the tunnel cost never disappears because VRAM is emptied on unload. [INFERRED] Terms used here (the measuring sibling uses the same three): cold = file not in the page cache (the kernel’s RAM cache of file data), so disk is read; warm = file in the page cache but model not in VRAM, so only the RAM-to-VRAM copy remains; hot = model already resident in VRAM, so no load happens. “Keep-warm” in this file’s title means keeping the model hot; the decision guide below says “keep resident” so it does not clash with “warm” (page cache).
  3. mmap (memory-mapped file I/O) makes loads look fast or slow depending on what you measure. With mmap the kernel maps the file into the process address space and reads pages on demand, so the loader may return before pages are read; cost then appears as slow time to first token (TTFT) or as page faults during upload. llama.cpp itself notes mmap “may report incorrect progress on some platforms”. [SOURCED https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md]
  4. Residency is a policy. Keeping the model in VRAM (Ollama keep-alive, LM Studio TTL (time to live: idle time before unload), no idle sleep) converts a multi-second per-request tax into a memory reservation. Ollama, LM Studio and llama-server each expose a different knob (see sections below).
  5. Bytes moved = file size, not parameter count. Quantization (storing weights in fewer bits) and file size set the numerator; link rate is the denominator. Smaller quant = proportionally faster load. [INFERRED]
  6. Load speed is not inference speed. Once weights are resident, the link is largely idle for fully GPU-resident models (per-token traffic arithmetic: blackwell-sm120-llm-inference-stack-linux.md) [INFERRED]; do not conflate a slow load with slow tokens/s (see the measuring sibling).

The Load Path Stages

Stage What happens Bound by Notes
1 Disk read NVMe or SATA read of a GGUF (llama.cpp’s single-file model format) or safetensors (tensor file format used by vLLM) file device sequential read rate Cold only. Queue depth and readahead matter for mmap page-fault access. [INFERRED]
2 Page cache Kernel keeps file pages in free RAM RAM size, eviction Warm loads start here. drop_caches empties it (see Page Cache and Storage).
3 Host staging Runtime copies/maps tensors into host buffers (pinned, if the backend uses them) CPU memcpy, RAM bandwidth, pinning limits Exact buffering is backend-specific. Whether ggml’s CUDA backend stages through pinned buffers, and how large they are, was not verified: [UNVERIFIED] for ggml CUDA internals here. Pinned vs pageable copy rates: measuring sibling; host RAM bandwidth: ddr5-host-tuning-cpu-moe-inference.md.
4 Tunnel DMA (direct memory access) over Thunderbolt-tunnelled PCIe to the GPU link rate (spec 32 Gbps tunnel; ~3 GB/s practical estimate, unmeasured) Usually the warm-load bottleneck on this box. [INFERRED]
5 VRAM Tensors resident; KV (key/value attention) cache and compute buffers allocated VRAM capacity (16 GB) If weights + KV do not fit, layers stay on CPU (-ngl), which changes runtime speed, not just load.

Which stage bounds:

llama.cpp Loading Behaviour

All flags below are from the current master server README, not from an installed build; released builds may differ, so run llama-server --help on the installed build to confirm each option name [SOURCED https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/README.md].

Ollama and LM Studio Residency

Ollama [SOURCED https://raw.githubusercontent.com/ollama/ollama/main/docs/faq.mdx, https://raw.githubusercontent.com/ollama/ollama/main/docs/api.md, https://raw.githubusercontent.com/ollama/ollama/main/envconfig/config.go]:

LM Studio [SOURCED https://lmstudio.ai/docs/developer/core/ttl-and-auto-evict]:

Headless rule: on this box treat any of these idle timers as a decision, not a default. The default 5m Ollama keep-alive means a service idle for six minutes pays the full reload.

Page Cache and Storage

Formats and Loaders

Cold-Start Budget (arithmetic on assumed rates, not measurements)

Nothing in this section was measured. Each cell is a model size divided by an assumed rate, using the formulas below. Replace every rate with your own measurement (link and disk: measuring-a-thunderbolt-egpu-bandwidth-and-inference-linux.md, devops-linux-internals hub) and recompute. Read the seconds as order of magnitude only.

Formulas (S = file size in GB, B = rate in GB/s; weights only, excludes KV/compute-buffer allocation and runtime init, which add a fixed overhead you must measure):

Assumptions [INFERRED, not measured]: B_link(Thunderbolt) = 3 GB/s; B_link(internal PCIe) = 25 GB/s (assumed practical rate for a x16 Gen4-class slot); B_disk(NVMe) = 3 GB/s; B_disk(SATA SSD) = 0.5 GB/s. Sizes S are typical file sizes for 16 GB-card model classes (approximate; check the actual GGUF). All four rates are unmeasured assumptions [UNVERIFIED], with these provenance notes:

All cells are seconds.

Model class (file S) Warm, Thunderbolt: S/3 Cold, NVMe, Thunderbolt: overlapped S/3 / serial S/3+S/3 Cold, SATA, Thunderbolt, serial: S/0.5+S/3 Warm, internal PCIe: S/25 Cold, NVMe, internal PCIe, serial: S/3+S/25
~8B Q4 (5 GB) 1.7 1.7 / 3.3 11.7 0.2 1.9
~14B Q4 (9 GB) 3.0 3.0 / 6.0 21.0 0.4 3.4
~24B-class Q4 (14 GB) 4.7 4.7 / 9.3 32.7 0.6 5.2
Just-fits Q4-Q5 (15 GB) 5.0 5.0 / 10.0 35.0 0.6 5.6

Reading it (every figure is an output of the assumptions above): a warm 9 GB reload is about 3 s, which is small next to prompt processing for long contexts but visible as a per-request tax; cold NVMe is 1x to 2x warm (overlapped vs serial); a slow disk dominates everything. With an NVMe as fast as the assumed link, overlapped cold loads over internal PCIe are disk-bound too (S/3, the same as the overlapped Thunderbolt cell; the internal column above shows only the serial bound). The tunnel penalty in seconds is S/3 - S/25 (about 2.6 s for 9 GB) in the warm and serial-cold cases and zero in the overlapped-cold case; relative to total time it is largest on warm loads. Sensitivity: if the measured link is 2 GB/s instead of 3, the Thunderbolt warm and overlapped-cold cells scale by 1.5 (S/2 instead of S/3); the serial and SATA cells do not (S/3+S/2 is only 1.25x the serial NVMe cell, and SATA is dominated by S/0.5), so recompute them from the formulas. All figures [INFERRED]; overwrite with measurements.

Residency Decision Guide (keep resident vs reload)

For a headless service. The load costs quoted here come from the unmeasured budget above, so redo item 6 with your own measured t_load before you commit to a policy.

  1. Requests arrive at least once per idle-timeout window, single model: keep resident. Set Ollama keep_alive to -1 (or a long duration) and LM Studio TTL long; leave llama-server --sleep-idle-seconds at default -1. These change a running service: they are OPTIONAL, state-changing steps, and each undo is listed under “Changing residency settings” below. Cost: the model’s VRAM (weights plus KV cache) stays reserved and cannot be shared with other GPU work; a just-fits model leaves almost nothing free on a 16 GB card.
  2. Bursty use, long idle gaps, GPU wanted for other work: reload per request or short keep-alive; accept t_warm (about 2-5 s by the unmeasured budget) or t_cold on first request; use a scheduled preload (empty request; OPTIONAL and state-changing, see “Changing residency settings”) shortly before expected use.
  3. Multiple models that do not co-fit in 16 GB: you will swap; make RAM large enough that all files stay page-cached (warm swaps), and set OLLAMA_MAX_LOADED_MODELS/--models-max consistently with VRAM reality. Otherwise every switch is a cold load.
  4. Latency SLO (service-level objective) tighter than t_warm + prefill: keep resident; fix eviction causes instead (competing models, restart).
  5. Reliability concern (eGPU drop, sleep/resume): after the GPU falls off the bus or the host resumes, VRAM is gone and the next request pays a full reload; plan health checks accordingly (see egpu-health-monitoring-and-automated-recovery-linux.md and egpu-suspend-resume-and-sleep-states-linux.md, both in the devops-linux-internals hub).
  6. Tie-break: the two sides use different units, so this is a judgment, not a formula. Reload tax = (loads per hour without keep-resident) x t_load, in seconds of added wait per hour; loads per hour is the number of idle gaps longer than the timeout, and t_load is your measured warm or cold load time. Price of keeping resident = idle watts and their cost (egpu-idle-power-and-energy-accounting-linux.md, devops-linux-internals hub) plus the VRAM the model and KV cache occupy. Neither side converts into the other, so decide by judgment; as a starting rule of thumb (not derived from the two quantities above), if the loads happen fewer than once per few hours, reload wins; otherwise keep resident. [INFERRED heuristic]

Changing residency settings (OPTIONAL, state-changing; undo beside each)

Read-only Measurement Snippet (time to first token, TTFT, after load)

Config-read-only, not state-neutral. It changes no setting and downloads nothing, but it sends a real inference request: if the model is not loaded, Ollama loads it into VRAM (on a 16 GB card that can evict another model) and the request resets that model’s idle timer. That is the point of the test, so do not run it against a service that is serving users or while other GPU work is running. Assumes Ollama on the default port (override with ADDR) and an already-pulled model; the model name is a placeholder. The API call should fail on a missing model rather than pull it, whereas ollama run may download it, so do not substitute ollama run [UNVERIFIED: recalled Ollama behaviour]. Needs curl, python3 and awk. Undo, if the request loaded a model you do not want resident: ollama stop "$MODEL". This snippet was run only against a local mock of the two endpoints, so confirm the output fields on your own install.

# save as ttft.sh and run with bash (the trap below removes the scratch file on exit)
MODEL=your-model:tag                   # plain name: no quotes or backslashes
ADDR=${ADDR:-localhost:11434}          # host:port of the Ollama API
OUT=$(mktemp)                          # scratch file for the response stream
trap 'rm -f "$OUT"' EXIT
# 1) is it resident? (read-only)
curl -sf --max-time 5 "$ADDR/api/ps" | python3 -m json.tool || echo "no valid JSON from $ADDR/api/ps (server down?)" >&2
# 2) TTFT via streaming; load_duration is in the final JSON line (ns)
start=$(date +%s.%N)                   # GNU date, as on Ubuntu
curl -sfN --max-time 300 "$ADDR/api/generate" \
  -d "{\"model\":\"$MODEL\",\"prompt\":\"Say hi\",\"stream\":true,\"options\":{\"num_predict\":8}}" \
 | { read -r first
     t1=$(date +%s.%N)
     if [ -z "$first" ]; then
       echo "no response (server down, or model not found)" >&2
     else
       awk -v a="$start" -v b="$t1" 'BEGIN { printf "ttft_s=%.3f\n", b - a }'
       { printf '%s\n' "$first"; cat; } > "$OUT"
     fi; }
python3 - "$OUT" <<'PY'
import json, sys
lines = [l for l in open(sys.argv[1]) if l.strip()]
if lines:
    try:
        d = json.loads(lines[-1])
    except ValueError:
        print("last line is not JSON:", lines[-1][:200])
    else:
        ld, td = d.get("load_duration"), d.get("total_duration")
        print("load_duration_s=", "absent" if ld is None else ld / 1e9,
              "total_s=", "absent" if td is None else td / 1e9)
PY

ttft_s covers the connection, any model load, prompt processing and the first streamed chunk (assumed to be the first generated token [INFERRED]). Run once cold-ish and once immediately again: the second run should show load_duration near zero because the model is now hot, and a warm first run’s load_duration is an upper bound on S/B_link, because it also includes the runtime overhead the budget excludes [INFERRED]; compare it with the budget above with that in mind. For page-cache-cold vs warm control, follow measuring-a-thunderbolt-egpu-bandwidth-and-inference-linux.md (do not drop caches unless you intend to). For llama-server, time the first /v1/chat/completions token after start with --load-mode none vs default and compare; each start is a service restart, so it is state-changing. [INFERRED]

Anti-patterns

Sources

  1. llama.cpp server README (flags, load-mode, sleep, router): https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/README.md
  2. llama.cpp completion README (load-mode note on RAM): https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/completion/README.md
  3. Ollama FAQ: https://raw.githubusercontent.com/ollama/ollama/main/docs/faq.mdx
  4. Ollama API docs: https://raw.githubusercontent.com/ollama/ollama/main/docs/api.md
  5. Ollama envconfig source (OLLAMA_LOAD_TIMEOUT etc.): https://raw.githubusercontent.com/ollama/ollama/main/envconfig/config.go
  6. LM Studio TTL and Auto-Evict: https://lmstudio.ai/docs/developer/core/ttl-and-auto-evict
  7. vLLM load config source (load_format, safetensors strategies): https://raw.githubusercontent.com/vllm-project/vllm/main/vllm/config/load.py
  8. vLLM Run:ai Model Streamer: https://docs.vllm.ai/en/latest/models/extensions/runai_model_streamer.html
  9. Linux kernel vm sysctl (drop_caches): https://www.kernel.org/doc/html/latest/admin-guide/sysctl/vm.html