Local LLM model load path over a Thunderbolt eGPU

Parent: Thunderbolt eGPU on Linux for local LLM inference · Published reference · snapshot 2026-09-08 · skill ai-llm-model-layer/references/local-llm-model-load-path-over-thunderbolt-linux.md

Also known as: cold start llm, llama.cpp mmap, model load time egpu, ollama keep_alive, page cache llm model

↓ Facts as markdown↓ Download this reference fileall context files

What happens between a model file on NVMe and weights resident in VRAM when the host link is slow: the stages and which one bounds cold and warm loads, llama.cpp load modes, Ollama and LM Studio keep-

These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.

Local LLM Model Load Path Over a Thunderbolt eGPU (Linux) and Keep-Warm Policy

Core Concepts

The Load Path Stages

llama.cpp Loading Behaviour

Ollama and LM Studio Residency

Page Cache and Storage

Formats and Loaders

Cold-Start Budget (arithmetic on assumed rates, not measurements)

Residency Decision Guide (keep resident vs reload)

Changing residency settings (OPTIONAL, state-changing; undo beside each)

Read-only Measurement Snippet (time to first token, TTFT, after load)

Anti-patterns

Sources

Children

Frontier under this node: Cold-start budget: file size over link rate, Keep-warm versus reload decision, LM Studio JIT loading, TTL and auto-evict, Ollama keep-alive, preload and eviction, Page-cache prewarm with vmtouch and fincore, drop_caches for an honest cold test, llama-server idle sleep and router mode, llama.cpp load modes: mmap, mlock and direct I/O, mmap hides load cost: measuring time to first token, vLLM load formats: Run:ai streamer, InstantTensor, prefetch

← the whole tree · 3D view· how to read this page