Two llama.cpp or llama-swap models loaded concurrently on one GPU: stability
Parent: Mac local LLMs: GPU stability and kernel panics · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Whether llama.cpp router mode unloads the evicted model before or after loading the new one; the one report suggests the load can overlap the residents.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Whether llama.cpp router mode unloads the evicted model before or after loading the new one; the one report suggests the load can overlap the residents. [source]
- Stability of two concurrent llama.cpp or llama-swap models is already covered by multi-model-serving-on-a-large-mac-with-llama-swap.md and metal-gpu-hangs-with-multiple-concurrent-engines.md; no contradiction found. [source]
- With `--models-max 2` on a 64 GB M1 Max, two resident models plus a third swap-in briefly overlapped and pushed wired memory past 64 GB; the resulting Metal OOM left the loading child poisoned. Peak memory during a swap can therefore be three models, not `--models-max`. [source]
- A single-node guide says multi-process serving on one Apple Silicon machine is rarely the right scale-out because two processes contend for the same memory bandwidth and the same GPU. [source]
Children
- No children recorded.