<!-- llms-explorer concept facts · https://llms-explorer.com/tree/two-llama-cpp-or-llama-swap-models-loaded-concur/ · pack 2026-10-05 · ~397 tokens -->

# Two llama.cpp or llama-swap models loaded concurrently on one GPU: stability

> Whether llama.cpp router mode unloads the evicted model before or after loading the new one; the one report suggests the load can overlap the residents.

Parent: [Mac local LLMs: GPU stability and kernel panics](https://llms-explorer.com/tree/mac-local-llms-gpu-stability-and-kernel-panics/) · 1 facets · 4 facts · page: https://llms-explorer.com/tree/two-llama-cpp-or-llama-swap-models-loaded-concur/

## Facts

- Whether llama.cpp router mode unloads the evicted model before or after loading the new one; the one report suggests the load can overlap the residents. — source: `asserted`
- Stability of two concurrent llama.cpp or llama-swap models is already covered by multi-model-serving-on-a-large-mac-with-llama-swap.md and metal-gpu-hangs-with-multiple-concurrent-engines.md; no contradiction found. — [source](https://huggingface.co/blog/ggml-org/model-management-in-llamacpp)
- With `--models-max 2` on a 64 GB M1 Max, two resident models plus a third swap-in briefly overlapped and pushed wired memory past 64 GB; the resulting Metal OOM left the loading child poisoned. Peak memory during a swap can therefore be three models, not `--models-max`. — [source](https://github.com/ggml-org/llama.cpp/issues/27309)
- A single-node guide says multi-process serving on one Apple Silicon machine is rarely the right scale-out because two processes contend for the same memory bandwidth and the same GPU. — [source](https://medium.com/@michael.hannecke/tuning-llama-server-on-apple-silicon-9b3e778ab100)
