Mac local LLMs: Serving ops and multi-model
Parent: Running LLM models locally on a Mac · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
LaunchDaemon (/Library/LaunchDaemons) loads at system start; LaunchAgent only at user login; a job starts only with RunAtLoad or KeepAlive. Use `sudo launchctl bootstrap system x.plist` / `bootout`, `gui/<uid>` for agents.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
launchd and boot
- LaunchDaemon (/Library/LaunchDaemons) loads at system start; LaunchAgent only at user login; a job starts only with RunAtLoad or KeepAlive. Use `sudo launchctl bootstrap system x.plist` / `bootout`, `gui/<uid>` for agents. [source]
- Env vars from .zshrc never reach launchd: put OLLAMA_HOST etc. in plist EnvironmentVariables. `launchctl setenv` does not survive reboot. [source]
- ExitTimeOut: launchd.plist(5) says the default is system-defined (corrects "20 s" and an issue reporter's "5 s"); 0 means infinity and can stall shutdown forever. Set it explicitly for a 400 GB-class MLX server so teardown finishes. [source]
- Daemon without UserName runs as root: set HOME or models land in root's home. Ollama LaunchDaemon with UserName/GroupName/ExitTimeOut 30 works (issue 2955; corrects the earlier "no pre-login daemon" claim). [source]
- `brew upgrade ollama` replaces the Cellar plist and drops edits; copy it to ~/Library/LaunchAgents after stopping the brew service. [source]
- oMLX: `brew services start omlx`; idle eviction 5 min unless pinned; defaults ~/.omlx/models, port 8000. OLLAMA_KEEP_ALIVE=-1 pins on a dedicated box (default 5 min). [source]
- llmster: calling `lms` before `lms daemon up` finishes spawns a second daemon; loser exits 0 ("Exiting due to PID lock loss"), looking like a KeepAlive loop. [source]
- launchd has no job ordering (sysctl vs server daemon can race); KeepAlive after a panic relaunches the workload that caused it; metal-guard's shell guard skips launchd, call `metal-guard panic-gate`. [source]
Pre-login, FileVault, TCC, headless
- FileVault on: nothing runs after an unplanned reboot until unlock. `sudo fdesetup authrestart` is one-use, planned reboots only, not after OS updates. macOS 26+ on Apple silicon: SSH unlock with Remote Login on. Only Secure Token users can unlock. [source]
- Astropad: FileVault off plus auto-login for home use; `sudo pmset -a autorestart 1`; sleep off via `pmset -a sleep 0`, `disksleep 0`, `womp 1`, `powernap 0`; lid-closed `disablesleep 1`. [source]
- TCC binds daemons even as root, with no prompt: denial is silent (empty `ls`, Console "System Policy: ... deny(1) file-read-data", Errno 1 "Operation not permitted"). Keep binaries, models, logs out of ~/Desktop, Documents, Downloads, iCloud, /Volumes. [source]
Health and liveness
- llama-server `/health`: 503 "Loading model", then 200 `{"status":"ok"}`; public even with `--api-key`. [source]
- Poisoned backend (issue 27309): after Metal OOM `/health` stays 200, completions return 500 "Compute error." ("backend is in error state"), process stays alive, so KeepAlive never fires. Probe with `max_tokens:1` completion (send the key), SIGKILL on failure; router children ignore SIGTERM, `kill -9` autoloads a fresh child. PR 28242 (fail init) unreviewed. [source]
- mlx_lm.server: dead generation thread logs "mlx_lm.server generation thread died", `/health` gives 503 but process keeps serving; poll `/health` and restart. `--kv-bits` disables batching. Error `[metal::malloc] Resource limit (499000) exceeded` is a live-resource cap; recycle by request count. [source]
- `--sleep-idle-seconds`: `/health`, `/props`, `/models`, `/metrics` neither wake nor reset it; issue 29689: a request arriving as it sleeps hangs or SIGSEGVs; a 1-token completion first held 123/123. [source]
Routers and multi-model
- llama.cpp router: `--models-max` default 4 (0 unlimited) counts models not bytes, includes sleeping; on current master unload precedes load so peak is models_max (corrects the earlier "three models" claim). Set it on the CLI, not the INI. [source]
- Router presets: CLI > model section > `[*]`; `stop-timeout` 10 s; `GET /models?reload=1`. [source]
- llama-swap (v262, 2026-10-03): `routing.router.use: group|matrix`; groups `swap` true, `exclusive` true, `persistent` false; matrix sums `evict_costs`. `globalTTL` 0, `unloadTimeout` 10 s, `healthCheckTimeout` 120 s, concurrency limit gives 429. Preload without a group makes models evict each other. [source]
- One mlx_lm.server per model under llama-swap keeps memory measurable; cold start 2.4 s (3 GB), 7.7 s (20 GB); swaps lose in-memory prefix caches. [source]
- oMLX: one process, LRU, pinning, per-model TTL, `--memory-guard safe|balanced` (default) or `--memory-guard-gb N`, `--max-concurrent-requests` 8. [source]
- vllm-mlx registry: `--memory-budget-gb` (PR 640) overrides weights-only `manager.memory_budget_gb`; budget error "models-config manager.memory_budget_gb is required"; 80-100 GB suggested on 128 GB; `wait_then_fail` for user-facing. [source]
- LM Studio: Auto-Evict affects only JIT models per its doc (a blog says otherwise); bug 2051 LM Link models bypass it. [source]
- Budget rule: weights + KV (ctx x slots) + prefix cache + OS headroom vs the wired limit, not RAM; `footprint -p <pid>` shows IOAccelerator, undercounts mmapped GGUF. [source]
Concurrency and batching
- vllm-mlx default serves one request at a time: pass `--continuous-batching`; `--max-num-seqs` default 256 too high (16 used on 128 GB). 16 concurrent: 3.7x (0.6B), 2.6x (8B), not the 4.3x summaries claim. [source]
- llama-server `--parallel N` splits ctx: `--ctx-size 32768 --parallel 8` leaves 4,096 per slot; Qwen3.6-27B KV about 4 GB per 64K slot. [source]
- mlx-lm 0.31.0 issue 965: 16 concurrent requests cross-contaminated answers (fixed v0.31.2); throughput does not prove correctness. LM Studio 0.4.2+ batches MLX. [source]
Security
- llama-server default has no key. Prefer `--api-key-file` (env LLAMA_ARG_API_KEY_FILE) over argv; with a key `/metrics` returns 401 "Invalid API Key" (PR 28915 open). [source]
Open questions
- Does `--sleep-idle-seconds` free Metal memory? Untested. Headless virtual display on M1+ (Astropad) vs empty `screencapture` until BetterDisplay (jock.pl): unresolved. [source]
Corrections and disagreements
- CONTRADICTS: Towards AI and the paper's contribution list say 4.3x aggregate at 16 concurrent, while the paper's measured Figure 2 text gives 3.7x (0.6B) and 2.6x (8B); existing dossiers hold only the 87 to 215 tok/s (2.6x) figure. [source]
- LM Studio pinning. Hannecke (2026-03-06) says Auto-Evict "doesn't distinguish pinned from JIT-loaded models", so turn it off when pinning with `ttl=0`. CONTRADICTS lm-studio-on-mac.md and LM Studio's own TTL doc: Auto-Evict affects only JIT-loaded models and "non-JIT loaded models are not affected"; bug 2051 shows the reverse failure (manually loaded models never evicted). [source]
- OpenCode PR 28352 evidence. CONTRADICTS the framing that it measures per-model cache slot eviction, and sharpens prompt-cache-invalidation-by-agent-clients.md: the PR is a client-side fix (stop sending `reasoning_content: ""` on assistant turns without reasoning) tested against llama-swap in front of llama-server. Its own verification line reads "WITH reasoning_content -> 97.2% cache hit, WITHOUT -> 0%", which is the opposite wording of its description (the empty string breaks prefix matching); read it as field-present-with-text vs the pre-fix path, not as slot behavior. Additional measured numbers: multi-turn 0% on turn 1, 52.6% on turn 2, 91.0% on turn 3; 0% hits on 196K-token prompts in production captures. The PR linked issue 19081 and was opened 2026-05-19. Nothing in it measures eviction between models. [source]
- CONTRADICTS continuous-batching-on-mlx.md: the startup memory report now includes MLLM engines' memory-aware prefix cache (issue 712 closed via PR 713 on 2026-09-05) rather than ignoring it. [source]
- CONTRADICTS the LM Studio TTL doc: Hannecke claims Auto-Evict evicts pinned models as well and advises turning it off when pinning with `ttl=0`. [source]
- CONTRADICTS: ollama-on-macos.md line 28 and 77. A working Ollama LaunchDaemon exists in issue 2955 (UserName, GroupName, OLLAMA_HOST in EnvironmentVariables, ExitTimeOut 30, KeepAlive, `/opt/homebrew/bin/ollama serve`, installed with `sudo cp ollama.plist /Library/LaunchDaemons/`) and is reported to start on boot. [source]
- Does a headless Apple silicon Mac have a working display by default? Side A (Astropad): M1 and later Mac minis create a virtual display automatically, no hardware needed, default 1920 by 1080. Side B (jock.pl): on his headless Mac mini `screencapture` returned empty files and UI automation failed until a BetterDisplay virtual screen existed. The existing dossier windowserver-gpu-contention-and-display-sleep-mi.md asserts Side A ("A headless Apple silicon Mac still runs WindowServer against a virtual display"). CONTRADICTS (partly): windowserver-gpu-contention-and-display-sleep-mi.md. The two accounts can both be true if the automatic display exists for remote-desktop sessions but not for `screencapture` of the console session; neither source says which Mac model or macOS version the second account used. [source]
- CONTRADICTS: two-llama-cpp-or-llama-swap-models-loaded-concur.md ("Peak memory during a swap can therefore be three models, not --models-max") for current master: the source serializes unload and load, so peak is `models_max` children; the cited report gives no router build. [source]
- CONTRADICTS launchd-service-setup-for-local-llm-servers.md line 53 (ExitTimeOut defaults to 20 s): launchd.plist(5) says only that the default value is system-defined. [source]
- CONTRADICTS omlx-fatal-teardown-watchdog-and-wired-memory-st.md ("default 5 s launchd kill window"): the man page names no 5 s default; the figure is the issue reporter's observation, not a documented value. [source]
Concepts in this cluster
- Continuous batching on MLX [source]
- Multi-model serving on a large Mac with llama-swap [source]
- launchd service setup for local LLM servers [source]
- Continuous batching and aggregate throughput on MoE [source]
- Continuous batching and concurrent-request serving on Apple silicon [source]
- Multi-model serving and memory budgeting in vllm-mlx and oMLX registries [source]
- Per-model memory measurement with footprint and IOAccelerator [source]
- Service accounts and API-key hardening for LAN-exposed local LLM servers [source]
- fdesetup authrestart and FileVault pre-boot unlock for headless Macs [source]
- llama-server readiness and health semantics after fatal Metal init OOM [source]
- macOS TCC Files and Folders consent for boot-time daemons [source]
- Headless Apple silicon Mac virtual display and the interactivity watchdog [source]
- Long-running MLX server OOM recovery and process recycling [source]
- Supervisor-side liveness probes for llama-server (completion probe plus SIGKILL) [source]
- llama-server router mode child instance exposure and auth [source]
- router --models-max swap ordering (unload before load) peak memory [source]
- llama.cpp --sleep-idle-seconds router sleeping state keeps its models-max slot [source]
- launchd ExitTimeOut and graceful stop discipline for large MLX models [source]
Children
- macOS TCC Files and Folders consent for boot-time daemons
- Multi-model serving and memory budgeting in vllm-mlx and oMLX registries
- Multi-model serving on a large Mac with llama-swap
- Per-model memory measurement with footprint and IOAccelerator
- router --models-max swap ordering (unload before load) peak memory
- Service accounts and API-key hardening for LAN-exposed local LLM servers
- Supervisor-side liveness probes for llama-server (completion probe plus SIGKILL)
- Continuous batching and aggregate throughput on MoE
- Continuous batching and concurrent-request serving on Apple silicon
- Continuous batching on MLX
- fdesetup authrestart and FileVault pre-boot unlock for headless Macs
- Headless Apple silicon Mac virtual display and the interactivity watchdog
- launchd ExitTimeOut and graceful stop discipline for large MLX models
- launchd service setup for local LLM servers
- llama.cpp --sleep-idle-seconds router sleeping state keeps its models-max slot
- llama-server readiness and health semantics after fatal Metal init OOM
- llama-server router mode child instance exposure and auth
- Long-running MLX server OOM recovery and process recycling