<!-- llms-explorer concept facts · https://llms-explorer.com/tree/supervisor-side-liveness-probes-for-llama-server/ · pack 2026-10-05 · ~324 tokens -->

# Supervisor-side liveness probes for llama-server (completion probe plus SIGKILL)

> Whether a supervisor should probe each child or only the router port (a router-level probe with the child's `model` name exercises the child).

Parent: [Mac local LLMs: Serving ops and multi-model](https://llms-explorer.com/tree/mac-local-llms-serving-ops-and-multi-model/) · 1 facets · 3 facts · page: https://llms-explorer.com/tree/supervisor-side-liveness-probes-for-llama-server/

## Facts

- Whether a supervisor should probe each child or only the router port (a router-level probe with the child's `model` name exercises the child). — source: `asserted`
- The completion-probe plus SIGKILL design, the `/health` blind spot, the SIGTERM-absorbed behaviour and the launchd KeepAlive interaction are already covered by llama-server-readiness-and-health-semantics-afte.md; nothing contradicts it. — [source](https://github.com/ggml-org/llama.cpp/issues/27309)
- In the router-mode report the poisoned child was never seen to decode once in its life, so a probe that waits for recovery never succeeds; the trigger was a third model swap-in overlapping two residents under `--models-max 2` on a 64 GB M1 Max, pushing wired memory past the limit. — [source](https://github.com/ggml-org/llama.cpp/issues/27309)
