<!-- llms-explorer concept facts · https://llms-explorer.com/tree/llama-server-router-mode-child-instance-exposure/ · pack 2026-10-05 · ~978 tokens -->

# llama-server router mode child instance exposure and auth

> The router forwards each request to the child that serves the `model` named in the body or query. The router also makes two internal calls to children (`POST /v1/streams/lookup` and `DELETE /v1/stream`) that carry no key.

Parent: [Mac local LLMs: Serving ops and multi-model](https://llms-explorer.com/tree/mac-local-llms-serving-ops-and-multi-model/) · 1 facets · 15 facts · page: https://llms-explorer.com/tree/llama-server-router-mode-child-instance-exposure/

## Facts

- The router forwards each request to the child that serves the `model` named in the body or query. The router also makes two internal calls to children (`POST /v1/streams/lookup` and `DELETE /v1/stream`) that carry no key. — source: `asserted`
- Children inherit the router's command-line arguments and environment; the router removes or overwrites host, port, API key, HF repo and model alias when it spawns them. — [source](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md)
- 15 Sep 2026: a contributor ran a before/after matrix on PR 28938 and found the fix removes the child 401s, with the trade that a child port is reachable without a key. — [source](https://github.com/ggml-org/llama.cpp/pull/28938)
- Before the fix, with `--api-key-file` only, a direct call to a child's `/v1/streams/lookup` returned 401 and `DELETE /v1/stream` returned 401, so the router's own internal stream calls could 401. — [source](https://github.com/ggml-org/llama.cpp/pull/28938)
- A contributor proposed rejecting `--api-key` and `--api-key-file` when both are passed, since he saw no case needing both. — [source](https://github.com/ggml-org/llama.cpp/pull/28938)
- Whether children bind loopback regardless of the router's `--host`; no cached source states the child bind address. — source: `asserted`
- RESOLVES the open question in service-accounts-and-api-key-hardening-for-lan-e.md ("whether the LAN-exposed router still accepts unauthenticated direct requests on child ports"): after PR 28938 a child port is reachable without a key, marked "by design" in a tested matrix. — [source](https://github.com/ggml-org/llama.cpp/pull/28938)
- With `--api-key` alone (no key file), a direct `GET /v1/models` on a child returned 200 without a key both before and after the change, because the CLI key never reached the child; the matrix author labels this "pre-existing". — [source](https://github.com/ggml-org/llama.cpp/pull/28938)
- With `--api-key-file` alone, a direct child `GET /v1/models` returned 401 before the change and 200 after it. — [source](https://github.com/ggml-org/llama.cpp/pull/28938)
- Before the change, with `--api-key-file`, direct child calls to `POST /v1/streams/lookup` and `DELETE /v1/stream` returned 401; after it they returned 200 and 204. — [source](https://github.com/ggml-org/llama.cpp/pull/28938)
- A key set in a preset INI model section was forwarded to the child before the change and stripped after it. — [source](https://github.com/ggml-org/llama.cpp/pull/28938)
- After the change the router logs a CORS warning for children; the matrix author calls it cosmetic. — [source](https://github.com/ggml-org/llama.cpp/pull/28938)
- The server README says host, port, API key, HF repo and model alias are controlled by the router and are removed or overwritten when a preset model loads. — [source](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md)
- The server README says router model instances inherit both command-line arguments and environment variables from the router. — [source](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md)
- Practical consequence: with a router bound to a LAN address, authentication protects only the router port; anyone who can reach a child port can call it with no key (inferred from the matrix; child bind address untested). — source: `asserted`
