<!-- llms-explorer concept facts · https://llms-explorer.com/tree/service-accounts-and-api-key-hardening-for-lan-e/ · pack 2026-10-05 · ~1935 tokens -->

# Service accounts and API-key hardening for LAN-exposed local LLM servers

> Key sources: `--api-key KEY` (comma-separated list allowed; env LLAMA_API_KEY) and `--api-key-file FNAME` (one key per line, lines starting with # are comments; env LLAMA_ARG_API_KEY_FILE). Default is none, meaning no authentication.

Parent: [Mac local LLMs: Serving ops and multi-model](https://llms-explorer.com/tree/mac-local-llms-serving-ops-and-multi-model/) · 1 facets · 29 facts · page: https://llms-explorer.com/tree/service-accounts-and-api-key-hardening-for-lan-e/

## Facts

- Key sources: `--api-key KEY` (comma-separated list allowed; env LLAMA_API_KEY) and `--api-key-file FNAME` (one key per line, lines starting with # are comments; env LLAMA_ARG_API_KEY_FILE). Default is none, meaning no authentication. — [source](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md)
- Use the file or env form in a plist: `--api-key KEY` puts the secret in the process argument list, visible in `ps` to every local user (inferred from how argv works; the oMLX guide in the existing dossier avoids it with `$(cat file)`, which still expands into argv). Prefer `--api-key-file` with a mode-600 file owned by the service account, or LLAMA_ARG_API_KEY_FILE in EnvironmentVariables. — source: `asserted`
- Anthropic-style clients may send the key as `x-api-key`; the README's /v1/messages example uses that header besides the Bearer form. — [source](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md)
- Public endpoints that skip the key: `/health` and `/v1/health` (documented), static WebUI assets (PRs 22038 and 21269 skip verification for static files, both closed in April 2026). `/props` is read-only unless `--props` is set; POST /props is disabled by default. — [source](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md)
- `/metrics` is NOT public: with an API key set, a Prometheus scrape of `/metrics` gets 401 "Invalid API Key". PR 28915 (14 Sep 2026) would add it to the public allowlist; it was open and had been converted to draft then back to ready after a bot flagged the PR template and AI-written description. Scrapers need to send the key until that merges. — [source](https://github.com/ggml-org/llama.cpp/pull/28915)
- Apr 2026: PRs 21180 (API key option work), 22038 and 21269 (skip key for WebUI static files). June 2026: PR 24825 proposed constant-time key comparison because `std::find` over `std::string::operator==` leaks matched-prefix length through timing (CWE-208, localhost most exploitable); it was closed without merge and also bundled an unrelated pickle fix. As of the sources read, the comparison is unchanged. — [source](https://github.com/ggml-org/llama.cpp/pull/24825)
- Aug 2026: PR 27027 (open) reports that in router mode the `/models` path was treated as public regardless of method, so unauthenticated `POST /models` and `DELETE /models?model=...` reached the router handlers (400 and 500 responses prove they passed auth); the fix limits public access to GET and adds a test. — [source](https://github.com/ggml-org/llama.cpp/pull/27027)
- Sep 2026: issue 28820, router mode: `--api-key` was not forwarded to spawned children but `--api-key-file` was, so a client using the CLI key could list models and then got 401 on every chat completion. PR 28837 (open) forwards LLAMA_API_KEY; PR 28938 (merged 21 Sep 2026, commit 982a332) instead strips LLAMA_ARG_API_KEY_FILE too, so no key reaches child argv and router-internal calls (`POST /v1/streams/lookup`, `DELETE /v1/stream`) stop returning 401. The maintainer's test matrix after the merge: remaining NOK cases are a key set via `LLAMA_ARG_API_KEY_FILE` in the environment (still forwarded) and a cosmetic CORS warning in the router log. — [source](https://github.com/ggml-org/llama.cpp/pull/28938)
- Router-mode design principle stated in PR 28938: authentication belongs to the router; children should receive no keys. This also keeps keys out of child process arguments. — [source](https://github.com/ggml-org/llama.cpp/pull/28938)
- Child instances bind their own ports on the same host; if the router is bound to 0.0.0.0 the children may still be reachable directly with no key after the 28938 change (inferred from "no API keys reach child instances"; no source tests direct child access). — source: `asserted`
- A LAN service account still needs a TCC-clean home for its models, logs and key file (see the TCC dossier); `_svcuser`-style hidden accounts have shell /usr/bin/false and cannot log in. — source: `asserted`
- On Apple silicon, an account created only for services has no Secure Token; with FileVault it cannot unlock the disk, so the service account must not be the unlock account (see the fdesetup dossier). — source: `asserted`
- llama.app (LlamaBarn): `defaults write app.llama.Llama extraServerArgs -string "--api-key secret"` puts the secret in a preferences plist and in argv; this is weaker than a key file (already recorded in the existing dossier; the risk assessment is added here). — source: `asserted`
- PR 28837 (forward LLAMA_API_KEY to children) and PR 28938 (strip file keys from children) pick opposite directions for the same bug; the maintainers approved and merged the strip approach. — [source](https://github.com/ggml-org/llama.cpp/pull/28938)
- Whether the LAN-exposed router still accepts unauthenticated direct requests on child ports. — source: `asserted`
- Whether PR 27027 (non-GET /models) is merged; it was open when fetched. — source: `asserted`
- Whether the key comparison remains non-constant-time on current master (PR 24825 closed unmerged). — source: `asserted`
- llama-server accepts multiple API keys as a comma-separated list through `--api-key` (env LLAMA_API_KEY). — [source](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md)
- `--api-key-file` reads one key per line, treats # lines as comments, and has env LLAMA_ARG_API_KEY_FILE. — [source](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md)
- The default is no API key. — [source](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md)
- `/health` and `/v1/health` bypass the key check by design. — [source](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md)
- With an API key set, `GET /metrics` returns 401 "Invalid API Key"; PR 28915 to exempt it was open on 4 Oct 2026. — [source](https://github.com/ggml-org/llama.cpp/pull/28915)
- PR 24825 reported a timing side channel in API-key comparison (std::find, short-circuit on first mismatching byte) and was closed unmerged on 19 Jun 2026. — [source](https://github.com/ggml-org/llama.cpp/pull/24825)
- In router mode, unauthenticated POST and DELETE on `/models` passed the key middleware before PR 27027 (open). — [source](https://github.com/ggml-org/llama.cpp/pull/27027)
- In router mode with `--api-key-file` set, children validated only file keys, so CLI-key clients got 401 on chat completions (issue 28820). — [source](https://github.com/ggml-org/llama.cpp/pull/28837)
- PR 28938 (merged 21 Sep 2026) unsets LLAMA_ARG_API_KEY_FILE in `unset_reserved_args()` so no keys reach router children. — [source](https://github.com/ggml-org/llama.cpp/pull/28938)
- After PR 28938 a key supplied via the LLAMA_ARG_API_KEY_FILE environment variable is still forwarded to children. — [source](https://github.com/ggml-org/llama.cpp/pull/28938)
- A key given as `--api-key VALUE` is visible in the process argument list; a key file is not (inferred). — source: `asserted`
- A LAN-exposed llama-server needs an external probe or reverse proxy for `/metrics` and `/health` because those two behave differently under a key (inferred). — source: `asserted`
