<!-- llms-explorer concept facts · https://llms-explorer.com/tree/litellm-proxy-deployment-and-configuration/ · pack 2026-09-30 · ~1367 tokens -->

# LiteLLM proxy deployment and configuration

> 12 source-anchored research claims on LiteLLM proxy deployment and configuration, grouped by facet. Original confidence and source-owner limits are retained.

Parent: [LiteLLM gateway and SDK engineering](https://llms-explorer.com/tree/litellm-gateway-sdk-engineering/) · 6 facets · 12 facts · page: https://llms-explorer.com/tree/litellm-proxy-deployment-and-configuration/

## Structure and components

- Separate config.yaml concerns: model_list defines deployments, router_settings configures the router, litellm_settings configures SDK-level behavior, general_settings configures the server, and environment_variables supplies environment values. Start the proxy with --config pointing to that file. — [source](https://docs.litellm.ai/docs/proxy/configs) *(confidence low; single-owner LiteLLM evidence; qualify exact deployment · confidence: low)*

## How it works

- Use shared Redis when running multiple proxy instances so counters, router state and response caching are shared. Without it, the production guide describes per-instance limit enforcement and local cache hits; adding replicas can change effective allowances. — [source](https://docs.litellm.ai/docs/proxy/deploy) *(confidence high; single-owner LiteLLM evidence · confidence: high)*

## Parameters and configuration

- Treat model_name as the client-facing alias and litellm_params.model as the provider/model identifier. Requests name the alias, so inspect model_list when a requested model routes to an unexpected provider deployment. — [source](https://docs.litellm.ai/docs/proxy/configs) *(confidence high; single-owner LiteLLM evidence · confidence: high)*

## How-to and procedures

- Pin a LiteLLM release image for repeatable upgrades and rollback. For an immutable artifact, record its Docker digest; version tags identify releases but Docker's digest form fixes the exact image content. — [source](https://docs.litellm.ai/docs/proxy/deploy) *(confidence high; native contracts do not independently certify LiteLLM implementation · confidence: high)*
- Place adapter-specific api_base, model and credential references under a deployment's litellm_params. The documented os.environ/NAME syntax reads the named environment value; keep provider credentials outside committed YAML and inject them through the deployment environment. — [source](https://docs.litellm.ai/docs/proxy/configs) *(confidence high; single-owner LiteLLM evidence · confidence: high)*
- For a local-only Docker gateway, bind the published port to 127.0.0.1 explicitly rather than copying -p 4000:4000 unchanged. Docker publishes unqualified ports beyond the host by default; its documentation also notes a localhost-isolation caveat for releases older than 28.0.0. — [source](https://docs.litellm.ai/docs/proxy/docker_quick_start) *(confidence medium; native contracts do not independently certify LiteLLM implementation · confidence: medium)*
- Use /health/liveliness for process health and /health/readiness for configured readiness dependencies. These lightweight endpoints do not make LLM calls. A separate authenticated request through the intended alias is needed to validate the full provider path. — [source](https://docs.litellm.ai/docs/proxy/health) *(confidence high; single-owner LiteLLM evidence · confidence: high)*
- On Kubernetes, map liveness to process recovery and readiness to traffic eligibility. A failing readiness probe removes the pod from service traffic; avoid using a provider-inference check as a process liveness probe, because an upstream failure is not proof the proxy process needs restart. — [source](https://docs.litellm.ai/docs/proxy/health) *(confidence medium; native contracts do not independently certify LiteLLM implementation · confidence: medium)*
- Assign schema migrations to the deployment's migration job or Helm PreSync hook, and set DISABLE_SCHEMA_UPDATE=true on serving pods. Current production documentation says Prisma migrate deploy otherwise runs on startup; coordinate migrations separately from replica rollout. — [source](https://docs.litellm.ai/docs/proxy/deploy) *(confidence high; single-owner LiteLLM evidence · confidence: high)*

## Problems, failure modes and limitations

- The current database-free quickstart warns that litellm_settings.max_budget does not enforce a spend cap without a database: the global check lacks stored spend and requests continue. A configured dollar value alone is not evidence of enforcement. — [source](https://docs.litellm.ai/docs/proxy/docker_quick_start) *(confidence low; single-owner LiteLLM evidence; qualify exact deployment · confidence: low)*
- Preserve LITELLM_SALT_KEY after adding database-backed models. The deployment and production guides state that it encrypts stored provider credentials and changing it makes them unreadable; it is not interchangeable with an ordinary rotating client key. — [source](https://docs.litellm.ai/docs/proxy/deploy) *(confidence high; single-owner LiteLLM evidence · confidence: high)*

## Comparisons and alternatives

- A database-free proxy can serve the OpenAI-compatible API using a config file and master-key authentication. It does not provide database-backed virtual keys, persistent spend tracking or Admin UI model management; those capabilities require the stateful deployment path. — [source](https://docs.litellm.ai/docs/proxy/docker_quick_start) *(confidence high; single-owner LiteLLM evidence · confidence: high)*
