LiteLLM routing retries and fallbacks
Parent: LiteLLM gateway and SDK engineering · Published reference · snapshot 2026-09-30 · skill ai-llm-model-layer/references/litellm-routing-retries-and-fallbacks.md
↓ Facts as markdownall context files
11 source-anchored research claims on LiteLLM routing retries and fallbacks, grouped by facet. Original confidence and source-owner limits are retained.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Definitions
- Deployments sharing model_name form a model group. Router retries stay within the failed group; configured fallbacks move to other groups after the retry path is exhausted. [source] — confidence low; single-owner LiteLLM evidence; qualify exact deployment · confidence: low
Parameters and configuration
- Fallback lists are ordered model-group chains. Configure Router fallbacks as a list of source-group mappings to target-group lists, for example fallbacks=[{"bad-model": ["my-good-model"]}]. ContextWindowExceededError and ContentPolicyViolationError use context_window_fallbacks and content_policy_fallbacks instead of the generic error chain. [source] — confidence low; single-owner LiteLLM evidence; qualify exact deployment · confidence: low
- allowed_fails and cooldown_time control deployment exclusion after failures; AllowedFailsPolicy customizes thresholds by error type independently of RetryPolicy attempt counts. [source] — confidence low; single-owner LiteLLM evidence; qualify exact deployment · confidence: low
How-to and procedures
- Rate-limit retries should respect provider retry-after and exponential backoff. LiteLLM exposes retry_after as a minimum delay and reads exception response headers when computing retry sleep. [source] — confidence medium; native contracts do not independently certify LiteLLM implementation · confidence: medium
- Configure finite retry, fallback, and request-timeout budgets together; router num_retries, max_fallbacks, and provider/client timeout settings describe different stages of one request. [source] — confidence low; single-owner LiteLLM evidence; qualify exact deployment · confidence: low
- Use an error-specific retry policy. Authentication and malformed-input failures require configuration or payload repair; transient connection, rate-limit, and server failures can warrant bounded backoff. [source] — confidence medium; native contracts do not independently certify LiteLLM implementation · confidence: medium
- For proxy qualification, force a real retryable failure in a test backend. Incoming proxy mock_testing_* fallback flags are stripped; they remain usable only in direct Router tests. [source] — confidence low; single-owner LiteLLM evidence; qualify exact deployment · confidence: low
Problems, failure modes and limitations
- Inference from independent retry loops: client retries can repeat an entire gateway retry/fallback chain. Set caller max_retries deliberately and measure total upstream attempts. [source] — confidence medium; native contracts do not independently certify LiteLLM implementation · confidence: medium
- The inspected Anthropic stream router declines fallback after real content reaches the client. Test pre-content failure separately from partial-answer failure; do not assume all endpoint streams restart identically. [source] — confidence low; single-owner LiteLLM evidence; qualify exact deployment · confidence: low
- A 429 can indicate an exhausted spend cap rather than replenishing rate capacity. Check provider error details and retry-after before repeatedly retrying or choosing a fallback. [source] — confidence medium; single-owner Anthropic evidence · confidence: medium
Comparisons and alternatives
- Router num_retries differs from provider SDK max_retries. The inspected Router initializes provider retries to zero, and the routing docs direct proxy users to control attempts with num_retries. [source] — confidence low; single-owner LiteLLM evidence; qualify exact deployment · confidence: low
Children
- Endpoint-specific streaming failover (frontier)
- LiteLLM cooldown policies (frontier)
- LiteLLM model groups and deployment selection (frontier)
- LiteLLM retry ownership and budgets (frontier)
Frontier under this node: Endpoint-specific streaming failover, LiteLLM cooldown policies, LiteLLM model groups and deployment selection, LiteLLM retry ownership and budgets