LiteLLM observability and caching
Parent: LiteLLM gateway and SDK engineering · Published reference · snapshot 2026-09-30 · skill ai-llm-model-layer/references/litellm-observability-and-caching.md
↓ Facts as markdownall context files
12 source-anchored research claims on LiteLLM observability and caching, grouped by facet. Original confidence and source-owner limits are retained.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Parameters and configuration
- Semantic caching scopes buckets by key, team and organization. End users sharing a virtual key share that bucket unless semantic_cache_scope: end_user and an authenticated end-user identifier are present; absent end-user identity falls back to key scope. [source] — confidence medium; single-owner LiteLLM evidence · confidence: medium
- ttl sets entry retention and s-maxage limits acceptable cached age. Redis expiration removes keys when their lifetime elapses; neither setting establishes answer freshness after upstream documents, tools or authorization change. [source] — confidence medium; native contracts do not independently certify LiteLLM implementation · confidence: medium
- OTel v2 is opt-in with LITELLM_OTEL_V2=true. OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT defaults to no_content; span_only, event_only and span_and_event select capture destinations. LITELLM_OTEL_INTEGRATION_ENABLE_METRICS and LITELLM_OTEL_INTEGRATION_ENABLE_EVENTS are separate boolean flags that default to false at the inspected commit. Pin the integration generation and verify exported traces instead of copying v1 assumptions into v2. [source] — confidence medium; single-owner LiteLLM evidence · confidence: medium
How-to and procedures
- A namespace prefixes cache keys; it is not itself an access-control rule. Use Redis ACL key patterns and credentials to restrict cache access, then confirm the configured namespace covers every proxy key class. [source] — confidence medium; native contracts do not independently certify LiteLLM implementation · confidence: medium
- Normalize total input-token usage across cached and uncached provider fields. OpenTelemetry specifies inclusive input totals; Anthropic documents input_tokens plus cache_creation_input_tokens plus cache_read_input_tokens. Do not label uncached input alone as total prompt size. [source] — confidence medium; native contracts do not independently certify LiteLLM implementation · confidence: medium
Problems, failure modes and limitations
- Redis maxmemory-policy can evict entries before their TTL expires. noeviction instead rejects memory-growing writes; investigate memory policy and failed writes when cache hit rate falls rather than assuming only TTL caused misses. [source] — confidence low; single-owner Redis evidence; qualify exact deployment · confidence: low
- Exact cache keys hash recognized LLM request parameters; provider-specific optional parameters need enable_caching_on_provider_specific_optional_params. Tenant identity is not automatically included for exact caching. Add a trusted tenant namespace and test cross-tenant identical requests. [source] — confidence medium; single-owner LiteLLM evidence · confidence: medium
- Langfuse mask_input and mask_output request metadata do not redact the legacy OpenTelemetry logger. Use explicit NO_CONTENT or turn_off_message_logging, and inspect exported spans and events for residual metadata. [source] — confidence medium; single-owner LiteLLM evidence · confidence: medium
- OTel v2 attribute include_list or exclude_list filters metric labels, not span attributes. Excluding high-cardinality metric fields does not establish a trace-content privacy policy. [source] — confidence low; single-owner LiteLLM evidence; qualify exact deployment · confidence: low
- Semantic caching excludes conversation content from its bucket key and searches similarity within that bucket. The guide warns that multi-turn agents can replay stale responses, including tool calls; retain exact caching or disable semantic replay unless correctness is demonstrated. [source] — confidence medium; single-owner LiteLLM evidence · confidence: medium
Comparisons and alternatives
- LiteLLM response caching reuses a stored response and can avoid a new model call. Semantic caching selects similar embedded prompts; provider prompt caching reuses input processing instead of replaying a complete answer. Keep their hit and correctness metrics separate. [source] — confidence medium; single-owner LiteLLM evidence · confidence: medium
- Local and disk response caches are not shared across workers or replicas. Redis is the documented shared-cache choice for multi-worker deployments; compare hit rates per worker before interpreting latency as model performance. [source] — confidence low; single-owner LiteLLM evidence; qualify exact deployment · confidence: low
Children
- LiteLLM OpenTelemetry generations and privacy (frontier)
- LiteLLM exact cache identity and tenant namespace (frontier)
- LiteLLM semantic replay correctness for agents (frontier)
- Provider usage normalization with cached tokens (frontier)
- Redis cache expiration eviction and ACLs (frontier)
Frontier under this node: LiteLLM OpenTelemetry generations and privacy, LiteLLM exact cache identity and tenant namespace, LiteLLM semantic replay correctness for agents, Provider usage normalization with cached tokens, Redis cache expiration eviction and ACLs