LiteLLM tool calls and SSE streaming
Parent: LiteLLM gateway and SDK engineering · Published reference · snapshot 2026-09-30 · skill ai-llm-model-layer/references/litellm-tool-calls-and-sse-streaming.md
↓ Facts as markdownall context files
12 source-anchored research claims on LiteLLM tool calls and SSE streaming, grouped by facet. Original confidence and source-owner limits are retained.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Definitions
- For application-defined client tools, the model proposes a structured call and application code executes it, then returns the result in a subsequent request. LiteLLM's function-call example follows that loop; distinguish it from provider-hosted server tools that execute on provider infrastructure. [source] — confidence high; native contracts do not independently certify LiteLLM implementation · confidence: high
Structure and components
- With stream=True, consume LiteLLM's returned iterator; for asynchronous streaming, await acompletion and then async-iterate it. Text is carried in choices[].delta, so consumers must also handle structured tool deltas and terminal chunks instead of assuming every chunk contains text. [source] — confidence medium; single-owner LiteLLM evidence · confidence: medium
How it works
- SSE framing is UTF-8 text with events dispatched at blank lines, not one JSON document per network read. Its parser joins multiple data lines and ignores colon-prefixed comment lines; preserve that framing before interpreting provider-specific JSON payloads. [source] — confidence low; single-owner WHATWG evidence; qualify exact deployment · confidence: low
Parameters and configuration
- Request streaming usage deliberately. LiteLLM documents stream_options.include_usage; OpenAI's Chat Completions contract places total usage in a final chunk with choices=[] before [DONE], and an interrupted stream may omit it. Treat absent final usage as unknown, not zero. [source] — confidence medium; native contracts do not independently certify LiteLLM implementation · confidence: medium
How-to and procedures
- For OpenAI Chat Completions tool streams, accumulate argument fragments by the tool-call index. id, function.name and type may appear only in the first delta; preserve them while appending function.arguments, then parse the completed argument JSON. [source] — confidence low; single-owner OpenAI evidence; qualify exact deployment · confidence: low
- LiteLLM exposes stream_chunk_builder to reconstruct a completed response from accumulated chunks. Its documented availability is not a substitute for checking preservation of tool IDs, argument JSON, finish reasons and provider-specific metadata in the selected streaming path. [source] — confidence low; single-owner LiteLLM evidence; qualify exact deployment · confidence: low
- Validate function.arguments as JSON and against the expected tool schema before execution. LiteLLM's example warns about invalid JSON, and OpenAI's API reference warns about invalid JSON or invented arguments; a successful completion response is not a validated invocation. [source] — confidence high; native contracts do not independently certify LiteLLM implementation · confidence: high
- For OpenAI-style Chat Completions, append the original assistant tool-call message, then append a role=tool result with the matching tool_call_id for each returned call before requesting the next completion. Do not reduce this conversation state to message.content alone. [source] — confidence high; native contracts do not independently certify LiteLLM implementation · confidence: high
- Use supports_function_calling and supports_parallel_function_calling as capability checks for the exact model/provider. The inspected helpers read model metadata; they do not execute a tool round trip, so qualify the actual deployment before enabling an autonomous tool loop. [source] — confidence high; single-owner LiteLLM evidence · confidence: high
Problems, failure modes and limitations
- LiteLLM's inspected repetition guard tracks consecutive identical text deltas longer than two characters and raises InternalServerError at REPEATED_STREAMING_CHUNK_LIMIT. This text-repetition safeguard does not replace explicit request timeouts or output-token limits. [source] — confidence high; single-owner LiteLLM evidence · confidence: high
- Handle errors during both stream creation and consumption. LiteLLM's exception example encloses iteration in its handler, and Anthropic documents error events inside an established SSE stream; successful headers alone do not establish successful completion. [source] — confidence medium; native contracts do not independently certify LiteLLM implementation · confidence: medium
Changes and history
- An upstream report for LiteLLM 1.98.0-rc.1 with a vLLM Chat Completions backend describes tool-argument deltas arriving in an end-of-stream burst when include_usage=true. The report was still open at inspection; it was not reproduced here and does not establish a defect in every version or provider. [source] — confidence low; single-owner LiteLLM evidence; qualify exact deployment · confidence: low
Children
- LiteLLM stream compatibility fixtures (frontier)
- LiteLLM streamed argument assembly (frontier)
- LiteLLM tool result correlation (frontier)
- LiteLLM usage and stream termination (frontier)
Frontier under this node: LiteLLM stream compatibility fixtures, LiteLLM streamed argument assembly, LiteLLM tool result correlation, LiteLLM usage and stream termination