vLLM parser engine grammar (vllm/parser) and structural-tag strict tool calling
Parent: Mac local LLMs: Chat templates, reasoning and tool calling · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
`vllm/parser/AGENTS.md` says `Parser` unifies reasoning extraction and tool-call extraction behind one object that serving calls through `parse`, `parse_delta` and `is_reasoning_end`.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- `vllm/parser/AGENTS.md` says `Parser` unifies reasoning extraction and tool-call extraction behind one object that serving calls through `parse`, `parse_delta` and `is_reasoning_end`. [source]
- The `engine/` directory holds `parser_engine.py`, `streaming_parser_engine.py`, `parser_engine_config.py`, `token_id_scanner.py` (pre-lexes special-token ids), `incremental_lexer.py` (lexes text), `events.py` and `adapters.py` (exposes an engine through the legacy reasoning and tool parser interfaces). [source]
- The guidelines say text emitted in a state with no `content_events` entry is discarded, special tokens reach the engine as token ids not text, and a text terminal that spans a special token never matches while streaming and must be listed in `token_id_terminals`. [source]
- `DelegatingParser` exists to wrap a legacy `ReasoningParser` plus `ToolParser` pair for backward compatibility and must not be used for new formats. [source]
- On main (2026-10-04) `vllm/parser/` has per-model files for cohere_command, deepseek_v32, deepseek_v4, deepseek_v41, gemma4, glm47_moe, granite, granite_thinking, harmony, inkling, kimi_k2, kimi_k3, ling3, mimo, minimax_m2, mistral, nemotron_v3, plamo3, qwen3, seed_oss and step3p5, plus `response_template.py`. [source]
- Open pull requests port the Hermes parser (51937), llama3_json and llama4_json (51577), Kimi K3 (50229), MiniMax M3 (59743), Olmo3 (48160) and MuseGlimmer (54585) to the engine. [source]
- vLLM's docs split structural-tag enforcement into call obligation and grammar activation: named and `"required"` always activate a structural tag; `"auto"` activates one only when a tool sets `strict: true` or `--tool-strict-level` is `function` or `parameter`; `"none"` disables tool calling. [source]
- The tool-call envelope (markup and function name) is constrained whenever a tag applies, but a tool's argument schema is pinned only when that tool sets `strict: true` or the server uses `--tool-strict-level parameter`; a tool that omits `strict` keeps unconstrained arguments even when another tool in the request is strict. [source]
- `--tool-strict-level` accepts `auto` (default), `function` (constrain the envelope for every request with tools) and `parameter` (also pin every tool's schema), and never relaxes a constraint the request would already get. [source]
- `VLLM_ENFORCE_STRICT_TOOL_CALLING` defaults to true; set to false it removes structural tags regardless of `strict` and takes precedence over `--tool-strict-level`, and it leaves schema-derived structured outputs for named and required calls unchanged. [source]
- The `strict` field is supported on Chat Completion, Responses and Anthropic Messages, and the docs recommend OpenAI strict-schema style (`additionalProperties: false`, all properties required, optional fields as `["string","null"]`). [source]
- The docs say most clients never set `strict`, so with `tool_choice="auto"` malformed markup can leak into the response, which is why `--tool-strict-level function` exists. [source]
- The `hf` parser reads a checkpoint's `response_template` from `tokenizer_config.json` (`--tool-call-parser hf --reasoning-parser hf`), fails at startup if the template or its `thinking` or `tool_calls` field is missing, and rejects strict tools, required or named choice and `parallel_tool_calls=false` because it does no constrained decoding. [source]
- With `tool_choice="none"` vLLM still puts tool definitions in the prompt unless `--exclude-tools-when-tool-choice-none` is set. [source]
- The structural-tag RFC (issue 32142) argued that forcing `tool_choice="required"` through a JSON-schema array may force a first token of `[` and skip the model's native tool-call start token and its thinking. [source]
- Issue 53745 (closed 2026-08-26): with `--reasoning-parser qwen3 --tool-call-parser qwen3_coder`, `ParserManager.get_parser` returned the shared engine class directly (`parser_manager.py:142-143`), whose `adjust_request` only sets `skip_special_tokens = False`, so no structural tag attached and `strict: true` was a silent no-op. [source]
- The production symptom was malformed argument keys (`objctive`, `obj`, `objecTive`) from an agent client, and v0.27.1 had attached a populated triggered-tags tag for the same request. [source]
- A commenter confirmed PR 52830 (merged 2026-08-26) removed the short-circuit so `get_parser` composes a `DelegatingParser` and `_apply_structural_tag` runs when the tool parser has `structural_tag_model`, but said its tests do not build a strict request or use `qwen3_coder`, and PR 53752 with that test was still open. [source]
- Issue 54808 (closed 2026-09-29): on vLLM 0.28.0 with Qwen3.8-27B, `qwen3_coder` and `qwen3_xml` ignored `tool_choice` `"required"` and a named function, returning `finish_reason: "tool_calls"` with `tool_calls: null`, even with `strict: true`; the reporter links the same root cause to gemma4 (50477), poolside_v1 (49712) and step3p5 (51804). [source]
- A commenter stated 54808 was fixed on main by PR 52830 (v0.29.0 and later), and the issue was closed on 2026-09-29. [source]
- A commenter on 54808 reported `tool_choice: "none"` with no tools declared still let Qwen3.8 emit raw `<tool_call>` XML in `content` with `finish_reason: "tool_calls"` on `--tool-call-parser qwen3_coder`. [source]
- Issue 55152 (open): with `--structured-outputs-config '{"backend": "guidance"}'`, any request whose parser emits a structural tag fails with `VLLMValidationError: Invalid grammar specification: 'triggers'`, because tool parsers build new-format `{"type":"structural_tag","format":{...}}` tags via `vllm/tool_parsers/structural_tag_registry.py` while `serialize_guidance_grammar` only reads the legacy `triggers`/`structures` shape. [source]
- All 42 combinations of 14 registered `structural_tag_model` families by auto, required and named fail on guidance; PR 55153 makes hermes work in all three modes and llama and qwen_3 work for required, and rejects the rest with an error naming the shape (SequenceFormat, OptionalFormat, `excludes`). [source]
- With `VLLM_USE_RUST_FRONTEND=1`, `tool_choice: "required"` and named calls returned HTTP 500 on v0.30.0 with DeepSeek-V4.1-Flash and no structured-outputs flag, because the Rust frontend ignores `structured_outputs_config` and the requests reach the guidance path. [source]
- The author of vLLM's structured-output refactor RFC (48197) states xgrammar is the default for `backend=auto` and the strict tool-calling tags are built with xgrammar, and that spec-decode plus structured-output fixes (PRs 44297 and 44993) should make xgrammar at least as stable as guidance. [source]
- PR 59629 (open): when reasoning ends on a tool-call opener (GLM-4.7 and 5.x), the structured-output gate started the constraint one token past the opener, so the grammar re-forced a second opener and required or named calls came back empty; the fix feeds the opener to the grammar for content terminators, for the `glm47_moe`, `qwen3`, `gemma4` and `mistral` configs. [source]
- PR 56113 (open): with speculative decoding, `</think>` and a tool call can arrive in one chunk and `DelegatingParser.parse_delta` kept only the `</think>` text, losing the function name and arguments; stale `delta_token_ids` produced malformed arguments such as `{"command=ls": ""}`. [source]
- PR 58407 (open, labelled security): a closing tag copied inside a Qwen parameter value ended the value early and the model's own closers produced a second forged tool call; the fix detects more closers than openers and keeps the text in the first parameter. [source]
- Inferred: for a Qwen3.8 agent stack on vLLM 0.28 or earlier, strict and forced tool calling cannot be trusted; check `adjust_request(...).structured_outputs` for a non-empty `structural_tag` before relying on it. [source]
Children
- No children recorded.