<!-- llms-explorer concept facts · https://llms-explorer.com/tree/mlxguidedgeneration-and-xgrammar-constrained-dec/ · pack 2026-10-05 · ~1331 tokens -->

# MLXGuidedGeneration and XGrammar constrained decoding on Apple Silicon

> MLXGuidedGeneration standalone setup is four steps: extract a vocab with `TokenizerVocabExtractor.extractForGrammar`, build a `GrammarTokenizer` (vocab, vocabType, eosTokenId), build a `GrammarConstraint` from a JSON schema string, then call `GuidedGenerationLoop.run` with maxTokens and vocabSize

Parent: [Mac local LLMs: Apple Foundation Models and Core AI](https://llms-explorer.com/tree/mac-local-llms-apple-foundation-models-and-core-ai/) · 1 facets · 15 facts · page: https://llms-explorer.com/tree/mlxguidedgeneration-and-xgrammar-constrained-dec/

## Facts

- MLXGuidedGeneration standalone setup is four steps: extract a vocab with `TokenizerVocabExtractor.extractForGrammar`, build a `GrammarTokenizer` (vocab, vocabType, eosTokenId), build a `GrammarConstraint` from a JSON schema string, then call `GuidedGenerationLoop.run` with maxTokens and vocabSize — [source](https://github.com/ml-explore/mlx-swift-lm/blob/main/Libraries/MLXGuidedGeneration/README.md)
- `GrammarConstraint` takes a `fastForward: true` option and a `hostTokenizer` argument in the README example — [source](https://github.com/ml-explore/mlx-swift-lm/blob/main/Libraries/MLXGuidedGeneration/README.md)
- README warning: the first `GuidedGenerationLoop.run` for a schema and tokenizer can block for hundreds of milliseconds because grammar compile and token-mask build do not yield; do not call it on `@MainActor`, run it in `Task.detached`, or pre-warm by building a throwaway `GrammarConstraint` on a background task — [source](https://github.com/ml-explore/mlx-swift-lm/blob/main/Libraries/MLXGuidedGeneration/README.md)
- Later calls that reuse the same compiled grammar and tokenizer skip the compile — [source](https://github.com/ml-explore/mlx-swift-lm/blob/main/Libraries/MLXGuidedGeneration/README.md)
- Through MLXFoundationModels guided generation is requested with the `.guidedGeneration` capability on `#huggingFaceLanguageModel(configuration:capabilities:)` and a `@Generable` type passed to `respond(generating:)` — [source](https://github.com/ml-explore/mlx-swift-lm/blob/main/Libraries/MLXGuidedGeneration/README.md)
- XGrammar's `get_model_structural_tag` builds tool-call grammars for these styles: llama, qwen_3, qwen_3_5, qwen_3_coder, kimi, kimi_k3, deepseek_r1, deepseek_v3_1, deepseek_v3_2, deepseek_v4, deepseek_v4_1, harmony, minimax, minimax_m3, glm_4_7, mimo, cohere, exaone, gemma_4 — [source](https://xgrammar.mlc.ai/docs/latest/structural_tag/tool_calling_and_reasoning.html)
- `get_model_structural_tag` takes `reasoning` as "enabled", "disabled" or "auto" (booleans deprecated); "auto" means the prompt does not prefill the reasoning opener and the model may emit one block or go straight to an answer or tool call; default is "enabled" — [source](https://xgrammar.mlc.ai/docs/latest/structural_tag/tool_calling_and_reasoning.html)
- `get_model_structural_tag` options include `tool_choice` (auto, none, required, forced function, forced builtin, allowed_tools), `parallel_tool_calls` (default true; false allows at most one call and generation must end after it), `any_order` for argument property order, and `max_whitespace_cnt` to bound whitespace runs that blow up compile or match time — [source](https://xgrammar.mlc.ai/docs/latest/structural_tag/tool_calling_and_reasoning.html)
- For the harmony style, builtin tools such as `web_search_preview` can be listed beside function tools, with a model-output `name` like `browser.search` — [source](https://xgrammar.mlc.ai/docs/latest/structural_tag/tool_calling_and_reasoning.html)
- XGrammar mask generation runs on the CPU and can overlap the GPU forward pass; grammar compile can overlap prefill — [source](https://xgrammar.mlc.ai/docs/latest/using_xgrammar/engine_integration.html)
- XGrammar per-step integration is three calls: `fill_next_token_bitmask` (returns `need_apply`; false means skip masking), `apply_token_bitmask_inplace` (sets invalid logits to -inf), and `accept_token` (returns False only if the mask was not applied) — [source](https://xgrammar.mlc.ai/docs/latest/using_xgrammar/engine_integration.html)
- XGrammar offers jump-forward decoding (`find_jump_forward_string` returns text that must follow, appended without decoding), `rollback` and `fork` of matcher state, and `traverse_draft_tree` for speculative decoding (API first proposed by the SGLang team) — [source](https://xgrammar.mlc.ai/docs/latest/using_xgrammar/engine_integration.html)
- XGrammar lists SGLang, vLLM, TensorRT-LLM and MLC-LLM as engines with integrations; MLX and llama.cpp are not listed there — [source](https://xgrammar.mlc.ai/docs/latest/using_xgrammar/engine_integration.html)
- Inferred: on Apple Silicon the CPU-side mask overlap with the GPU forward pass should hold because CPU and GPU share unified memory, but no source measures masking cost or the effect of `fastForward` on MLX decode speed. — source: `asserted`
- Inferred: XGrammar's gemma_4 and harmony structural tags could constrain the tool-call text of Gemma 4 and gpt-oss on MLX, which would address the malformed-call loops reported for Gemma 4, but nothing read shows MLXGuidedGeneration exposing structural tags for tool calls beyond the README line that it accepts a structural tag. — source: `asserted`
