MLXGuidedGeneration and XGrammar constrained decoding on Apple Silicon
Parent: Mac local LLMs: Apple Foundation Models and Core AI · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
MLXGuidedGeneration standalone setup is four steps: extract a vocab with `TokenizerVocabExtractor.extractForGrammar`, build a `GrammarTokenizer` (vocab, vocabType, eosTokenId), build a `GrammarConstraint` from a JSON schema string, then call `GuidedGenerationLoop.run` with maxTokens and vocabSize
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- MLXGuidedGeneration standalone setup is four steps: extract a vocab with `TokenizerVocabExtractor.extractForGrammar`, build a `GrammarTokenizer` (vocab, vocabType, eosTokenId), build a `GrammarConstraint` from a JSON schema string, then call `GuidedGenerationLoop.run` with maxTokens and vocabSize [source]
- `GrammarConstraint` takes a `fastForward: true` option and a `hostTokenizer` argument in the README example [source]
- README warning: the first `GuidedGenerationLoop.run` for a schema and tokenizer can block for hundreds of milliseconds because grammar compile and token-mask build do not yield; do not call it on `@MainActor`, run it in `Task.detached`, or pre-warm by building a throwaway `GrammarConstraint` on a background task [source]
- Later calls that reuse the same compiled grammar and tokenizer skip the compile [source]
- Through MLXFoundationModels guided generation is requested with the `.guidedGeneration` capability on `#huggingFaceLanguageModel(configuration:capabilities:)` and a `@Generable` type passed to `respond(generating:)` [source]
- XGrammar's `get_model_structural_tag` builds tool-call grammars for these styles: llama, qwen_3, qwen_3_5, qwen_3_coder, kimi, kimi_k3, deepseek_r1, deepseek_v3_1, deepseek_v3_2, deepseek_v4, deepseek_v4_1, harmony, minimax, minimax_m3, glm_4_7, mimo, cohere, exaone, gemma_4 [source]
- `get_model_structural_tag` takes `reasoning` as "enabled", "disabled" or "auto" (booleans deprecated); "auto" means the prompt does not prefill the reasoning opener and the model may emit one block or go straight to an answer or tool call; default is "enabled" [source]
- `get_model_structural_tag` options include `tool_choice` (auto, none, required, forced function, forced builtin, allowed_tools), `parallel_tool_calls` (default true; false allows at most one call and generation must end after it), `any_order` for argument property order, and `max_whitespace_cnt` to bound whitespace runs that blow up compile or match time [source]
- For the harmony style, builtin tools such as `web_search_preview` can be listed beside function tools, with a model-output `name` like `browser.search` [source]
- XGrammar mask generation runs on the CPU and can overlap the GPU forward pass; grammar compile can overlap prefill [source]
- XGrammar per-step integration is three calls: `fill_next_token_bitmask` (returns `need_apply`; false means skip masking), `apply_token_bitmask_inplace` (sets invalid logits to -inf), and `accept_token` (returns False only if the mask was not applied) [source]
- XGrammar offers jump-forward decoding (`find_jump_forward_string` returns text that must follow, appended without decoding), `rollback` and `fork` of matcher state, and `traverse_draft_tree` for speculative decoding (API first proposed by the SGLang team) [source]
- XGrammar lists SGLang, vLLM, TensorRT-LLM and MLC-LLM as engines with integrations; MLX and llama.cpp are not listed there [source]
- Inferred: on Apple Silicon the CPU-side mask overlap with the GPU forward pass should hold because CPU and GPU share unified memory, but no source measures masking cost or the effect of `fastForward` on MLX decode speed. [source]
- Inferred: XGrammar's gemma_4 and harmony structural tags could constrain the tool-call text of Gemma 4 and gpt-oss on MLX, which would address the malformed-call loops reported for Gemma 4, but nothing read shows MLXGuidedGeneration exposing structural tags for tool calls beyond the README line that it accepts a structural tag. [source]
Children
- No children recorded.