oMLX per-model chat template kwargs and cache-key construction
Parent: Mac local LLMs: Chat templates, reasoning and tool calling · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Merge order, lowest to highest (model_settings.py `merge_chat_template_kwargs`): (1) the model's stored `chat_template_kwargs`; (2) the dedicated model fields `enable_thinking` and `preserve_thinking` when not None; (3) request `chat_template_kwargs`, except keys listed in the model's `forced_ct_...
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Merge order, lowest to highest (model_settings.py `merge_chat_template_kwargs`): (1) the model's stored `chat_template_kwargs`; (2) the dedicated model fields `enable_thinking` and `preserve_thinking` when not None; (3) request `chat_template_kwargs`, except keys listed in the model's `forced_ct_kwargs`; (4) `enable_thinking=true` when a thinking budget is active and the key is still unset; (5) `preserve_thinking=true` when the model's detected default is true, `enable_thinking` is not false, and the key is still unset. [source]
- `forced_ct_kwargs` is a per-model list of keys a request may not override; request values for those keys are dropped silently. [source]
- Detection of the preserve default (model_discovery.py `detect_preserve_thinking`): read `chat_template.jinja` from the model directory, else the `chat_template` string in `tokenizer_config.json`; if the text contains the substring `preserve_thinking`, the model's default is true. This is a plain substring test, so a template that mentions the name only in a comment also qualifies. The default is stored in the model entry as `preserve_thinking_default`. [source]
- Consequence: for Qwen3.6 and Qwen3.8 templates oMLX turns preserved thinking on by itself, which is what the held dossiers listed as an open question. The same code sets the output-caching flag (see reasoning-token-cache-invalidation-across-turns.md). [source]
- API reasoning effort enters through `merge_reasoning_effort_chat_template_kwargs` with `setdefault`, so an explicit kwarg in settings or request wins over the API field. [source]
- Render call: `tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True, [tools], **kwargs)`. If it raises `TypeError`, oMLX removes every key the user supplied, `tools`, and `enable_thinking`, and renders again; a template with strict signature therefore silently loses the user's kwargs on the retry. The reasoning-effort path has its own retry (see qwen-3-8-chat-template-fixes-and-reasoning-level.md). [source]
- Cache key (omlx/cache/paged_cache.py `compute_block_hash`): SHA-256 over the model name (when given), the parent block hash (or the constant `omlx-root` for the first block), the block's token ids rendered as a tuple string, and optional `extra_keys` (used for LoRA and multimodal feature hashes). Chat-template kwargs, effort, profile name and `cache_control` markers are not inputs. [source]
- Because each block hash includes its parent's hash, one changed token invalidates that block and every later block of the same prefix chain. [source]
- Prefix matching in the scheduler compares token sequences (`_common_prefix_len`) and floors reuse to the paged block size (a reviewer cites 2048 for a Qwen3.8 27B hybrid). [source]
- Admin UI: a saved profile template matches a profile by its `chat_template_kwargs` regardless of nested key order, and a changed value breaks the match (tests/test_admin_model_settings_template.py). [source]
- Per-model kwargs, profiles and `forced_ct_kwargs` are on main as of 2026-10-04; introduction dates were not found. [source]
- Two profiles of one model that render the same tokens share cache blocks only if `model_name` in the hash is the base model name; whether it is the base name or the profile alias was not read. [source]
- Setting `preserve_thinking=false` per model overrides the auto default and also turns off output caching, so the setting changes both the prompt bytes and what oMLX stores. [source]
- A template mentioning `preserve_thinking` only in a comment gets preserved-thinking forced on; a template using a differently named key (for example `clear_thinking` or `drop_thinking`, the other names llama.cpp aliases) gets no auto default. [source]
- The TypeError retry drops all user kwargs, not just the one that failed, so a typo in one kwarg can silently disable `enable_thinking` and `preserve_thinking` and change the prompt. [source]
- Because the cache key ignores kwargs, switching `enable_thinking` between requests does not need explicit invalidation: the changed generation prompt changes the token ids at the tail only, and earlier blocks stay shared. [source]
- Whether `model_name` in the block hash is the base model or the profile alias. [source]
- Whether the Anthropic route and the OpenAI route pass identical merged kwargs (both call `merge_chat_template_kwargs`, but the Anthropic adapter was not read line by line). [source]
- Whether profiles saved through the admin panel are applied to API-key-scoped requests differently from the default profile. [source]
- oMLX merges model chat_template_kwargs, then dedicated enable_thinking and preserve_thinking toggles, then request kwargs except forced keys, then a thinking-budget enable_thinking default, then the detected preserve_thinking default [source]
- ModelSettings has a forced_ct_kwargs list of chat-template keys that API requests cannot override [source]
- detect_preserve_thinking reads chat_template.jinja, else tokenizer_config.json, and returns true when the text contains the substring preserve_thinking [source]
- oMLX sets preserve_thinking=true automatically when the model's detected default is true, enable_thinking is not false and the key is unset [source]
- merge_reasoning_effort_chat_template_kwargs uses setdefault so explicit kwargs override the API reasoning effort [source]
- On TypeError from apply_chat_template oMLX removes all user kwargs, tools and enable_thinking and renders again [source]
- compute_block_hash hashes model name, parent block hash (or omlx-root), the block's token ids and optional extra keys with SHA-256 [source]
- A reviewer on 3608 reads cache_control as unused by oMLX's prefix comparison, which works on token sequences [source]
- compute_block_hash takes no chat-template kwargs, effort or profile argument [source]
- A reviewer on 3608 reads the scheduler as flooring reused prefix length to the paged block size, 2048 for the reporter's Qwen3.8 27B model [source]
- oMLX's mid-system probe cache key includes the frozen chat-template kwargs [source]
- Chain hashing means one changed token invalidates every later block of that prefix [source]
- The substring test for preserve_thinking cannot tell a live use of the variable from a mention in a comment [source]
Children
- No children recorded.