Chat-template patching for cache stability
Parent: Mac local LLMs: Prompt cache and persistent KV · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
mlx-lm: some Qwen3 2507 repos ship chat_template.jinja only, so tokenizer.chat_template is None after conversion; check the file exists in the local model dir before assuming a patch applies.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- mlx-lm: some Qwen3 2507 repos ship chat_template.jinja only, so tokenizer.chat_template is None after conversion; check the file exists in the local model dir before assuming a patch applies. [source]
- Ollama: TEMPLATE is Go text/template, not Jinja, so a Jinja fix cannot be pasted; community Jinja fixes must be re-expressed. [source]
- No source documents a render-twice-and-diff procedure, a sort_keys guard on tool arguments in any shipped community template, or an mlx_lm.server flag for a template file. [source]
- Whether LM Studio's Prompt Template box takes effect for MLX-engine models (as opposed to GGUF) is not stated in its docs. [source]
- Trust ranking of community templates cannot be sourced beyond authorship and license facts below. [source]
- llama-server accepts --chat-template-file JINJA_TEMPLATE_FILE; the model's embedded metadata template is the default, and --jinja is needed before non-built-in templates are accepted [source]
- llama-server's README says tool calling may require a --chat-template-file override to get a tool-compatible template [source]
- llama-server exposes the active template via the props endpoint fields chat_template and chat_template_caps, which lets a client confirm which template is loaded [source]
- llama-server requests accept chat_template_kwargs, e.g. {"enable_thinking": false} [source]
- Ollama overrides a template with the Modelfile TEMPLATE instruction, written in Go text/template with variables .System, .Prompt, .Response [source]
- The Ollama TEMPLATE variables documented are single-turn (.System, .Prompt, .Response); the Modelfile page lists no tool-call or reasoning variables [source]
- Ollama's stock qwen3-coder template was reported (issue #11621, 2025-08) as lacking tools and FIM support, with 66+ thumbs-up [source]
- LM Studio lets you override a model's prompt template in My Models > gear icon, in Jinja or "Manual" role prefix/suffix form; the box is hidden unless the model lacks template metadata or you enable "Always Show Prompt Template" [source]
- LM Studio's own doc says that in most cases you do not need to change the template [source]
- mlx-lm issue #340 (2025-07-31, closed) records Qwen3 2507 models shipping chat_template.jinja instead of a tokenizer_config.json entry, giving tokenizer.chat_template == None after MLX conversion [source]
- froggeric's own install steps: llama.cpp uses `--jinja --chat-template-file chat_template.jinja --reasoning-format deepseek`; LM Studio uses paste into the Prompt Template box and Save; MLX/oMLX uses overwrite chat_template.jinja in the local model directory; vLLM uses replace the chat_template string in tokenizer_config.json [source]
- froggeric template v22.5 is one file for Qwen 3.5, 3.6 and 3.8, Apache-2.0, three named contributors plus a barubary/spiritbuun C++ AST contribution, and ships scripts/check_applied.py to inspect model dirs and GGUFs for the active template version [source]
- froggeric's test suite claims 105 verification cells including "prefix KV cache stability" (run via scripts/test_v22.py); the cache-hit rate itself is not measured in the page [source]
- froggeric's template renders JSON-string tool arguments in history without crashing and merges consecutive leading system messages into one system turn [source]
- A community mlx-community Qwen3-30B-A3B discussion (2025-04-29) offers a replacement template as a quick fix for a Jinja error in LM Studio; it renders tools with `tool | tojson` and no sort_keys argument [source]
- Trust level of community templates is single-maintainer, unaudited by Qwen or runtime vendors; verify by diffing against the official template and rendering your own transcript before adopting [source]
- Prefix-stability check: render the same message list through apply_chat_template at turn N and turn N+1 (turn N+1 = turn N plus the assistant reply and next user message, add_generation_prompt false for the first), and assert the turn-N string is a prefix of the turn-N+1 string; the first differing character index locates the offender [source]
- Ordering guard: in a patched template, pass tool arguments through `tojson(sort_keys=True)` (or iterate `dictsort`) only if the client never reorders keys itself; sorting changes bytes versus earlier-cached prompts, so the first request after patching is a full re-prefill [source]
- Empty-think guard: emit a `<think>` wrapper for a past assistant turn only when its reasoning text is non-empty, so tool-call-only turns render identically each time [source]
Children
- No children recorded.