<!-- llms-explorer concept facts · https://llms-explorer.com/tree/chat-template-patching-for-cache-stability/ · pack 2026-10-05 · ~1509 tokens -->

# Chat-template patching for cache stability

> mlx-lm: some Qwen3 2507 repos ship chat_template.jinja only, so tokenizer.chat_template is None after conversion; check the file exists in the local model dir before assuming a patch applies.

Parent: [Mac local LLMs: Prompt cache and persistent KV](https://llms-explorer.com/tree/mac-local-llms-prompt-cache-and-persistent-kv/) · 1 facets · 24 facts · page: https://llms-explorer.com/tree/chat-template-patching-for-cache-stability/

## Facts

- mlx-lm: some Qwen3 2507 repos ship chat_template.jinja only, so tokenizer.chat_template is None after conversion; check the file exists in the local model dir before assuming a patch applies. — source: `asserted`
- Ollama: TEMPLATE is Go text/template, not Jinja, so a Jinja fix cannot be pasted; community Jinja fixes must be re-expressed. — source: `asserted`
- No source documents a render-twice-and-diff procedure, a sort_keys guard on tool arguments in any shipped community template, or an mlx_lm.server flag for a template file. — source: `asserted`
- Whether LM Studio's Prompt Template box takes effect for MLX-engine models (as opposed to GGUF) is not stated in its docs. — source: `asserted`
- Trust ranking of community templates cannot be sourced beyond authorship and license facts below. — source: `asserted`
- llama-server accepts --chat-template-file JINJA_TEMPLATE_FILE; the model's embedded metadata template is the default, and --jinja is needed before non-built-in templates are accepted — [source](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md)
- llama-server's README says tool calling may require a --chat-template-file override to get a tool-compatible template — [source](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md)
- llama-server exposes the active template via the props endpoint fields chat_template and chat_template_caps, which lets a client confirm which template is loaded — [source](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md)
- llama-server requests accept chat_template_kwargs, e.g. {"enable_thinking": false} — [source](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md)
- Ollama overrides a template with the Modelfile TEMPLATE instruction, written in Go text/template with variables .System, .Prompt, .Response — [source](https://docs.ollama.com/modelfile)
- The Ollama TEMPLATE variables documented are single-turn (.System, .Prompt, .Response); the Modelfile page lists no tool-call or reasoning variables — [source](https://docs.ollama.com/modelfile)
- Ollama's stock qwen3-coder template was reported (issue #11621, 2025-08) as lacking tools and FIM support, with 66+ thumbs-up — [source](https://github.com/ollama/ollama/issues/11621)
- LM Studio lets you override a model's prompt template in My Models > gear icon, in Jinja or "Manual" role prefix/suffix form; the box is hidden unless the model lacks template metadata or you enable "Always Show Prompt Template" — [source](https://lmstudio.ai/docs/app/advanced/prompt-template)
- LM Studio's own doc says that in most cases you do not need to change the template — [source](https://lmstudio.ai/docs/app/advanced/prompt-template)
- mlx-lm issue #340 (2025-07-31, closed) records Qwen3 2507 models shipping chat_template.jinja instead of a tokenizer_config.json entry, giving tokenizer.chat_template == None after MLX conversion — [source](https://github.com/ml-explore/mlx-lm/issues/340)
- froggeric's own install steps: llama.cpp uses `--jinja --chat-template-file chat_template.jinja --reasoning-format deepseek`; LM Studio uses paste into the Prompt Template box and Save; MLX/oMLX uses overwrite chat_template.jinja in the local model directory; vLLM uses replace the chat_template string in tokenizer_config.json — [source](https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates)
- froggeric template v22.5 is one file for Qwen 3.5, 3.6 and 3.8, Apache-2.0, three named contributors plus a barubary/spiritbuun C++ AST contribution, and ships scripts/check_applied.py to inspect model dirs and GGUFs for the active template version — [source](https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates)
- froggeric's test suite claims 105 verification cells including "prefix KV cache stability" (run via scripts/test_v22.py); the cache-hit rate itself is not measured in the page — [source](https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates)
- froggeric's template renders JSON-string tool arguments in history without crashing and merges consecutive leading system messages into one system turn — [source](https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates)
- A community mlx-community Qwen3-30B-A3B discussion (2025-04-29) offers a replacement template as a quick fix for a Jinja error in LM Studio; it renders tools with `tool | tojson` and no sort_keys argument — [source](https://huggingface.co/mlx-community/Qwen3-30B-A3B-4bit/discussions/1)
- Trust level of community templates is single-maintainer, unaudited by Qwen or runtime vendors; verify by diffing against the official template and rendering your own transcript before adopting — source: `asserted`
- Prefix-stability check: render the same message list through apply_chat_template at turn N and turn N+1 (turn N+1 = turn N plus the assistant reply and next user message, add_generation_prompt false for the first), and assert the turn-N string is a prefix of the turn-N+1 string; the first differing character index locates the offender — source: `asserted`
- Ordering guard: in a patched template, pass tool arguments through `tojson(sort_keys=True)` (or iterate `dictsort`) only if the client never reorders keys itself; sorting changes bytes versus earlier-cached prompts, so the first request after patching is a full re-prefill — source: `asserted`
- Empty-think guard: emit a `<think>` wrapper for a past assistant turn only when its reasoning text is non-empty, so tool-call-only turns render identically each time — source: `asserted`
