<!-- llms-explorer concept facts · https://llms-explorer.com/tree/non-canonical-tokenization-robustness-of-instruc/ · pack 2026-10-05 · ~992 tokens -->

# Non-canonical tokenization robustness of instruction-tuned models

> A July 2026 study of 27 languages and six downstream tasks finds the English invariance does not generalize: instruction-tuned models lose on average 23.7% relative performance for Llama-3.1-8B, 11.4% for Qwen3-8B and 9.9% for Gemma-3-12B under alternative tokenizations.

Parent: [Mac local LLMs: Quantization evaluation](https://llms-explorer.com/tree/mac-local-llms-quantization-evaluation/) · 1 facets · 18 facts · page: https://llms-explorer.com/tree/non-canonical-tokenization-robustness-of-instruc/

## Facts

- A July 2026 study of 27 languages and six downstream tasks finds the English invariance does not generalize: instruction-tuned models lose on average 23.7% relative performance for Llama-3.1-8B, 11.4% for Qwen3-8B and 9.9% for Gemma-3-12B under alternative tokenizations. — [source](https://arxiv.org/html/2607.26831v1)
- Sensitivity tracks token fragmentation: languages whose canonical tokenization splits into more tokens are significantly more sensitive. — [source](https://arxiv.org/html/2607.26831v1)
- The effect reaches embedding tasks: cross-lingual sentence retrieval is also hurt. — [source](https://arxiv.org/html/2607.26831v1)
- LoRA fine-tuning on multi-tokenization data mitigates it; fine-tuning on English alone already improves robustness across languages, and sampling diverse non-canonical tokenizations works best. — [source](https://arxiv.org/html/2607.26831v1)
- The authors name tokenizer adaptation (vocabulary pruning, tokenizer transfer) as a real source of non-canonical inputs after pretraining. — [source](https://arxiv.org/html/2607.26831v1)
- For a Mac harness this means a token-array prompt path or a retokenize round trip is safe for English instruct models by the 2025 result but not assumed safe for non-English prompts. — source: `asserted`
- No source tests whether quantization (GGUF or MLX) changes tokenization robustness; both papers evaluate full-precision models. — source: `asserted`
- 2025: Zheng et al., "Broken Tokens?" (NeurIPS 2025 spotlight; 20 benchmarks). 2026-07: multilingual follow-up above. — [source](https://neurips.cc/virtual/2025/poster/117561)
- The two papers use different models and task sets, so 93.4% retention (English, Qwen-2.5-7B-Instruct per the covering file) and 9.9% to 23.7% drops (multilingual, three other models) are not directly comparable. — source: `asserted`
- The new paper's figures are relative drops averaged over languages; per-language drops were not read beyond the abstract. — source: `asserted`
- English-only robustness (Zheng et al.: instruct models "secretly handle" non-canonical input) versus multilingual fragility (the 2607.26831 abstract: "Models are broken in multilingual settings"). Both hold for their own scope. — [source](https://arxiv.org/html/2607.26831v1)
- Whether a 4-bit quant of Llama-3.1-8B, Qwen3-8B or Gemma-3-12B drops more under non-canonical tokenization than its BF16 base. — source: `asserted`
- Whether llama-server or mlx-lm re-tokenization of a chat history changes ids often enough to matter for non-English agents. — source: `asserted`
- Instruction-tuned models showed average relative performance drops of 23.7% (Llama-3.1-8B), 11.4% (Qwen3-8B) and 9.9% (Gemma-3-12B) under non-canonical tokenization across 27 languages and six tasks. — [source](https://arxiv.org/html/2607.26831v1)
- Languages with higher token fragmentation are more sensitive to non-canonical tokenization. — [source](https://arxiv.org/html/2607.26831v1)
- LoRA fine-tuning with multi-tokenization data mitigates the sensitivity. — [source](https://arxiv.org/html/2607.26831v1)
- The authors state Zheng et al. (2025) found English robustness to non-canonical tokenization and hypothesize it emerges in post-training. — [source](https://arxiv.org/html/2607.26831v1)
- The English result is already held in token-level-generation-versus-render-drift-measu.md. — source: `asserted`
