Non-canonical tokenization robustness of instruction-tuned models
Parent: Mac local LLMs: Quantization evaluation · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
A July 2026 study of 27 languages and six downstream tasks finds the English invariance does not generalize: instruction-tuned models lose on average 23.7% relative performance for Llama-3.1-8B, 11.4% for Qwen3-8B and 9.9% for Gemma-3-12B under alternative tokenizations.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- A July 2026 study of 27 languages and six downstream tasks finds the English invariance does not generalize: instruction-tuned models lose on average 23.7% relative performance for Llama-3.1-8B, 11.4% for Qwen3-8B and 9.9% for Gemma-3-12B under alternative tokenizations. [source]
- Sensitivity tracks token fragmentation: languages whose canonical tokenization splits into more tokens are significantly more sensitive. [source]
- The effect reaches embedding tasks: cross-lingual sentence retrieval is also hurt. [source]
- LoRA fine-tuning on multi-tokenization data mitigates it; fine-tuning on English alone already improves robustness across languages, and sampling diverse non-canonical tokenizations works best. [source]
- The authors name tokenizer adaptation (vocabulary pruning, tokenizer transfer) as a real source of non-canonical inputs after pretraining. [source]
- For a Mac harness this means a token-array prompt path or a retokenize round trip is safe for English instruct models by the 2025 result but not assumed safe for non-English prompts. [source]
- No source tests whether quantization (GGUF or MLX) changes tokenization robustness; both papers evaluate full-precision models. [source]
- 2025: Zheng et al., "Broken Tokens?" (NeurIPS 2025 spotlight; 20 benchmarks). 2026-07: multilingual follow-up above. [source]
- The two papers use different models and task sets, so 93.4% retention (English, Qwen-2.5-7B-Instruct per the covering file) and 9.9% to 23.7% drops (multilingual, three other models) are not directly comparable. [source]
- The new paper's figures are relative drops averaged over languages; per-language drops were not read beyond the abstract. [source]
- English-only robustness (Zheng et al.: instruct models "secretly handle" non-canonical input) versus multilingual fragility (the 2607.26831 abstract: "Models are broken in multilingual settings"). Both hold for their own scope. [source]
- Whether a 4-bit quant of Llama-3.1-8B, Qwen3-8B or Gemma-3-12B drops more under non-canonical tokenization than its BF16 base. [source]
- Whether llama-server or mlx-lm re-tokenization of a chat history changes ids often enough to matter for non-English agents. [source]
- Instruction-tuned models showed average relative performance drops of 23.7% (Llama-3.1-8B), 11.4% (Qwen3-8B) and 9.9% (Gemma-3-12B) under non-canonical tokenization across 27 languages and six tasks. [source]
- Languages with higher token fragmentation are more sensitive to non-canonical tokenization. [source]
- LoRA fine-tuning with multi-tokenization data mitigates the sensitivity. [source]
- The authors state Zheng et al. (2025) found English robustness to non-canonical tokenization and hypothesize it emerges in post-training. [source]
- The English result is already held in token-level-generation-versus-render-drift-measu.md. [source]
Children
- No children recorded.