llama.cpp reasoning-budget sampler and lazy grammar trigger gating
Parent: Mac local LLMs: Chat templates, reasoning and tool calling · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
forge issue #54 (opened 2026-04-13): llama.cpp builds after 2026-04-10 (commit d7ff074) register thinking start and end tags for Gemma 4, which activates the reasoning-budget sampler
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- forge issue #54 (opened 2026-04-13): llama.cpp builds after 2026-04-10 (commit d7ff074) register thinking start and end tags for Gemma 4, which activates the reasoning-budget sampler [source]
- forge #54: with `--reasoning-format auto` the auto-parser detects `<think>` tags from the template at run time (Qwen 3.5 27B and 35B confirmed), and Ministral Reasoning uses explicit `[THINK]` tags, so all three activate the sampler [source]
- forge #54: with no `--reasoning-budget` given, the default budget is INT_MAX (2147483647 tokens), and the server log line reads 'reasoning-budget: activated, budget=2147483647 tokens' [source]
- forge #54 symptom: some multi-turn tool-calling runs hung non-deterministically, filled the KV cache until it spilled to CPU RAM or swap, and crashed llama-server on memory-constrained hosts [source]
- forge #54 measurements (dual 5070 Ti rig): Gemma 4 31B took 340 s unbounded versus 100 s with `--reasoning-budget 0` with identical quality on the task; a Qwen 3.5 27B run hung over 50 minutes versus about 75 s normally [source]
- forge's fix was to default `--reasoning-budget 0` for all models in its server launcher, documented in BACKEND_SETUP.md and MODEL_GUIDE.md (commit 75953ab, 2026-04-16); on 2026-04-28 the author noted further llama.cpp commits and said they would re-evaluate [source]
- Inferred: an update of llama.cpp can change inference behaviour for an unchanged command line, because the sampler arms itself whenever a model's template exposes think tags; the unbounded default is also the setting llama.cpp #24181 says keeps the sampler initialised so lazy grammar triggers inside thinking are suppressed, so `--reasoning-budget 0` also removes thinking for those models. [source]
Children
- No children recorded.