<!-- llms-explorer concept facts · https://llms-explorer.com/tree/llama-cpp-reasoning-budget-sampler-and-lazy-gram/ · pack 2026-10-05 · ~652 tokens -->

# llama.cpp reasoning-budget sampler and lazy grammar trigger gating

> forge issue #54 (opened 2026-04-13): llama.cpp builds after 2026-04-10 (commit d7ff074) register thinking start and end tags for Gemma 4, which activates the reasoning-budget sampler

Parent: [Mac local LLMs: Chat templates, reasoning and tool calling](https://llms-explorer.com/tree/mac-local-llms-chat-templates-reasoning-and-tool-calling/) · 1 facets · 7 facts · page: https://llms-explorer.com/tree/llama-cpp-reasoning-budget-sampler-and-lazy-gram/

## Facts

- forge issue #54 (opened 2026-04-13): llama.cpp builds after 2026-04-10 (commit d7ff074) register thinking start and end tags for Gemma 4, which activates the reasoning-budget sampler — [source](https://github.com/antoinezambelli/forge/issues/54)
- forge #54: with `--reasoning-format auto` the auto-parser detects `<think>` tags from the template at run time (Qwen 3.5 27B and 35B confirmed), and Ministral Reasoning uses explicit `[THINK]` tags, so all three activate the sampler — [source](https://github.com/antoinezambelli/forge/issues/54)
- forge #54: with no `--reasoning-budget` given, the default budget is INT_MAX (2147483647 tokens), and the server log line reads 'reasoning-budget: activated, budget=2147483647 tokens' — [source](https://github.com/antoinezambelli/forge/issues/54)
- forge #54 symptom: some multi-turn tool-calling runs hung non-deterministically, filled the KV cache until it spilled to CPU RAM or swap, and crashed llama-server on memory-constrained hosts — [source](https://github.com/antoinezambelli/forge/issues/54)
- forge #54 measurements (dual 5070 Ti rig): Gemma 4 31B took 340 s unbounded versus 100 s with `--reasoning-budget 0` with identical quality on the task; a Qwen 3.5 27B run hung over 50 minutes versus about 75 s normally — [source](https://github.com/antoinezambelli/forge/issues/54)
- forge's fix was to default `--reasoning-budget 0` for all models in its server launcher, documented in BACKEND_SETUP.md and MODEL_GUIDE.md (commit 75953ab, 2026-04-16); on 2026-04-28 the author noted further llama.cpp commits and said they would re-evaluate — [source](https://github.com/antoinezambelli/forge/issues/54)
- Inferred: an update of llama.cpp can change inference behaviour for an unchanged command line, because the sampler arms itself whenever a model's template exposes think tags; the unbounded default is also the setting llama.cpp #24181 says keeps the sampler initialised so lazy grammar triggers inside thinking are suppressed, so `--reasoning-budget 0` also removes thinking for those models. — source: `asserted`
