<!-- llms-explorer concept facts · https://llms-explorer.com/tree/llama-fit-params-tool-and-freezing-fitted-flags/ · pack 2026-10-05 · ~1167 tokens -->

# llama-fit-params tool and freezing fitted flags into a launch script

> Fit adjusts `llama_model_params` and `llama_context_params`; the tool prints the changed arguments after "Printing fitted CLI arguments to stdout...".

Parent: [Mac local LLMs: llama.cpp internals](https://llms-explorer.com/tree/mac-local-llms-llama-cpp-internals/) · 1 facets · 21 facts · page: https://llms-explorer.com/tree/llama-fit-params-tool-and-freezing-fitted-flags/

## Facts

- Fit adjusts `llama_model_params` and `llama_context_params`; the tool prints the changed arguments after "Printing fitted CLI arguments to stdout...". — source: `asserted`
- Logs go to the console stream separate from stdout, which is why `| tee args.txt` captures only the arguments. — source: `asserted`
- Feeding the stored arguments to llama-server makes its own fit a no-op ("will leave 1199 >= 1024 MiB of free device memory, no changes needed"). — source: `asserted`
- 2025-12-15 (PR 16653): tool and README created. — source: `asserted`
- 2026-05-21 (PR 23459): fit-params, batched-bench, quantize and perplexity built as `app` subcommands. — source: `asserted`
- 2026-08-22 (PR 27496): fit-params source updated for `n_streams` and an optional second (draft/MTP) model. — source: `asserted`
- Frozen flags encode the free memory seen at that moment (24077 MiB free in the README run); they go stale if free memory, wired limit, KV type or slot flags change. — source: `asserted`
- `xargs` treats backslashes and parentheses specially; the printed `-ot blk\.N\.ffn_(up|down|gate)_(ch|)exps=CPU` strings may be altered by a naive `cat args.txt | xargs`. — source: `asserted`
- On Mac, generated `-ot ...=CPU` expert placements do not free GPU-visible memory because weights and Metal share one pool (held in llama-cpp-fit-auto-memory-fitting.md). — source: `asserted`
- README flow: fit once with the tool, then launch. Operators who hand-tune offload recommend `-fit off` (held elsewhere); a launch script that freezes output sits between the two. — source: `asserted`
- Whether the tool honors `-fitt` and `-fitc` on Metal identically to the server (no Mac log found). — source: `asserted`
- Whether freezing is worth it on a Mac, where the printed fit often says no changes are needed. — source: `asserted`
- The fit-params README says llama.cpp binaries can fit projected memory to free device memory at runtime via `-fit`/`--fit` arguments, and that `llama-fit-params` "prints the CLI arguments corresponding to these adjustments to stdout". — [source](https://github.com/ggml-org/llama.cpp/tree/master/tools/fit-params)
- The README example is `./build/bin/llama-fit-params --model <gguf> | tee args.txt` followed by `cat args.txt | xargs ./build/bin/llama-server --model <gguf>`. — [source](https://github.com/ggml-org/llama.cpp/tree/master/tools/fit-params)
- In the README run (Qwen3-30B-A3B F16, RTX 4090 24077 MiB free) fit projected 61807 MiB, reduced context from 40960 to 4096 (saving 3456 MiB) and printed `-c 4096 -ngl 48` plus one `-ot blk\.N\.ffn_(up|down|gate)_(ch|)exps=CPU` rule per overflowing layer (layers 14 to 47). — [source](https://github.com/ggml-org/llama.cpp/tree/master/tools/fit-params)
- Fitting took 1.15 s in the tool run; the later llama-server run with the stored flags logged "will leave 1199 >= 1024 MiB of free device memory, no changes needed" and fitting took 0.28 s. — [source](https://github.com/ggml-org/llama.cpp/tree/master/tools/fit-params)
- The tool's `main.cpp` entered the tree through PR 23459 (2026-05-21, "app : add batched-bench, fit-params, quantize & perplexity"). — [source](https://github.com/ggml-org/llama.cpp/tree/master/tools/fit-params)
- The latest `fit-params.cpp` commit (2fb989b, 2026-08-22, PR 27496) takes `n_streams` into account and accepts an optional second model so a draft or MTP context is measured at the largest target context. — [source](https://github.com/ggml-org/llama.cpp/tree/master/tools/fit-params)
- PR 27496 makes `--fit -no-kvu -np 4` choose `n_ctx = n_ctx_train * 4`, so each slot receives the full trained context. — [source](https://github.com/ggml-org/llama.cpp/pull/27496)
- Inferred: run `llama-fit-params` with the same `-fa`, `-ctk/-ctv`, `-np`, `--kv-unified` and `-fitt` the server will use, because each changes projected memory; the printed flags are only valid for that combination. — source: `asserted`
- Inferred: in a launch script, store the printed line in a file next to the model, re-generate it when macOS memory settings (`iogpu.wired_limit_mb`) or KV flags change, and pass it through `xargs -0` or a bash array to avoid backslash mangling of the `-ot` regexes. — source: `asserted`
