llama-fit-params tool and freezing fitted flags into a launch script
Parent: Mac local LLMs: llama.cpp internals · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Fit adjusts `llama_model_params` and `llama_context_params`; the tool prints the changed arguments after "Printing fitted CLI arguments to stdout...".
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Fit adjusts `llama_model_params` and `llama_context_params`; the tool prints the changed arguments after "Printing fitted CLI arguments to stdout...". [source]
- Logs go to the console stream separate from stdout, which is why `| tee args.txt` captures only the arguments. [source]
- Feeding the stored arguments to llama-server makes its own fit a no-op ("will leave 1199 >= 1024 MiB of free device memory, no changes needed"). [source]
- 2025-12-15 (PR 16653): tool and README created. [source]
- 2026-05-21 (PR 23459): fit-params, batched-bench, quantize and perplexity built as `app` subcommands. [source]
- 2026-08-22 (PR 27496): fit-params source updated for `n_streams` and an optional second (draft/MTP) model. [source]
- Frozen flags encode the free memory seen at that moment (24077 MiB free in the README run); they go stale if free memory, wired limit, KV type or slot flags change. [source]
- `xargs` treats backslashes and parentheses specially; the printed `-ot blk\.N\.ffn_(up|down|gate)_(ch|)exps=CPU` strings may be altered by a naive `cat args.txt | xargs`. [source]
- On Mac, generated `-ot ...=CPU` expert placements do not free GPU-visible memory because weights and Metal share one pool (held in llama-cpp-fit-auto-memory-fitting.md). [source]
- README flow: fit once with the tool, then launch. Operators who hand-tune offload recommend `-fit off` (held elsewhere); a launch script that freezes output sits between the two. [source]
- Whether the tool honors `-fitt` and `-fitc` on Metal identically to the server (no Mac log found). [source]
- Whether freezing is worth it on a Mac, where the printed fit often says no changes are needed. [source]
- The fit-params README says llama.cpp binaries can fit projected memory to free device memory at runtime via `-fit`/`--fit` arguments, and that `llama-fit-params` "prints the CLI arguments corresponding to these adjustments to stdout". [source]
- The README example is `./build/bin/llama-fit-params --model <gguf> | tee args.txt` followed by `cat args.txt | xargs ./build/bin/llama-server --model <gguf>`. [source]
- In the README run (Qwen3-30B-A3B F16, RTX 4090 24077 MiB free) fit projected 61807 MiB, reduced context from 40960 to 4096 (saving 3456 MiB) and printed `-c 4096 -ngl 48` plus one `-ot blk\.N\.ffn_(up|down|gate)_(ch|)exps=CPU` rule per overflowing layer (layers 14 to 47). [source]
- Fitting took 1.15 s in the tool run; the later llama-server run with the stored flags logged "will leave 1199 >= 1024 MiB of free device memory, no changes needed" and fitting took 0.28 s. [source]
- The tool's `main.cpp` entered the tree through PR 23459 (2026-05-21, "app : add batched-bench, fit-params, quantize & perplexity"). [source]
- The latest `fit-params.cpp` commit (2fb989b, 2026-08-22, PR 27496) takes `n_streams` into account and accepts an optional second model so a draft or MTP context is measured at the largest target context. [source]
- PR 27496 makes `--fit -no-kvu -np 4` choose `n_ctx = n_ctx_train * 4`, so each slot receives the full trained context. [source]
- Inferred: run `llama-fit-params` with the same `-fa`, `-ctk/-ctv`, `-np`, `--kv-unified` and `-fitt` the server will use, because each changes projected memory; the printed flags are only valid for that combination. [source]
- Inferred: in a launch script, store the printed line in a file next to the model, re-generate it when macOS memory settings (`iogpu.wired_limit_mb`) or KV flags change, and pass it through `xargs -0` or a bash array to avoid backslash mangling of the `-ot` regexes. [source]
Children
- No children recorded.