Ollama launcher flag set passed to llama-server
Parent: Mac local LLMs: Ollama internals · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Always: `--log-verbosity 4 --no-log-prefix --no-log-timestamps`, so startup memory and offload lines stay visible for Ollama's scheduler accounting.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Always: `--log-verbosity 4 --no-log-prefix --no-log-timestamps`, so startup memory and offload lines stay visible for Ollama's scheduler accounting. [source]
- Batch: for chat models `-b N -ub N` only when `NumBatch > 0`; for embedding models `--embedding` plus `-b`/`-ub` from an embedding batch size capped by `NumCtx * numParallel`. [source]
- Multimodal: `--mmproj <file>` for the first projector; `--no-mmproj-offload` when the projector is not worth putting on the GPU. [source]
- Context: `--context-shift` (plus `--keep N` when `NumKeep > 0`) only when the model config enables context shift. [source]
- GPU choice: `--split-mode none --main-gpu N` only when a main GPU is set. [source]
- Threads: `-t N` only when set. [source]
- Qwen vision families: `--image-min-tokens 1024`. [source]
- Speculation: `--spec-type`, `--spec-draft-n-max`, `--spec-draft-backend-sampling` and `--spec-draft-model` (see the MTP dossier). [source]
- Environment: `LLAMA_MEDIA_MARKER` set to a per-process random `<__ollama_media_...__>` string, plus device-selection variables. [source]
- 2026-05-29: Ollama removes its own GGML engines and uses this launcher for all GGUF models; the flag list has grown by helper functions since (mmproj, draft, load mode, flash attention). [source]
- The README for llama bumps tells maintainers to re-check "launch args and defaults" and log lines parsed from llama-server on every bump. [source]
- `--no-mmproj-offload` is chosen for reasons `cpu-only`, `partial-text-offload`, `limited-vram` and `startup-oom-retry`. [source]
- Gemma3n's projector is the exception: it silently produces corrupted image embeddings on the CPU backend, so offload is kept whenever a GPU is in play. [source]
- The compat clip allowlist (`gemma3`, `gemma4`, `qwen35`, `qwen35moe`, `qwen25vl`, `qwen3vl`, `qwen3vlmoe`, `mistral3`, `deepseekocr`, `glmocr`, `llama4`, `nemotron_h_omni`) decides whether a model file with inline vision tensors is passed to `--mmproj` as itself; other architectures would make the clip loader abort on untranslated Ollama tensor names. [source]
- Nothing in the list sets `--ctx-checkpoints`, `--cache-ram`, `--lazy-mode`, `--threads-batch` or `--parallel` other than `-np`; llama-server defaults apply. [source]
- Whether `OLLAMA_NUM_PARALLEL` plus context scaling (`-c NumCtx*numParallel`) interacts with llama-server's unified KV default as expected on Metal. [source]
- Which flags change when `ollama launch` integrations start a server (not read). [source]
- The launcher always appends `--log-verbosity 4`, `--no-log-prefix` and `--no-log-timestamps`, with the comment "Keep startup memory/offload lines visible for scheduler accounting". [source]
- For non-embedding models `-b` and `-ub` are both set to `NumBatch` only when it is above zero. [source]
- For embedding models the launcher adds `--embedding` and sets `-b` and `-ub` to an embedding batch size that is capped at `NumCtx * max(numParallel, 1)`. [source]
- `--mmproj` receives the first projector path, and `--no-mmproj-offload` is added when `mmprojOffloadDisabled` returns true. [source]
- Projector offload is disabled for `cpu-only` (`NumGPU == 0`), `partial-text-offload` (`0 < NumGPU < model layers`), `limited-vram` (free or total memory below projector size plus a headroom constant) and `startup-oom-retry`. [source]
- The gemma3n architecture is exempt from projector offload disabling because its MobileNetV5 projector "silently produces corrupted image embeddings" on the CPU backend. [source]
- `--context-shift` is passed only when the model config enables it, with `--keep N` when `NumKeep` is above zero. [source]
- `--split-mode none --main-gpu N` is passed only when `MainGPU` is set, and `-t N` only when `NumThread` is above zero. [source]
- For architectures `qwen2vl`, `qwen25vl`, `qwen3vl` and `qwen3vlmoe` the launcher adds `--image-min-tokens 1024` because upstream mtmd warns that Qwen-VL needs at least 1024 image tokens for correct grounding and counting. [source]
- The launcher sets `LLAMA_MEDIA_MARKER` in the subprocess environment to a random `<__ollama_media_...__>` string. [source]
- The compat clip allowlist maps the architectures gemma3, gemma4, qwen35, qwen35moe, qwen25vl, qwen3vl, qwen3vlmoe, mistral3, deepseekocr, glmocr, llama4 and nemotron_h_omni to "pass the model file as its own mmproj". [source]
- Flash attention resolves to `on` or `off` when `OLLAMA_FLASH_ATTENTION` is set and to `auto` when unset and supported, and to `off` when the devices do not support it. [source]
- The stall timeout for model load comes from `envconfig.LoadTimeout()`. [source]
- The `llama/README.md` bump checklist names "launch args and defaults" and "scheduler-sensitive flags consumed by llm/llama_server.go or server/sched.go". [source]
- Observed on this Mac (Ollama 0.34.4, macOS arm64) for a running embedding model: `llama-server --model <blob> --port 57772 --host 127.0.0.1 --no-webui --offline -c 2048 -np 1 --log-verbosity 4 --no-log-prefix --no-log-timestamps --flash-attn auto --embedding -b 2048 -ub 2048 --context-shift --keep 4`. [source]
- The same observed command line has no `-ngl`, `--load-mode` or `--cache-type-*` flag, matching the rule that these are added only on explicit options. [source]
- Consequence: Ollama users cannot reach llama-server flags outside this list (checkpoints, cache RAM, lazy mode, threads for batch) except through the environment variables that llama-server itself reads, if they reach the subprocess. [source]
Children
- No children recorded.