Ollama default 262144 context and num_ctx memory cost on Mac
Parent: Mac local LLMs: Ollama internals · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Ollama docs: "Tasks which require large context like web search, agents, and coding tools should be set to at least 64000 tokens."
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Ollama docs: "Tasks which require large context like web search, agents, and coding tools should be set to at least 64000 tokens." [source]
- Ollama docs: cloud models are set to their maximum context length by default. [source]
- Ollama docs example: `ollama ps` columns are NAME, ID, SIZE, PROCESSOR, CONTEXT, UNTIL, e.g. gemma4:latest 9.6 GB, 100% GPU, context 131072. [source]
- Ollama docs advise using the model's maximum context for best performance while keeping PROCESSOR at 100% GPU (no CPU offload). [source]
- Per-token KV cost differs by model: a maintainer measured 32k context at 6.2 GB for phi4:14b and 12 GB for gemma3:12b (gemma3 also needs memory for its image projector). [source]
- A maintainer showed OLLAMA_NUM_PARALLEL=3 turned a 128000 request into --ctx-size 384000, needing 181.6 GiB KV and leaving only 12 of 63 layers on GPU. [source]
- Changing num_ctx causes a model reload, which is why the first request after a change is slow. [source]
- If a model is sharded across several devices the VRAM needed rises because each device runs a copy of the compute graph. [source]
Children
- No children recorded.