<!-- llms-explorer concept facts · https://llms-explorer.com/tree/ollama-default-262144-context-and-num-ctx-memory/ · pack 2026-10-05 · ~472 tokens -->

# Ollama default 262144 context and num_ctx memory cost on Mac

> Ollama docs: "Tasks which require large context like web search, agents, and coding tools should be set to at least 64000 tokens."

Parent: [Mac local LLMs: Ollama internals](https://llms-explorer.com/tree/mac-local-llms-ollama-internals/) · 1 facets · 8 facts · page: https://llms-explorer.com/tree/ollama-default-262144-context-and-num-ctx-memory/

## Facts

- Ollama docs: "Tasks which require large context like web search, agents, and coding tools should be set to at least 64000 tokens." — [source](https://docs.ollama.com/context-length)
- Ollama docs: cloud models are set to their maximum context length by default. — [source](https://docs.ollama.com/context-length)
- Ollama docs example: `ollama ps` columns are NAME, ID, SIZE, PROCESSOR, CONTEXT, UNTIL, e.g. gemma4:latest 9.6 GB, 100% GPU, context 131072. — [source](https://docs.ollama.com/context-length)
- Ollama docs advise using the model's maximum context for best performance while keeping PROCESSOR at 100% GPU (no CPU offload). — [source](https://docs.ollama.com/context-length)
- Per-token KV cost differs by model: a maintainer measured 32k context at 6.2 GB for phi4:14b and 12 GB for gemma3:12b (gemma3 also needs memory for its image projector). — [source](https://github.com/ollama/ollama/issues/9890)
- A maintainer showed OLLAMA_NUM_PARALLEL=3 turned a 128000 request into --ctx-size 384000, needing 181.6 GiB KV and leaving only 12 of 63 layers on GPU. — [source](https://github.com/ollama/ollama/issues/9890)
- Changing num_ctx causes a model reload, which is why the first request after a change is slow. — [source](https://github.com/ollama/ollama/issues/9890)
- If a model is sharded across several devices the VRAM needed rises because each device runs a copy of the compute graph. — [source](https://github.com/ollama/ollama/issues/9890)
