<!-- llms-explorer concept facts · https://llms-explorer.com/tree/ollama-on-macos/ · pack 2026-10-05 · ~3257 tokens -->

# Ollama on macOS

> App install: mount ollama.dmg, drag to /Applications; on first start the app links the CLI into /usr/local/bin. Brew install runs a launchd service via `brew services`. The curl install script also exists for macOS.

Parent: [Mac local LLMs: Runtime selection and frontends](https://llms-explorer.com/tree/mac-local-llms-runtime-selection-and-frontends/) · 1 facets · 61 facts · page: https://llms-explorer.com/tree/ollama-on-macos/

## Facts

- App install: mount ollama.dmg, drag to /Applications; on first start the app links the CLI into /usr/local/bin. Brew install runs a launchd service via `brew services`. The curl install script also exists for macOS. — source: `asserted`
- The app is the launchd-less "login item": env vars reach it only through `launchctl setenv`, followed by an app restart. — source: `asserted`
- State lives under ~/.ollama (models, logs, server.json, config.json, id_ed25519 keypair). — source: `asserted`
- Runners: this box (0.34.4) spawns `llama-server` (llama.cpp) subprocesses with GGUF blobs and ships `mlx_metal_v3` / `mlx_metal_v4` dirs for the MLX path; llama.cpp Metal init logs `ggml_metal_*`. — source: `asserted`
- 0.19 (2026-03-30): MLX preview on Apple Silicon, NVFP4, cache reuse/checkpoints; the preview asked for a Mac with more than 32 GB unified memory. — source: `asserted`
- Several Ollama installs can coexist (app, /usr/local/bin symlink, brew) and each runs its own server; on this box 4 `ollama serve` processes were seen. Only one can own 11434. — source: `asserted`
- MLX runner leak report (issue 17792) was closed as completed 2026-08-26 via PR 17798 after maintainers could not reproduce on a Mac, so the existing note should say it is resolved/unreproduced. — source: `asserted`
- Docker Desktop on macOS has no GPU passthrough; Ollama in Docker on Mac is CPU only. — source: `asserted`
- Intel Macs are CPU only. — source: `asserted`
- Context default: FAQ says 4096 flat; context-length page says VRAM-tiered 4k/32k/256k. (Already in existing file.) The CLI `ollama serve --help` text in 0.34.4 agrees with the tiered version. — source: `asserted`
- Medium article (2025) says `ollama start` is a distinct background command; 0.34.4 `ollama serve --help` lists `start` only as an alias of `serve` (foreground). — source: `asserted`
- Whether MLX is chosen automatically for GGUF-free safetensors tags only (`-mlx`, `-nvfp4`) or also via a flag: runner flag `--mlx-engine` was seen only in a Linux issue transcript. — source: `asserted`
- No pre-login (LaunchDaemon) officially supported path found; issue 2955 (2024) still describes it as a workaround problem. — source: `asserted`
- Minimum macOS for the Ollama app is Sonoma (14); Apple M-series gets CPU+GPU, Intel x86 is CPU only. — [source](https://docs.ollama.com/macos)
- The documented install path is ollama.dmg drag-to-Applications; on first launch the app prompts to create a link to the CLI in /usr/local/bin if `ollama` is not on PATH. — [source](https://docs.ollama.com/macos)
- The CLI binary lives at Ollama.app/Contents/Resources/ollama; to install the app elsewhere, decline the "Move to Applications?" prompt and put the CLI or a symlink on PATH. — [source](https://docs.ollama.com/macos)
- ollama.com/download/mac now leads with `curl -fsSL https://ollama.com/install.sh | sh` for macOS and linux, with the dmg as "Download manually". — [source](https://ollama.com/download/mac)
- Homebrew formula `ollama` is a CLI/server formula, stable 0.35.1 on 2026-10-04, bottles for Apple Silicon on Sequoia, Tahoe and "golden gate" (macOS 27); there is no Intel macOS bottle listed. — [source](https://formulae.brew.sh/formula/ollama)
- The brew formula depends on the `mlx-c` formula (0.7.0) at runtime. — [source](https://formulae.brew.sh/formula/ollama)
- Brew service commands: `brew services start ollama` (start now and at login) or `$HOMEBREW_PREFIX/opt/ollama/bin/ollama serve` (foreground). — [source](https://formulae.brew.sh/formula/ollama)
- The Homebrew install on this box ships `sh.brew.ollama.plist` and `sh.brew.ollama.service` in the Cellar and keeps its runner at /opt/homebrew/Cellar/ollama/0.34.4/libexec/lib/ollama/llama-server. [src: local /opt/homebrew/Cellar/ollama/0.34.4] — source: `asserted`
- Homebrew and Ollama.app versions can drift: the cask-less brew formula is at 0.35.1 while this box's brew install was 0.34.4 and the CLI warned "client version is 0.35.1" against the 0.34.4 server. [src: local `ollama -v`] — source: `asserted`
- When Ollama runs as the macOS app, environment variables must be set with `launchctl setenv VAR value` and the Ollama app restarted; the FAQ's example is `launchctl setenv OLLAMA_HOST "0.0.0.0:11434"`. — [source](https://docs.ollama.com/faq)
- `launchctl setenv` values do not survive reboot (issue reporter had to repeat it after each reboot, and ran only after login). — [source](https://github.com/ollama/ollama/issues/2955)
- Verified on this box: `launchctl getenv OLLAMA_HOST` returns http://127.0.0.1:11434 (a LaunchAgent sets it for GUI apps). [src: local launchctl] — source: `asserted`
- The macOS app and the server register as a Login Item; disable it in System Settings > Login Items, "Allow in the Background"; the setting persists across upgrades unless the app is uninstalled. — [source](https://docs.ollama.com/faq)
- The macOS app updates itself automatically and applies the update from the menu-bar item "Restart to update"; Linux re-runs the install script. — [source](https://docs.ollama.com/faq)
- Models are stored in ~/.ollama/models on macOS; `OLLAMA_MODELS` relocates them. — [source](https://docs.ollama.com/faq)
- Logs on macOS: ~/.ollama/logs/app.log (GUI) and ~/.ollama/logs/server.log (server), with rotated app-N.log and server-N.log. — [source](https://docs.ollama.com/macos)
- The Ollama public key on macOS is ~/.ollama/id_ed25519.pub; sign-in is in the Mac app Settings or `ollama signin`. — [source](https://docs.ollama.com/faq)
- Local-only mode: `OLLAMA_NO_CLOUD=1` or `{"disable_ollama_cloud": true}` in ~/.ollama/server.json disables cloud models and web search; the log then prints `Ollama cloud disabled: true`. — [source](https://docs.ollama.com/faq)
- Cloud models use a `:cloud` name suffix (e.g. gemma4:cloud, free plan) and the Anthropic-compatible endpoint also works against https://ollama.com with `Authorization: Bearer` (x-api-key alone is rejected). — [source](https://ollama.com/download/mac)
- Full uninstall removes /Applications/Ollama.app, /usr/local/bin/ollama, ~/Library/Application Support/Ollama, ~/Library/Saved Application State/com.electron.ollama.savedState, ~/Library/Caches/com.electron.ollama, ~/Library/Caches/ollama, ~/Library/WebKit/com.electron.ollama and ~/.ollama. — [source](https://docs.ollama.com/macos)
- GPU acceleration on Apple devices is via the Metal API; the GPU doc has nothing Mac-specific beyond that sentence. — [source](https://docs.ollama.com/gpu)
- Docker Desktop on macOS cannot give Ollama GPU acceleration (no GPU passthrough or emulation). — [source](https://docs.ollama.com/faq)
- Keep-alive default is 5 minutes; `OLLAMA_KEEP_ALIVE` accepts a duration string, seconds, any negative number (forever) or 0 (unload immediately); the per-request `keep_alive` on /api/generate and /api/chat overrides the env var; `ollama stop <model>` unloads now. — [source](https://docs.ollama.com/faq)
- Preload with an empty request: `curl localhost:11434/api/generate -d '{"model":"m"}'` or `ollama run m ""`. — [source](https://docs.ollama.com/faq)
- Concurrency defaults: OLLAMA_MAX_LOADED_MODELS = 3 x GPUs (3 for CPU), OLLAMA_NUM_PARALLEL = 1 default, OLLAMA_MAX_QUEUE = 512; overflow returns HTTP 503. — [source](https://docs.ollama.com/faq)
- Parallel requests multiply the allocated context: memory scales with OLLAMA_NUM_PARALLEL x OLLAMA_CONTEXT_LENGTH (2K x 4 parallel = 8K). — [source](https://docs.ollama.com/faq)
- On Apple Silicon the GPU shares unified memory, so a loaded model counts against both, and a new model that does not fit alongside loaded ones makes requests queue until older models unload. — [source](https://docs.ollama.com/faq)
- `ollama ps` PROCESSOR column shows 100% GPU, 100% CPU or a CPU/GPU split; recent versions also show a CONTEXT column with the allocated context length. — [source](https://docs.ollama.com/context-length)
- The macOS app has a context-length slider in Settings; for non-app use `OLLAMA_CONTEXT_LENGTH=64000 ollama serve`. Coding tools and agents should use at least 64000 tokens. — [source](https://docs.ollama.com/context-length)
- `OLLAMA_ORIGINS` must include chrome-extension://*, moz-extension://* and safari-web-extension://* for browser extensions; the default allows only 127.0.0.1 and 0.0.0.0 origins. — [source](https://docs.ollama.com/faq)
- For model pulls behind a proxy set `HTTPS_PROXY`; avoid `HTTP_PROXY`, which can break client connections to the server. — [source](https://docs.ollama.com/faq)
- `ollama serve --help` in 0.34.4 also lists variables absent from the existing file: OLLAMA_LOAD_TIMEOUT (default 5m), OLLAMA_MAX_TRANSFER_STREAMS (default 4, safetensors pulls/pushes), OLLAMA_NOPRUNE, OLLAMA_SCHED_SPREAD, OLLAMA_LLM_LIBRARY, OLLAMA_GPU_OVERHEAD, OLLAMA_IGPU_ENABLE, LLAMA_ARG_FIT (default on) and LLAMA_ARG_FIT_TARGET. [src: local `ollama serve --help` 0.34.4] — source: `asserted`
- In 0.34.4 `start` is an alias of `serve` (foreground server), not a separate background daemon command. [src: local `ollama serve --help` 0.34.4] — source: `asserted`
- Ollama's own docs list `OLLAMA_MAX_LOADED_MODELS` as "per GPU"; the FAQ prose says 3 x GPUs. [src: local `ollama serve --help` 0.34.4] — source: `asserted`
- MLX engine blog (2026-03-30): MLX preview in 0.19 benefits all Apple Silicon, with M5/M5 Pro/M5 Max Neural Accelerators speeding TTFT and decode; Qwen3.5-35B-A3B NVFP4 measured 1810 tok/s prefill and 112 tok/s decode vs 1154 and 58 on 0.18 Q4_K_M. — [source](https://ollama.com/blog/mlx)
- The MLX preview required "more than 32GB of unified memory" and launched with `qwen3.5:35b-a3b-coding-nvfp4` via `ollama launch claude --model ...`. — [source](https://ollama.com/blog/mlx)
- The 0.19 cache upgrade reuses KV cache across conversations, stores checkpoints at intelligent prompt locations and evicts shared prefixes last, which helps Claude Code style shared system prompts. — [source](https://ollama.com/blog/mlx)
- Issue 17792 (MLX runner stays resident after `ollama stop`) was filed 2026-08-15 against 0.32.9/0.32.13 with 7 of 7 reproductions, was not reproducible by two maintainers on Linux or Mac, was handled in PR 17798 and closed as completed on 2026-08-26. — [source](https://github.com/ollama/ollama/issues/17792)
- Maintainer reproduction output shows the MLX runner launched as `ollama runner --mlx-engine --model <tag> --port <n>` and exiting after `ollama stop`. — [source](https://github.com/ollama/ollama/issues/17792)
- On 0.34.4 on an M5 Max, `ollama ps` showed an embedding model at 100% GPU with context 40960 and the log printed `ggml_metal_device_init: recommendedMaxWorkingSetSize = 55662.79 MB`, i.e. Metal exposes about 55.7 GB working set on this machine. [src: local ~/.ollama/logs/server.log] — source: `asserted`
- On 0.34.4 the app bundles `llama-server`, `mlx_metal_v3` and `mlx_metal_v4` plus x86 `libggml-cpu-*.so` variants in Contents/Resources; llama.cpp-backed loads run as separate `llama-server --model <blob>` subprocesses. [src: local /Applications/Ollama.app/Contents/Resources] — source: `asserted`
- The 4k/32k/256k context tier applies per VRAM; on Apple Silicon the VRAM figure is derived from unified memory, so a 32 GB Mac defaults to 32k and a 16 GB Mac to 4k. — source: `asserted`
- Anthropic-compatible local usage needs only `ANTHROPIC_AUTH_TOKEN=ollama` and `ANTHROPIC_BASE_URL=http://localhost:11434` (no /v1 suffix); SDK clients pass `base_url='http://localhost:11434'`, api_key ignored. — [source](https://docs.ollama.com/api/anthropic-compatibility)
- Anthropic compat supports messages, streaming and function calling, but tool-choice controls, deferred tools and hosted web search are not fully supported. — [source](https://docs.ollama.com/api/anthropic-compatibility)
- `ANTHROPIC_*` variables exported in a shell (e.g. ~/.spawnrc) override Claude auth and produce 401s when Ollama endpoints are not intended. — source: `asserted`
- Issue 3581 (2024): app on macOS not reachable on LAN after setting OLLAMA_HOST=0.0.0.0, still reachable on localhost; closed. — [source](https://github.com/ollama/ollama/issues/3581)
- Issue 2955 (2024): no first-party way to run Ollama as a pre-login daemon on macOS; user tried launchd.conf, rc.common, LaunchAgents and LaunchDaemons plists without success. — [source](https://github.com/ollama/ollama/issues/2955)
- Practical workaround for headless/pre-login service on macOS is brew's launchd service (`brew services start ollama`) with a plist env block, which runs per-user at login, not pre-login. — source: `asserted`
