<!-- llms-explorer concept facts · https://llms-explorer.com/tree/llama-macos-menu-bar-app/ · pack 2026-10-05 · ~3774 tokens -->

# Llama-macOS menu-bar app

> The app runs `llama serve` in router mode (since 0.22.0, 15 Jan 2026). The server stays up in the background and loads a model when a request names it; no manual model selection is needed.

Parent: [Mac local LLMs: Runtime selection and frontends](https://llms-explorer.com/tree/mac-local-llms-runtime-selection-and-frontends/) · 1 facets · 74 facts · page: https://llms-explorer.com/tree/llama-macos-menu-bar-app/

## Facts

- The app runs `llama serve` in router mode (since 0.22.0, 15 Jan 2026). The server stays up in the background and loads a model when a request names it; no manual model selection is needed. — source: `asserted`
- Idle unload: "Unload when idle" was added in 0.22.0; the default sleep-idle time became 5 minutes in 0.23.0 (19 Jan 2026). llama.app states "unload after 5 minutes idle". Manual Unload lives on each model's page (0.38.0) and as a hover eject icon (0.35.0). — source: `asserted`
- Per-model config is generated into `models.ini` (moved to Application Support in 0.30.0; regenerated every launch). User overrides go in `~/.config/llama/models.user.ini` (added 0.41.0): section header `[org/repo:QUANT]`, keys are `llama serve` options without leading dashes; user values beat app values, including memory-derived ones; an unknown section is passed through, so it can point at local weights or a draft model; an unappliable key makes the app ignore the file and name the faulty option in the menu. A commented template is written on first launch. — source: `asserted`
- Hidden defaults: `defaults write app.llama.Llama extraServerArgs -string "--api-key secret"` appends CLI args after the app's flags (takes effect on next server start); `defaults write app.llama.Llama exposeToNetwork -string "<ip>"` pins a bind address and falls back to localhost if the address is on no interface. — source: `asserted`
- Auto-configuration (llama.app): picks highest precision quant that fits memory, only context sizes that fit, image input on for vision models, speculative decoding on where supported, larger batch on Macs with 32 GB or more. Since 0.38.0 the memory budget is measured from what macOS reports, not estimated from total RAM, and models offer every trained context length (not capped at 128k). The context picker shows per-tier memory cost (0.37.0). — source: `asserted`
- Speculative decoding: prompt-based speculation on by default since 0.30.0; native MTP heads (e.g. Qwen3.6) used automatically since 0.33.0, with an "MTP" chip on installed models (0.42.0). The Settings "Command" tab shows the exact server command (0.33.0). — source: `asserted`
- Engine versions track llama.cpp builds: b8648 (0.28.0), b9835 (0.33.0), b10217 (0.39.0), b10679 (0.42.0), b11200 (0.43.0). The footer/Settings show the in-use build; updates apply on next server start. — source: `asserted`
- Network exposure (0.42.0): setting has Off / Tailscale / This network. Tailscale binds the Tailscale address (offered only if installed and signed in); This network binds 0.0.0.0 with no password. Network access turns off automatically when the Mac changes network or Tailscale signs out. A QR code of the server address is shown inline (0.43.0). — source: `asserted`
- Agent mode: a Settings option passes `--agent` to the server (0.37.0). README warns never to combine agent mode with open network access on an untrusted network. — source: `asserted`
- 0.43.0 lets Settings → Web UI point at a user folder that Llama serves in place of the built-in chat. 0.40.0 added a Quick prompt panel with a user-set global shortcut. 0.41.0 added a "Build an API request" page (OpenAI chat, OpenAI Responses or Anthropic formats; thinking off, JSON schema, image, tool; copy as curl/Python/JS). — source: `asserted`
- Base URL is shown without `/v1` since 0.43.0 so more clients work; localhost keeps working when bound to Tailscale or another address. A green dot beside the menu title shows the server is serving. — source: `asserted`
- Model ids: since 0.35.0 every API model id is `{org}/{repo}:{TAG}`; `qwen3-0.6b:Q8_0` became `ggml-org/Qwen3-0.6B-GGUF:Q8_0` (breaking). Deep links: `llama://` (plus legacy `llamabarn://`) install a model from Hugging Face. — source: `asserted`
- Install size: 1 MB .dmg, 4 MB installed (0.31.0 cut the dmg from 7 MB to 1.6 MB when the engine was unbundled). — source: `asserted`
- Downloads: pause/resume/cancel, survive app restarts, auto-resume when connectivity returns, show transfer rate; an optional Hugging Face token setting exists (0.26.0); custom models-folder setting (0.25.0). — source: `asserted`
- API surface (llama.app docs, describing `llama serve`): OpenAI `/v1/chat/completions`, `/v1/completions`, `/v1/responses`, `/v1/embeddings`, `/v1/rerank`, `/v1/models`; Anthropic `POST /v1/messages`; structured output via `response_format`; tool calling; vision and audio content parts; reasoning returned in `message.reasoning_content`; a `timings` object including `cache_n`. — source: `asserted`
- Aug 2025 repo created as LlamaBarn. Oct-Nov 2025: curated catalog only, ~12 MB app, port 2276, engine bundled, web UI at localhost:2276; no embedding/completion models, no concurrent models, no parallel requests per its roadmap at that time. — source: `asserted`
- 0.14.0 (15 Dec 2025) prefers full-precision models and falls back to quantized only when they will not fit; 0.16.0 removed the memory-limit option for an automatic safety budget. — source: `asserted`
- 0.21.0 removed `--no-mmap` so memory mapping is on. — source: `asserted`
- 0.22.0 router mode; 0.24.0 removed click-to-load because models load on demand. — source: `asserted`
- 0.27.0 (1 Apr 2026) new downloads go to `~/.cache/huggingface/hub/` (standard HF layout); older `~/.llamabarn/` models still work. — source: `asserted`
- 0.29.0 (15 Apr 2026) sideloading: any GGUF in the HF cache appears in the installed list. 0.29.0 shipped unnotarized by a pipeline bug; 0.29.1 fixed notarization. — source: `asserted`
- 0.31.0 (11 Jun 2026) engine unbundled: the app uses a llama.cpp from the llama.app install script or Homebrew if present; if the app installs it, it is also on the terminal PATH. Built-in fixed catalog replaced by "Recommended for your Mac" plus a catalog on llama.app. — source: `asserted`
- 0.32.0 (17 Jun 2026) renamed to Llama; not auto-updating because of the rename (manual download); default port moved 2276 to 8080 to match llama.cpp. — source: `asserted`
- 0.40.0 (11 Aug 2026) default port moved 8080 to 9931 "following llama.cpp, which is preparing to switch its own default too". Port set manually is untouched. — source: `asserted`
- 0.39.0 fixed a bug where Gemma 4, Qwen 3.6 and Devstral 2 vision models installed text-only. — source: `asserted`
- A Windows 11 WinUI 3 port, ggml-org/Llama-Windows, exists (v0.11.0 download linked from llama.app; 0.12.0 bump in source). Linux gets only the CLI. — source: `asserted`
- Scripts pointed at 2276 or 8080 break across the 0.32.0 and 0.40.0 port changes; old model ids break at 0.35.0. — source: `asserted`
- "Orphaned `llama serve` processes no longer block the next launch" and "a clear error when another app holds the server port" were fixed in 0.33.0, so earlier builds could die on a stale server or port clash. — source: `asserted`
- Docs on llama.app for plain `llama serve` still say port 8080 (docs/api, docs/serve), while the app uses 9931. Do not assume the two match. — source: `asserted`
- The server has no password by default; This network option is unauthenticated. Use `extraServerArgs --api-key` or Tailscale. — source: `asserted`
- Open issues (4 at snapshot): custom llama-server args UI (#115), easy integration with OpenClaw/OpenCode/Codex (#55), WebUI link loads a blank page (#78), model name missing from request (#46, seen on 0.23.0). Gap: the app has no GUI for arbitrary server flags other than `models.user.ini` and `extraServerArgs`. — source: `asserted`
- Speech-to-text models were being listed under Installed until 0.39.0; a speculative draft head could displace the real model. — source: `asserted`
- Third-party coverage (Substack Nov 2025, SourcePulse) describes the curated-catalog, port-2276, bundled-engine, ~12 MB design. The current README/llama.app describe arbitrary HF GGUF install, port 9931, shared engine, 4 MB. Treat the older posts as superseded. — source: `asserted`
- HN debate (Apr 2026): one side says llama.cpp now has UX parity with Ollama (GUI, `-hf`, router); the other says Ollama's CLI and docs are still better and llama.cpp "moves fast and breaks things" (new architectures such as gemma4 fail on stale builds). Both stand; this app is the project's answer to the UX side. — source: `asserted`
- Minimum macOS version and Intel support are not stated on the pages read. — source: `asserted`
- Whether llama.cpp's own default port actually switched to 9931 was not confirmed. — source: `asserted`
- Whether the app can run more than one model at once (router `--models-max`) is not stated on app pages. — source: `asserted`
- Llama-macOS was LlamaBarn until release 0.32.0 on 17 Jun 2026; settings and models carried over. — [source](https://github.com/ggml-org/Llama-macOS/releases?page=2)
- 0.32.0 was not delivered by auto-update because of the rename. — [source](https://github.com/ggml-org/Llama-macOS/releases?page=2)
- LlamaBarn's default port was 2276, changed to 8080 in 0.32.0, then to 9931 in 0.40.0 (11 Aug 2026). — [source](https://github.com/ggml-org/Llama-macOS/releases)
- The 0.40.0 notes say the 9931 change follows llama.cpp, which was preparing to change its own default. — [source](https://github.com/ggml-org/Llama-macOS/releases)
- Docs for plain `llama serve` on llama.app still list `127.0.0.1:8080` as the default. — [source](https://llama.app/docs/serve)
- Router mode (`llama serve` with no model) loads/unloads models on demand; sources are the cache, a models dir, or `--models-preset` ini. — [source](https://llama.app/docs/serve)
- The app migrated to llama-server Router Mode in 0.22.0 (15 Jan 2026) and added "Unload when idle". — [source](https://github.com/ggml-org/Llama-macOS/releases?page=3)
- The default sleep-idle time was set to 5 minutes in 0.23.0 (19 Jan 2026). — [source](https://github.com/ggml-org/Llama-macOS/releases?page=3)
- llama.app states models unload after 5 minutes idle. — [source](https://llama.app)
- Since 0.31.0 the app no longer bundles llama.cpp; it reuses an install from the llama.app script or Homebrew and shows the build in the footer. — [source](https://github.com/ggml-org/Llama-macOS/releases?page=2)
- If the app installs llama.cpp, the same install is available in the terminal. — [source](https://github.com/ggml-org/Llama-macOS/releases?page=2)
- The .dmg is 1 MB (1.6 MB at 0.31.0, 7 MB before); installed app is 4 MB. — [source](https://llama.app)
- Engine build in 0.43.0 is b11200; in 0.28.0 (Gemma 4) b8648. — [source](https://github.com/ggml-org/Llama-macOS/releases)
- llama.app's screenshot shows build b9726 and models Qwen3.8 27B (19.0 GB), gpt-oss 20B (12.1 GB), gemma-4 E4B (4.59 GB). — [source](https://llama.app)
- Models are stored once in `~/.cache/huggingface/hub/` from 0.27.0 (1 Apr 2026); older `~/.llamabarn/` models keep working. — [source](https://github.com/ggml-org/Llama-macOS/releases?page=2)
- Any GGUF found in the HF cache (including subdirectories and split shards) is listed as installed since 0.29.0/0.30.0. — [source](https://github.com/ggml-org/Llama-macOS/releases?page=2)
- `models.ini` is regenerated every launch; user overrides belong in `~/.config/llama/models.user.ini`. — [source](https://github.com/ggml-org/Llama-macOS)
- Override keys are `llama serve` options without leading dashes, under a `[org/repo:QUANT]` header, e.g. `ctx-size = 32768`, `cache-type-k = q4_1`. — [source](https://github.com/ggml-org/Llama-macOS)
- Unknown sections in `models.user.ini` pass through, enabling local `model =` paths and `spec-draft-model`. — [source](https://github.com/ggml-org/Llama-macOS)
- `defaults write app.llama.Llama extraServerArgs -string "--api-key secret"` appends server args; `defaults write app.llama.Llama exposeToNetwork -string "<ip>"` binds a specific address. — [source](https://github.com/ggml-org/Llama-macOS)
- Network access options are Off, Tailscale and This network; This network binds 0.0.0.0 with no password. — [source](https://github.com/ggml-org/Llama-macOS)
- Network access switches off when the Mac changes network or Tailscale signs out (0.42.0). — [source](https://github.com/ggml-org/Llama-macOS/releases)
- API model ids are `{org}/{repo}:{TAG}` for all orgs since 0.35.0 (breaking change from short ggml-org ids). — [source](https://github.com/ggml-org/Llama-macOS/releases)
- The server API includes `/v1/responses`, `/v1/embeddings`, `/v1/rerank` and Anthropic `/v1/messages` in addition to chat completions. — [source](https://llama.app/docs/api)
- llama.app's Llama page lists OpenAI- and Anthropic-compatible endpoints, streaming, tool calling, structured output and vision. — [source](https://llama.app)
- The `llama` command (`llama cli`, `llama serve`) is the current CLI name on llama.app docs; installed by `curl -LsSf https://llama.app/install.sh | sh`. — [source](https://llama.app/docs/installation)
- `llama cli`/`llama serve` default to `--fit on`, adjusting unset options to fit device memory. — [source](https://llama.app/docs/cli)
- Native MTP speculative decoding is automatic in the app from 0.33.0 (e.g. Qwen3.6). — [source](https://github.com/ggml-org/Llama-macOS/releases?page=2)
- The memory budget is measured from what macOS reports since 0.38.0, and every trained context length is offered. — [source](https://github.com/ggml-org/Llama-macOS/releases)
- 0.43.0 lets a user folder replace the built-in chat web UI; 0.40.0 added a global-shortcut Quick prompt panel. — [source](https://github.com/ggml-org/Llama-macOS/releases)
- A Settings option enables agent mode by passing `--agent` to the server (0.37.0). — [source](https://github.com/ggml-org/Llama-macOS/releases)
- An open issue (#55) asks for easy integration with OpenClaw, OpenCode and Codex. — [source](https://github.com/ggml-org/Llama-macOS/issues)
- ggml-org/Llama-Windows is a WinUI 3 tray-app port of the Mac app. — [source](https://github.com/ggml-org/Llama-Windows)
- In Nov 2025 LlamaBarn used a curated catalog and port 2276 and could not load arbitrary downloaded models. — [source](https://remotebrowser.substack.com/p/llamabarn-no-frills-local-llms-for)
- Per the project roadmap quoted by SourcePulse, LlamaBarn then lacked embedding models, completion models, concurrent models and parallel requests. — [source](https://www.sourcepulse.org/projects/17342715)
- HN commenters (Apr 2026) report a stale llama.cpp build fails on new architectures ("unknown model architecture: 'gemma4'"). — [source](https://news.ycombinator.com/item?id=47789569)
- Comparison with Ollama (inferred): Llama keeps weights in the standard HF cache and uses stock llama.cpp, whereas Ollama uses its own blob store and bundles its own runners. — source: `asserted`
