<!-- llms-explorer concept facts · https://llms-explorer.com/tree/unsloth-desktop-studio-macos-app-running-gguf-an/ · pack 2026-10-05 · ~3031 tokens -->

# Unsloth Desktop Studio macOS app running GGUF and MLX

> GGUF inference runs through llama.cpp; the loaded model is exposed as an authenticated API served by `llama-server`. MLX models run through an MLX backend on Apple Silicon (experimental since 2026-05-18, expanded since).

Parent: [Mac local LLMs: Runtime selection and frontends](https://llms-explorer.com/tree/mac-local-llms-runtime-selection-and-frontends/) · 1 facets · 62 facts · page: https://llms-explorer.com/tree/unsloth-desktop-studio-macos-app-running-gguf-an/

## Facts

- GGUF inference runs through llama.cpp; the loaded model is exposed as an authenticated API served by `llama-server`. MLX models run through an MLX backend on Apple Silicon (experimental since 2026-05-18, expanded since). — source: `asserted`
- API surface: Anthropic-compatible `/v1/messages` (for Claude Code, OpenClaw and the Anthropic SDK) and OpenAI-compatible `/v1/chat/completions` and `/v1/responses`. Keys are created in Settings, API, look like `sk-unsloth-...`, are shown once, and a revoked key returns 401. — source: `asserted`
- `unsloth run --model <repo>:<quant>` loads a model with sampling, context, GPU-layer, threading, network and tool options, and prints an endpoint URL and API key. `unsloth studio -H 0.0.0.0 -p 8888` starts the web UI; the default bind is 127.0.0.1. `unsloth studio --secure` opens a free Cloudflare HTTPS tunnel. — source: `asserted`
- Server-side tools (web search, Python and terminal execution) run as the local user and are on by default; anyone holding the API key can run code on the Mac, so `--disable-tools` is the guard when exposing the server. A preview OS-level tool sandbox exists on supported Linux and macOS hosts, with network access still unrestricted. — source: `asserted`
- Model swapping: API requests can opt in to automatic switching between downloaded local GGUFs (2026-07-07); unknown model names keep using the current model. This makes Studio act like a llama-swap-style router. — source: `asserted`
- Update mechanics: re-running the install command updates Studio; Desktop updates in-app. — source: `asserted`
- 2026-03-17: Studio beta launch (Mac behaved like CPU: chat first; the install page later says Data Recipes work and "MLX training now works"). — source: `asserted`
- 2026-03-25: app shortcuts on macOS; Data Recipes on macOS and CPU. — source: `asserted`
- 2026-03-27 and 2026-04-01: auto-detect existing HF and LM Studio models, then user-selected folder. — source: `asserted`
- 2026-04-03: Intel Macs work. — source: `asserted`
- 2026-05-18: experimental MLX inference; API connections to OpenAI, Anthropic, vLLM, Ollama; MTP speculative decoding. — source: `asserted`
- 2026-05-19: "improve MTP not being faster on Macs, CPUs and GPUs". — source: `asserted`
- 2026-05-31: prebuilt llama.cpp binaries re-enabled for Apple Silicon M1-M4 on macOS 14, 15 and 26 (Tahoe); Apple Silicon on macOS 13 builds from source; Intel Macs on 13.3 to 26 use prebuilts. — source: `asserted`
- 2026-06-12: Studio uses fresh llama.cpp prebuilts across CUDA, ROCm, Windows, Linux and macOS; improved MLX model labels and generation speed stats. — source: `asserted`
- 2026-06-18: MLX training updates. — source: `asserted`
- 2026-07-07: macOS installs no longer need CMake or Homebrew when a prebuilt llama.cpp exists; Apple Silicon handling of paths with spaces and unified memory sizing improves. — source: `asserted`
- 2026-08-11: Unsloth Desktop (Beta) announced for Mac, Windows, Linux. — source: `asserted`
- 2026-08-20: custom llama.cpp build support; toggles for Cache RAM, Mmap, Mlock, Checkpoints, speculative-decoding KV cache and vision on/off. — source: `asserted`
- 2026-09-02: OpenAI-compatible APIs for video, audio and MLX-served models. — source: `asserted`
- 2026-09-17: MLX gains video input and optional MoE and decode optimizations. — source: `asserted`
- 2026-09-28: Apple Silicon improvements: batched serving, structured outputs and a TurboQuant KV cache. — source: `asserted`
- 2026-10-01 (release v0.1.902-beta): updated MLX support; Mac LoRA exports to PEFT or GGUF. — source: `asserted`
- macOS floor: manual Studio install requires macOS 12 Monterey or newer and a Python environment (uv, venv or conda); prebuilt llama.cpp covers macOS 14, 15, 26 on Apple Silicon, so macOS 13 compiles llama.cpp from source. — source: `asserted`
- The doc page for `unsloth run` shows the model argument `unsloth/qwen3.8-27B-GGUF-GGUF:UD-Q4_K_XL` with a doubled `-GGUF` suffix, which looks like a typo and would not resolve as written. — source: `asserted`
- Mac training depth is contested: the install page says Studio works on Mac for chat and "training and all features" with MLX training working, while a third-party review says training is the weaker path on Mac and NVIDIA is listed for training. — source: `asserted`
- Studio's Mac chat path historically lacked MLX tools and web search: the 2026-05-18 notes say "thinking, tools and web search soon" for MLX, so older MLX chats may not call tools. — source: `asserted`
- Desktop uninstall (drag to Trash) leaves the Studio data and HF model cache; manual removal uses `rm -rf ~/.unsloth/studio`, the launcher `~/Applications/Unsloth Studio.app`, and `~/.local/bin/unsloth`. — source: `asserted`
- `UNSLOTH_STUDIO_HOME` isolates a second install (own venv, auth, `studio.db`, cache and llama.cpp build). — source: `asserted`
- Bundled llama.cpp can be stale: Studio warns when the prebuilt is too old for MTP. — source: `asserted`
- What "Desktop" is relative to "Studio": Unsloth's docs call Desktop the easiest way to install Studio and a native app that Studio is packaged into; a third-party review calls Desktop a superset of Studio and says its CLI is `unsloth start`, while the Unsloth pages show `unsloth run` and `unsloth studio`. The docs are the primary source. — source: `asserted`
- Licensing and telemetry are stated by a third-party review (Apache 2.0 core, AGPL-3.0 for Studio components, no telemetry) and not on the pages read from unsloth.ai; treat as unconfirmed. — source: `asserted`
- Star count: review says about 70.2k, the GitHub releases page at fetch shows about 77.2k; the review is dated August. — source: `asserted`
- No source compares Desktop's MLX decode and prefill speed to mlx-lm, LM Studio or oMLX on the same model and chip. — source: `asserted`
- Which MLX build and mlx-lm version Desktop bundles, and how it pins them, is not stated. — source: `asserted`
- How Desktop's "TurboQuant KV cache" differs from llama.cpp's q8_0 or q4_0 cache-type options is not documented in the changelog entry. — source: `asserted`
- Unsloth Desktop (Beta) is a free, open-source app for running and training models on macOS, Windows and Linux, built on Tauri. — [source](https://unsloth.ai/docs/desktop)
- The Mac install is a `.dmg`: open it, drag Unsloth to Applications, launch and wait for installation to complete, then pick a model and quantization from the Select model dropdown or Model hub tab. — [source](https://unsloth.ai/docs/get-started/install/mac)
- `curl -fsSL https://unsloth.ai/install.sh | sh` installs Unsloth Studio (the browser UI), not Desktop or Unsloth Core, and running it again updates. — [source](https://unsloth.ai/docs/get-started/install/mac)
- `unsloth studio -H 0.0.0.0 -p 8888` launches Studio, and by default it binds to 127.0.0.1 only. — [source](https://unsloth.ai/docs/new/studio/install)
- `unsloth studio --secure` launches over HTTPS through a free Cloudflare tunnel on Windows, Mac and Linux. — [source](https://unsloth.ai/docs/new/studio)
- Studio's macOS requirement is macOS 12 Monterey or newer on Intel or Apple Silicon, inside a Python environment such as uv, venv or conda. — [source](https://unsloth.ai/docs/new/studio/install)
- The install page says Mac chat and Data Recipes work and "MLX training now works", and the intro says Training, MLX and GGUF inference all work inside Unsloth on macOS. — [source](https://unsloth.ai/docs/new/studio/install)
- Models loaded in Unsloth, including GGUFs, are exposed as an authenticated API through `llama-server`, with Anthropic-compatible `/v1/messages` and OpenAI-compatible `/v1/chat/completions` and `/v1/responses`. — [source](https://unsloth.ai/docs/basics/api)
- API keys start with `sk-unsloth-`, are created under Settings, API, are shown only once, and revoked keys return 401. — [source](https://unsloth.ai/docs/basics/api)
- `unsloth run` serves a model with context size, GPU layers, threading, sampling, networking and tool settings. — [source](https://unsloth.ai/docs/desktop)
- Server-side web search and Python and terminal execution run as the local user and are on by default; the install page says to pass `--disable-tools` when exposing Unsloth. — [source](https://unsloth.ai/docs/new/studio/install)
- Unsloth Desktop lets users pick sandboxed tool execution or direct file access, and asks approval for file access outside the sandbox. — [source](https://unsloth.ai/docs/desktop)
- Studio can connect to OpenAI, Anthropic, Ollama, llama.cpp and vLLM backends and supports parallel chatting. — [source](https://unsloth.ai/docs/desktop)
- 2026-05-18 changelog: experimental MLX inference lets MLX quants and models run on Macs, with thinking, tools and web search "coming soon". — [source](https://unsloth.ai/docs/new/changelog)
- 2026-05-31 changelog: prebuilt llama.cpp binaries re-enabled for Apple Silicon (M1-M4) on macOS 14, 15 and 26; Apple Silicon on macOS 13 is a source build; Intel Macs use prebuilts. — [source](https://unsloth.ai/docs/new/changelog)
- 2026-07-07 changelog: macOS installs no longer require CMake or Homebrew when a prebuilt llama.cpp is available, and Apple Silicon support handles paths with spaces and unified memory sizing better. — [source](https://unsloth.ai/docs/new/changelog)
- 2026-07-07 changelog: API requests can opt into automatic switching between downloaded local GGUFs and `/v1/models` returns clean model IDs. — [source](https://unsloth.ai/docs/new/changelog)
- 2026-08-20 changelog: custom llama.cpp builds and toggles for Cache RAM, Mmap, Mlock, Checkpoints, speculative-decoding KV cache and vision on/off. — [source](https://unsloth.ai/docs/new/changelog)
- 2026-09-02 changelog: new OpenAI-compatible APIs for video, audio and MLX-served models. — [source](https://unsloth.ai/docs/new/changelog)
- 2026-09-17 changelog: MLX gains video input, optional MoE and decode optimizations and more reliable multimodal chats; preview OS-level tool sandboxing arrives on supported Linux and macOS hosts. — [source](https://unsloth.ai/docs/new/changelog)
- 2026-09-28 changelog: Apple Silicon improvements include batched serving, structured outputs and a TurboQuant KV cache. — [source](https://unsloth.ai/docs/new/changelog)
- 2026-10-01 changelog (v0.1.902-beta): updated MLX support and Mac LoRA exports to PEFT or GGUF. — [source](https://unsloth.ai/docs/new/changelog)
- The 2026-08-11 changelog announces Unsloth Desktop for Mac, Windows and Linux supporting MLX, diffusion, audio and GGUF. — [source](https://unsloth.ai/docs/new/changelog)
- A Desktop uninstall is Finder, Applications, Move to Trash; manual Studio removal is `rm -rf ~/.unsloth/studio`, `rm -rf ~/Applications/Unsloth\ Studio.app` and `rm -f ~/.local/bin/unsloth`, and none of these touch downloaded HF model files. — [source](https://unsloth.ai/docs/get-started/install/mac)
- Setting `UNSLOTH_STUDIO_HOME` installs Studio into an isolated location with its own virtual env, auth, `studio.db`, cache and llama.cpp build. — [source](https://unsloth.ai/docs/new/studio/install)
- A third-party review states Unsloth Desktop ships Apache 2.0 core with AGPL-3.0 Studio components, no telemetry, and an `unsloth start` agent bridge, and calls Mac training "check before you plan around it". — [source](https://explainx.ai/blog/unsloth-desktop-local-ai-train-run-models-august-2026)
- On a Mac, Desktop's GGUF path is llama.cpp with Metal and its MLX path is a separate backend, so the same model can be run both ways in one app and compared for speed and fidelity locally. — source: `asserted`
