Unsloth Desktop Studio macOS app running GGUF and MLX
Parent: Mac local LLMs: Runtime selection and frontends · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
GGUF inference runs through llama.cpp; the loaded model is exposed as an authenticated API served by `llama-server`. MLX models run through an MLX backend on Apple Silicon (experimental since 2026-05-18, expanded since).
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- GGUF inference runs through llama.cpp; the loaded model is exposed as an authenticated API served by `llama-server`. MLX models run through an MLX backend on Apple Silicon (experimental since 2026-05-18, expanded since). [source]
- API surface: Anthropic-compatible `/v1/messages` (for Claude Code, OpenClaw and the Anthropic SDK) and OpenAI-compatible `/v1/chat/completions` and `/v1/responses`. Keys are created in Settings, API, look like `sk-unsloth-...`, are shown once, and a revoked key returns 401. [source]
- `unsloth run --model <repo>:<quant>` loads a model with sampling, context, GPU-layer, threading, network and tool options, and prints an endpoint URL and API key. `unsloth studio -H 0.0.0.0 -p 8888` starts the web UI; the default bind is 127.0.0.1. `unsloth studio --secure` opens a free Cloudflare HTTPS tunnel. [source]
- Server-side tools (web search, Python and terminal execution) run as the local user and are on by default; anyone holding the API key can run code on the Mac, so `--disable-tools` is the guard when exposing the server. A preview OS-level tool sandbox exists on supported Linux and macOS hosts, with network access still unrestricted. [source]
- Model swapping: API requests can opt in to automatic switching between downloaded local GGUFs (2026-07-07); unknown model names keep using the current model. This makes Studio act like a llama-swap-style router. [source]
- Update mechanics: re-running the install command updates Studio; Desktop updates in-app. [source]
- 2026-03-17: Studio beta launch (Mac behaved like CPU: chat first; the install page later says Data Recipes work and "MLX training now works"). [source]
- 2026-03-25: app shortcuts on macOS; Data Recipes on macOS and CPU. [source]
- 2026-03-27 and 2026-04-01: auto-detect existing HF and LM Studio models, then user-selected folder. [source]
- 2026-04-03: Intel Macs work. [source]
- 2026-05-18: experimental MLX inference; API connections to OpenAI, Anthropic, vLLM, Ollama; MTP speculative decoding. [source]
- 2026-05-19: "improve MTP not being faster on Macs, CPUs and GPUs". [source]
- 2026-05-31: prebuilt llama.cpp binaries re-enabled for Apple Silicon M1-M4 on macOS 14, 15 and 26 (Tahoe); Apple Silicon on macOS 13 builds from source; Intel Macs on 13.3 to 26 use prebuilts. [source]
- 2026-06-12: Studio uses fresh llama.cpp prebuilts across CUDA, ROCm, Windows, Linux and macOS; improved MLX model labels and generation speed stats. [source]
- 2026-06-18: MLX training updates. [source]
- 2026-07-07: macOS installs no longer need CMake or Homebrew when a prebuilt llama.cpp exists; Apple Silicon handling of paths with spaces and unified memory sizing improves. [source]
- 2026-08-11: Unsloth Desktop (Beta) announced for Mac, Windows, Linux. [source]
- 2026-08-20: custom llama.cpp build support; toggles for Cache RAM, Mmap, Mlock, Checkpoints, speculative-decoding KV cache and vision on/off. [source]
- 2026-09-02: OpenAI-compatible APIs for video, audio and MLX-served models. [source]
- 2026-09-17: MLX gains video input and optional MoE and decode optimizations. [source]
- 2026-09-28: Apple Silicon improvements: batched serving, structured outputs and a TurboQuant KV cache. [source]
- 2026-10-01 (release v0.1.902-beta): updated MLX support; Mac LoRA exports to PEFT or GGUF. [source]
- macOS floor: manual Studio install requires macOS 12 Monterey or newer and a Python environment (uv, venv or conda); prebuilt llama.cpp covers macOS 14, 15, 26 on Apple Silicon, so macOS 13 compiles llama.cpp from source. [source]
- The doc page for `unsloth run` shows the model argument `unsloth/qwen3.8-27B-GGUF-GGUF:UD-Q4_K_XL` with a doubled `-GGUF` suffix, which looks like a typo and would not resolve as written. [source]
- Mac training depth is contested: the install page says Studio works on Mac for chat and "training and all features" with MLX training working, while a third-party review says training is the weaker path on Mac and NVIDIA is listed for training. [source]
- Studio's Mac chat path historically lacked MLX tools and web search: the 2026-05-18 notes say "thinking, tools and web search soon" for MLX, so older MLX chats may not call tools. [source]
- Desktop uninstall (drag to Trash) leaves the Studio data and HF model cache; manual removal uses `rm -rf ~/.unsloth/studio`, the launcher `~/Applications/Unsloth Studio.app`, and `~/.local/bin/unsloth`. [source]
- `UNSLOTH_STUDIO_HOME` isolates a second install (own venv, auth, `studio.db`, cache and llama.cpp build). [source]
- Bundled llama.cpp can be stale: Studio warns when the prebuilt is too old for MTP. [source]
- What "Desktop" is relative to "Studio": Unsloth's docs call Desktop the easiest way to install Studio and a native app that Studio is packaged into; a third-party review calls Desktop a superset of Studio and says its CLI is `unsloth start`, while the Unsloth pages show `unsloth run` and `unsloth studio`. The docs are the primary source. [source]
- Licensing and telemetry are stated by a third-party review (Apache 2.0 core, AGPL-3.0 for Studio components, no telemetry) and not on the pages read from unsloth.ai; treat as unconfirmed. [source]
- Star count: review says about 70.2k, the GitHub releases page at fetch shows about 77.2k; the review is dated August. [source]
- No source compares Desktop's MLX decode and prefill speed to mlx-lm, LM Studio or oMLX on the same model and chip. [source]
- Which MLX build and mlx-lm version Desktop bundles, and how it pins them, is not stated. [source]
- How Desktop's "TurboQuant KV cache" differs from llama.cpp's q8_0 or q4_0 cache-type options is not documented in the changelog entry. [source]
- Unsloth Desktop (Beta) is a free, open-source app for running and training models on macOS, Windows and Linux, built on Tauri. [source]
- The Mac install is a `.dmg`: open it, drag Unsloth to Applications, launch and wait for installation to complete, then pick a model and quantization from the Select model dropdown or Model hub tab. [source]
- `curl -fsSL https://unsloth.ai/install.sh | sh` installs Unsloth Studio (the browser UI), not Desktop or Unsloth Core, and running it again updates. [source]
- `unsloth studio -H 0.0.0.0 -p 8888` launches Studio, and by default it binds to 127.0.0.1 only. [source]
- `unsloth studio --secure` launches over HTTPS through a free Cloudflare tunnel on Windows, Mac and Linux. [source]
- Studio's macOS requirement is macOS 12 Monterey or newer on Intel or Apple Silicon, inside a Python environment such as uv, venv or conda. [source]
- The install page says Mac chat and Data Recipes work and "MLX training now works", and the intro says Training, MLX and GGUF inference all work inside Unsloth on macOS. [source]
- Models loaded in Unsloth, including GGUFs, are exposed as an authenticated API through `llama-server`, with Anthropic-compatible `/v1/messages` and OpenAI-compatible `/v1/chat/completions` and `/v1/responses`. [source]
- API keys start with `sk-unsloth-`, are created under Settings, API, are shown only once, and revoked keys return 401. [source]
- `unsloth run` serves a model with context size, GPU layers, threading, sampling, networking and tool settings. [source]
- Server-side web search and Python and terminal execution run as the local user and are on by default; the install page says to pass `--disable-tools` when exposing Unsloth. [source]
- Unsloth Desktop lets users pick sandboxed tool execution or direct file access, and asks approval for file access outside the sandbox. [source]
- Studio can connect to OpenAI, Anthropic, Ollama, llama.cpp and vLLM backends and supports parallel chatting. [source]
- 2026-05-18 changelog: experimental MLX inference lets MLX quants and models run on Macs, with thinking, tools and web search "coming soon". [source]
- 2026-05-31 changelog: prebuilt llama.cpp binaries re-enabled for Apple Silicon (M1-M4) on macOS 14, 15 and 26; Apple Silicon on macOS 13 is a source build; Intel Macs use prebuilts. [source]
- 2026-07-07 changelog: macOS installs no longer require CMake or Homebrew when a prebuilt llama.cpp is available, and Apple Silicon support handles paths with spaces and unified memory sizing better. [source]
- 2026-07-07 changelog: API requests can opt into automatic switching between downloaded local GGUFs and `/v1/models` returns clean model IDs. [source]
- 2026-08-20 changelog: custom llama.cpp builds and toggles for Cache RAM, Mmap, Mlock, Checkpoints, speculative-decoding KV cache and vision on/off. [source]
- 2026-09-02 changelog: new OpenAI-compatible APIs for video, audio and MLX-served models. [source]
- 2026-09-17 changelog: MLX gains video input, optional MoE and decode optimizations and more reliable multimodal chats; preview OS-level tool sandboxing arrives on supported Linux and macOS hosts. [source]
- 2026-09-28 changelog: Apple Silicon improvements include batched serving, structured outputs and a TurboQuant KV cache. [source]
- 2026-10-01 changelog (v0.1.902-beta): updated MLX support and Mac LoRA exports to PEFT or GGUF. [source]
- The 2026-08-11 changelog announces Unsloth Desktop for Mac, Windows and Linux supporting MLX, diffusion, audio and GGUF. [source]
- A Desktop uninstall is Finder, Applications, Move to Trash; manual Studio removal is `rm -rf ~/.unsloth/studio`, `rm -rf ~/Applications/Unsloth\ Studio.app` and `rm -f ~/.local/bin/unsloth`, and none of these touch downloaded HF model files. [source]
- Setting `UNSLOTH_STUDIO_HOME` installs Studio into an isolated location with its own virtual env, auth, `studio.db`, cache and llama.cpp build. [source]
- A third-party review states Unsloth Desktop ships Apache 2.0 core with AGPL-3.0 Studio components, no telemetry, and an `unsloth start` agent bridge, and calls Mac training "check before you plan around it". [source]
- On a Mac, Desktop's GGUF path is llama.cpp with Metal and its MLX path is a separate backend, so the same model can be run both ways in one app and compared for speed and fidelity locally. [source]
Children
- No children recorded.