<!-- llms-explorer concept facts · https://llms-explorer.com/tree/ollama-0-40-mlx-by-default-rollout-and-opt-out/ · pack 2026-10-05 · ~2329 tokens -->

# Ollama 0.40 MLX-by-default rollout and opt-out

> The rc0 body is two sentences plus an example (`ollama pull qwen3.8` then `ollama run qwen3.8`, with no `-mlx` tag), so the default now applies to untagged names, not just `-mlx` tags.

Parent: [Mac local LLMs: Runtime selection and frontends](https://llms-explorer.com/tree/mac-local-llms-runtime-selection-and-frontends/) · 1 facets · 39 facts · page: https://llms-explorer.com/tree/ollama-0-40-mlx-by-default-rollout-and-opt-out/

## Facts

- The rc0 body is two sentences plus an example (`ollama pull qwen3.8` then `ollama run qwen3.8`, with no `-mlx` tag), so the default now applies to untagged names, not just `-mlx` tags. — source: `asserted`
- "During the pre-release we will be testing and enabling additional models" means the supported-architecture list is open and grows rc by rc. — source: `asserted`
- Evidence of the pre-change behaviour: on 0.32.x/0.34.x the untagged tag ran GGUF and the `-mlx` tag ran MLX (e.g. qwen3.5:9b 15 GB vs qwen3.5:9b-mlx 9 GB in `ollama ps`, both 262144 context). — source: `asserted`
- Diagnostic for the active runner is unchanged: `ollama runner --mlx-engine` in the process list vs `llama-server` subprocess. — source: `asserted`
- Version spine to 0.40: 0.19.0 (2026-03-30) MLX preview; 0.30.8 (2026-06-12); 0.32.6 (early Aug 2026) MLX engine uses Qwen3.5's MTP head for speculative decoding automatically; 0.32.11 recurrent-state fix; 0.32.14/0.32.15 current during the 32 GB swap reports (2026-08-29); 0.33.3; 0.34.0-0.34.4 (2026-09-23); 0.35.0 (2026-09-28); 0.35.1 (2026-09-29, listed as Latest on the releases page); 0.40.0-rc0 (2026-09-25, Pre-release, built from 75b9527; v0.34.4...v0.40.0-rc0 is the diff range). — source: `asserted`
- The version jump 0.35 to 0.40 and the release ordering (rc0 dated before 0.35.0/0.35.1) show 0.40 is a separate pre-release line cut from 0.34.4, not a successor of 0.35.x. — source: `asserted`
- On 2026-10-04 the releases page still showed 0.35.1 as Latest and 0.40.0 as pre-release; 13 commits had landed on main since rc0. — source: `asserted`
- Issue 18131 (the 32 GB M1 Max swap report) was filed against 0.32.14, so the prefix-cache budget problem predates the default flip; after 0.40 more Mac users land on MLX and meet it. — source: `asserted`
- Issue 18132 is a same-day duplicate of 18131 (author closed it as duplicate on 2026-08-29); the shared ask is OLLAMA_MLX_PREFIX_CACHE_BYTES or a memory-scaled default, suggested 2-4 GiB on 32 GB hosts and 8 GiB or more on 64/128 GB. — source: `asserted`
- In 18131 the context was only ~12.6K tokens while resident was ~28 GB, wired 22.7 GB, compressed 3.8 GB, 30.7 GB used of 32 GB, so swap came from model plus snapshot budget, not KV for the 73728 num_ctx. — source: `asserted`
- Upgrade side effect to expect: a name that ran GGUF with MTP (GGUF+MTP 11.08 tok/s in 18131) can now run MLX (16.58 tok/s), so decode, memory shape and swap behaviour all change with no tag change. — source: `asserted`
- MoE caveat: a third-party bake-off on a Mac Studio found Ollama (llama.cpp GGUF path) faster than MLX on a 30B-A3B MoE, 91.23 vs 68.06 tok/s at 512-token prompts and 57.76 vs 45.62 at 16K, so MLX-by-default is not uniformly faster for MoE models. — source: `asserted`
- Dense vs MoE: same bake-off shows MLX 41% faster on a dense 27B 4-bit (30.72 vs 20.86 tok/s at 512) and the lead vanishing at 8-bit, while Ollama wins MoE by 20-22%. Ollama's own blog numbers (NVFP4 about 20% over q4_K_M) sit between. Different models, quantisations and Ollama versions; unreconciled. — source: `asserted`
- 9B Qwen3.5 on M3 Ultra: Ollama MLX tag 90.5 tok/s vs Ollama GGUF 77.7, mlx-lm 107.5, rapid-mlx 105.9; the MLX tag's number includes automatic MTP speculative decoding the others did not use. — source: `asserted`
- Which architectures ship in final 0.40.0, and whether the rc series adds an opt-out; none found in rc0 text, in docs.ollama.com/faq or docs.ollama.com/gpu (neither mentions MLX). — source: `asserted`
- Whether OLLAMA_LLM_LIBRARY=mlx has any effect; no source found beyond the already-held "ignored on GGUF" note. — source: `asserted`
- Whether a prefix-cache budget variable exists in 0.40; the only evidence is the feature request in 18131. — source: `asserted`
- Rollback path: no source states one. Practical options inferred below. — source: `asserted`
- The v0.40.0 release page (tag v0.40.0-rc0, commit 75b952780f90807f651eb2f1f817e5a40126e81d, github-actions, 25 Sep 2026) is marked Pre-release and had 13 further commits on main since. — [source](https://github.com/ollama/ollama/releases/tag/v0.40.0-rc0)
- The rc0 notes give the example `ollama pull qwen3.8` / `ollama run qwen3.8` as running on MLX on Apple silicon with no tag suffix. — [source](https://github.com/ollama/ollama/releases/tag/v0.40.0-rc0)
- The rc0 notes name no supported architectures, no environment variable, flag or Modelfile option to disable MLX, and no upgrade or rollback guidance. — [source](https://github.com/ollama/ollama/releases/tag/v0.40.0-rc0)
- The rc0 compare range is v0.34.4...v0.40.0-rc0, so 0.40 branches from 0.34.4 rather than from 0.35.x. — [source](https://github.com/ollama/ollama/releases/tag/v0.40.0-rc0)
- On 2026-10-04 the releases page lists v0.35.1 as Latest, and v0.35.1 is the newest stable. — [source](https://github.com/ollama/ollama/releases)
- Ollama v0.32.6 (early August 2026) release note: "Qwen3.5 is faster on Apple GPUs: the MLX engine now uses the model's MTP head for speculative decoding automatically." — [source](https://terminalbytes.com/ollama-vs-llama-cpp-vs-mlx-mac-2026)
- On a Mac Studio M3 Ultra, Qwen3.5 9B 4-bit generated at Ollama GGUF 77.7, llama-bench 77.8, Ollama MLX 90.5, rapid-mlx 105.9 and mlx-lm 107.5 tok/s. — [source](https://terminalbytes.com/ollama-vs-llama-cpp-vs-mlx-mac-2026)
- `ollama ps` showed qwen3.5:9b-mlx at 9.0 GB and qwen3.5:9b (GGUF) at 15 GB, both with 262144 context, because Ollama loads Qwen3.5 at its full 262,144-token window by default. — [source](https://terminalbytes.com/ollama-vs-llama-cpp-vs-mlx-mac-2026)
- The same author advises setting `num_ctx` to a value you will use on short-memory Macs, and notes the Ollama-downloaded GGUF blob will not load in upstream llama.cpp. — [source](https://terminalbytes.com/ollama-vs-llama-cpp-vs-mlx-mac-2026)
- Ollama MLX tag processed a 34-token prompt at 26 to 37 tok/s versus 344 tok/s on its GGUF path, which the author reads as fixed startup cost not a rate. — [source](https://terminalbytes.com/ollama-vs-llama-cpp-vs-mlx-mac-2026)
- Zach Rattner bake-off (Mac Studio, 54 runs, 3 prompt lengths, 256 generated tokens): dense Qwen3.8-27B 4-bit decode was Ollama 20.86 / 22.07 / 19.00 vs MLX 30.72 / 29.61 / 27.41 tok/s at 512 / 4,096 / 16,384 prompt tokens. — [source](https://zachrattner.com/projects/ai-mac-cluster/mlx-vs-ollama)
- In the same bake-off the 30B-A3B MoE 4-bit decode was Ollama 91.23 / 80.50 / 57.76 vs MLX 68.06 / 60.01 / 45.62 tok/s, and Ollama led Rapid-MLX by 20-22% at every length. — [source](https://zachrattner.com/projects/ai-mac-cluster/mlx-vs-ollama)
- The author reports the MLX lead on the dense 27B vanishes at 8-bit (MLX 8-bit vs Ollama Q8_0) and that Rapid-MLX measured 1.40x-1.53x Ollama on dense and 0.79x on MoE against its 4.2x headline. — [source](https://zachrattner.com/projects/ai-mac-cluster/mlx-vs-ollama)
- Issue 18132 ("MLX prefix cache: fixed 8 GiB budget causes heavy swap on 32 GB Apple Silicon during agent workloads") was opened 2026-08-29 by Vr1155 and closed the same day as a duplicate of 18131. — [source](https://github.com/ollama/ollama/issues/18132)
- 18132 reports on M1 Max 32 GB, Ollama 0.32.14, OpenCode 1.18.25, alias qwen3.8-27b-72k-mlx with num_ctx 73728: process 18 GB after load, 20, 24, 28 GB later; Activity Monitor 30.7 GB used, Ollama 28.2 GB, wired 22.7 GB, compressed 3.8 GB, swap 10.1 GB, OpenCode context ~12.6K tokens. — [source](https://github.com/ollama/ollama/issues/18132)
- 18132 benchmarks: GGUF/Metal+MTP 11.08 vs MLX 16.58 tok/s; repo-analysis task 1341 s (~29.9K final context) vs 696 s (~25.3K final context). — [source](https://github.com/ollama/ollama/issues/18132)
- 18132 proposes `OLLAMA_MLX_PREFIX_CACHE_BYTES=<bytes>` and a unified-memory-scaled default (32 GB host 2-4 GiB; 64/128 GB keep 8 GiB or more), optionally reacting to macOS memory pressure. — [source](https://github.com/ollama/ollama/issues/18132)
- 18132 and 18131 carried no maintainer comment or assignee or label at fetch time (2026-10-04). — [source](https://github.com/ollama/ollama/issues/18132)
- docs.ollama.com/faq and docs.ollama.com/gpu, fetched 2026-10-04, contain no MLX opt-out or runner-selection setting. — [source](https://docs.ollama.com/faq)
- Inferred rollback options on macOS: pin the previous stable (0.35.1 or earlier; MLX was already opt-in by `-mlx` tag there) with `brew install ollama@<ver>`-style pinning or the app zip, or on 0.40 use `OLLAMA_KEEP_ALIVE=0` and a smaller `num_ctx` to bound memory; neither is documented. — source: `asserted`
- Inferred 32 GB guidance: before moving to 0.40, expect a 27B MLX model (18-19 GB) plus up to 8 GiB snapshot cache to swap under agent workloads; MoE models may be faster on the GGUF path, so test both before accepting the default. — source: `asserted`
