Mac local LLMs: Ollama internals

Parent: Running LLM models locally on a Mac · Published reference · snapshot 2026-10-05

↓ Facts as markdownall context files

0.40.0-rc0 (pre-release, 2026-09-25): models on MLX-supported architectures run on MLX automatically on Apple silicon; GGUF tags may switch runner; no opt-out found; final 0.40.0 unknown (release URL 404 on 2026-10-04).

These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.

Runner choice and diagnosis

MLX memory and concurrency (act on these)

Create, quantize, import

Context length

mmap

llama-server launcher and bundled llama.cpp

Renderer, parser, Modelfile (undocumented)

Template selection (Go TEMPLATE vs GGUF Jinja)

Speculative decoding (MLX and GGUF)

Open questions

Corrections and disagreements

Concepts in this cluster

Children

← the whole tree · 3D view· how to read this page