Running LLM models locally on a Mac
Parent: On-Device & Local LLM Runtimes · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Cited expert reference on running llm models locally on mac: 502 researched concepts in 21 clusters.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Search
Cross-cutting facts
- Ollama v0.40.0-rc0 runs Apple-silicon models on MLX with no tag suffix; no opt-out flag named. [source]
- llama-server `/health` stays 200 after a Metal failure; real failure is HTTP 500 "Compute error."; probe with a completion and SIGKILL. [source]
- Wired cap is `iogpu.wired_limit_mb`; a missing key silently does nothing, 0 restores default, resets on reboot. [source]
- Kernel panic `completeMemory() prepare count underflow` at `IOGPUMemory.cpp:550` on long prefill; mitigate with mlx-lm 0.31.0 and a wired limit below RAM. [source]
- Decode is bandwidth-bound: M5 Max 460 GB/s (32-core GPU) or 614 GB/s (40-core). [source]
- KLD: `llama-perplexity --kl-divergence-base f`, then `--kl-divergence`; compare at equal GiB (q4_0 4.34 GiB is 67% worse than q4_K_S 4.37). [source]
- Claude Code's `x-anthropic-billing-header ... cch=` breaks prefix caching; llama.cpp rewrites it (PR 21793); else `CLAUDE_CODE_ATTRIBUTION_HEADER=0` in settings env. [source]
- Codex needs `/v1/responses`; `model_auto_compact_token_limit_scope` is `total` (default) or `body_after_prefix`. [source]
- JACCL needs a full Thunderbolt mesh and `MLX_METAL_FAST_SYNCH=1`; mlx 0.32.3 fixes a fence deadlock (PR 4552). [source]
- ANE is FP16-only; activations above 65,504 overflow. [source]
- Pin llama.cpp commit and mlx-lm version, same bpw, greedy batch 1; report median of runs 2-4. [source]
Clusters
- Mac local LLMs: Runtime selection and frontends [source]
- Mac local LLMs: Serving ops and multi-model [source]
- Mac local LLMs: Memory and wired limits [source]
- Mac local LLMs: GPU stability and kernel panics [source]
- Mac local LLMs: Speed, bandwidth and prefill [source]
- Mac local LLMs: Benchmarking and comparisons [source]
- Mac local LLMs: Quantization formats and methods [source]
- Mac local LLMs: Quantization evaluation [source]
- Mac local LLMs: KV cache sizing and quantization [source]
- Mac local LLMs: Prompt cache and persistent KV [source]
- Mac local LLMs: Speculative decoding and MTP [source]
- Mac local LLMs: MoE streaming and offload [source]
- Mac local LLMs: Clusters, RDMA, exo and ds4 [source]
- Mac local LLMs: Agent clients, context and compaction [source]
- Mac local LLMs: Chat templates, reasoning and tool calling [source]
- Mac local LLMs: Apple Foundation Models and Core AI [source]
- Mac local LLMs: ANE and Core ML LLMs [source]
- Mac local LLMs: MLX kernels, numerics and internals [source]
- Mac local LLMs: llama.cpp internals [source]
- Mac local LLMs: Ollama internals [source]
- Mac local LLMs: oMLX, Rapid-MLX and related internals [source]
Children
- Mac local LLMs: Agent clients, context and compaction
- Mac local LLMs: ANE and Core ML LLMs
- Mac local LLMs: Apple Foundation Models and Core AI
- Mac local LLMs: Benchmarking and comparisons
- Mac local LLMs: Chat templates, reasoning and tool calling
- Mac local LLMs: Clusters, RDMA, exo and ds4
- Mac local LLMs: GPU stability and kernel panics
- Mac local LLMs: KV cache sizing and quantization
- Mac local LLMs: llama.cpp internals
- Mac local LLMs: Memory and wired limits
- Mac local LLMs: MLX kernels, numerics and internals
- Mac local LLMs: MoE streaming and offload
- Mac local LLMs: Ollama internals
- Mac local LLMs: oMLX, Rapid-MLX and related internals
- Mac local LLMs: Prompt cache and persistent KV
- Mac local LLMs: Quantization evaluation
- Mac local LLMs: Quantization formats and methods
- Mac local LLMs: Runtime selection and frontends
- Mac local LLMs: Serving ops and multi-model
- Mac local LLMs: Speculative decoding and MTP
- Mac local LLMs: Speed, bandwidth and prefill