<!-- llms-explorer concept facts · https://llms-explorer.com/tree/running-llm-models-locally-on-mac/ · pack 2026-10-05 · ~1030 tokens -->

# Running LLM models locally on a Mac

> Cited expert reference on running llm models locally on mac: 502 researched concepts in 21 clusters.

Parent: [On-Device & Local LLM Runtimes](https://llms-explorer.com/tree/on-device-local-llm-runtimes/) · 3 facets · 34 facts · page: https://llms-explorer.com/tree/running-llm-models-locally-on-mac/

## Search

- grep `topics/*/llms-facts.txt` for exact terms; read `topics/<key>/llms-full.txt` by line range. — source: `asserted`
- `llms-facts.txt` here holds every CONTRADICTS correction. — source: `asserted`

## Cross-cutting facts

- Ollama v0.40.0-rc0 runs Apple-silicon models on MLX with no tag suffix; no opt-out flag named. — [source](https://github.com/ollama/ollama/releases/tag/v0.40.0-rc0)
- llama-server `/health` stays 200 after a Metal failure; real failure is HTTP 500 "Compute error."; probe with a completion and SIGKILL. — [source](https://github.com/ggml-org/llama.cpp/issues/27309)
- Wired cap is `iogpu.wired_limit_mb`; a missing key silently does nothing, 0 restores default, resets on reboot. — [source](https://llmconfigurator.com/en/guides/troubleshooting/increase-metal-vram-limit-apple-silicon)
- Kernel panic `completeMemory() prepare count underflow` at `IOGPUMemory.cpp:550` on long prefill; mitigate with mlx-lm 0.31.0 and a wired limit below RAM. — [source](https://github.com/ml-explore/mlx/issues/3186)
- Decode is bandwidth-bound: M5 Max 460 GB/s (32-core GPU) or 614 GB/s (40-core). — [source](https://support.apple.com/en-us/126319)
- KLD: `llama-perplexity --kl-divergence-base f`, then `--kl-divergence`; compare at equal GiB (q4_0 4.34 GiB is 67% worse than q4_K_S 4.37). — [source](https://github.com/ggml-org/llama.cpp/blob/master/tools/perplexity/README.md)
- Claude Code's `x-anthropic-billing-header ... cch=` breaks prefix caching; llama.cpp rewrites it (PR 21793); else `CLAUDE_CODE_ATTRIBUTION_HEADER=0` in settings env. — [source](https://unsloth.ai/docs/basics/claude-code)
- Codex needs `/v1/responses`; `model_auto_compact_token_limit_scope` is `total` (default) or `body_after_prefix`. — [source](https://developers.openai.com/codex/config-reference)
- JACCL needs a full Thunderbolt mesh and `MLX_METAL_FAST_SYNCH=1`; mlx 0.32.3 fixes a fence deadlock (PR 4552). — [source](https://github.com/ml-explore/mlx/pull/4552)
- ANE is FP16-only; activations above 65,504 overflow. — [source](https://github.com/Anemll/Anemll/blob/main/docs/FP16_SCALING.md)
- Pin llama.cpp commit and mlx-lm version, same bpw, greedy batch 1; report median of runs 2-4. — [source](https://raw.githubusercontent.com/john-rocky/apple-silicon-llm-bench/main/methodology/fairness-rules.md)

## Clusters

- Mac local LLMs: Runtime selection and frontends — source: `asserted`
- Mac local LLMs: Serving ops and multi-model — source: `asserted`
- Mac local LLMs: Memory and wired limits — source: `asserted`
- Mac local LLMs: GPU stability and kernel panics — source: `asserted`
- Mac local LLMs: Speed, bandwidth and prefill — source: `asserted`
- Mac local LLMs: Benchmarking and comparisons — source: `asserted`
- Mac local LLMs: Quantization formats and methods — source: `asserted`
- Mac local LLMs: Quantization evaluation — source: `asserted`
- Mac local LLMs: KV cache sizing and quantization — source: `asserted`
- Mac local LLMs: Prompt cache and persistent KV — source: `asserted`
- Mac local LLMs: Speculative decoding and MTP — source: `asserted`
- Mac local LLMs: MoE streaming and offload — source: `asserted`
- Mac local LLMs: Clusters, RDMA, exo and ds4 — source: `asserted`
- Mac local LLMs: Agent clients, context and compaction — source: `asserted`
- Mac local LLMs: Chat templates, reasoning and tool calling — source: `asserted`
- Mac local LLMs: Apple Foundation Models and Core AI — source: `asserted`
- Mac local LLMs: ANE and Core ML LLMs — source: `asserted`
- Mac local LLMs: MLX kernels, numerics and internals — source: `asserted`
- Mac local LLMs: llama.cpp internals — source: `asserted`
- Mac local LLMs: Ollama internals — source: `asserted`
- Mac local LLMs: oMLX, Rapid-MLX and related internals — source: `asserted`
