<!-- llms-explorer concept facts · https://llms-explorer.com/tree/mac-local-llms-apple-foundation-models-and-core-ai/ · pack 2026-10-05 · ~3096 tokens -->

# Mac local LLMs: Apple Foundation Models and Core AI

> macOS 27 shipped 2026-09-14. SystemLanguageModel exists from macOS 26.0 (versions 26.0-26.3, 26.4, 27.0); 26.4 added `contextSize` and `tokenCount(for:)`.

Parent: [Running LLM models locally on a Mac](https://llms-explorer.com/tree/running-llm-models-locally-on-mac/) · 9 facets · 49 facts · page: https://llms-explorer.com/tree/mac-local-llms-apple-foundation-models-and-core-ai/

## System model, versions, context

- macOS 27 shipped 2026-09-14. SystemLanguageModel exists from macOS 26.0 (versions 26.0-26.3, 26.4, 27.0); 26.4 added `contextSize` and `tokenCount(for:)`. — [source](https://developer.apple.com/documentation/foundationmodels/systemlanguagemodel)
- On-device context is disputed: session 241 prints 8192; session 319 says "4k" vs 32K PCC; TN3193 says 4096. Read `SystemLanguageModel.default.contextSize`; overflow throws. — [source](https://developer.apple.com/videos/play/wwdc2026/319/)
- On-device: offline, no request limits. PCC: needs network, daily per-user limit, free under 2M first-time downloads with an entitlement. — [source](https://developer.apple.com/videos/play/wwdc2026/319/)
- Adapter toolkit ended at 26.0.0 (not OS 27 compatible); adapters ~160 MB. — [source](https://developer.apple.com/apple-intelligence/foundation-models-adapter/)

## fm CLI and fm serve

- /usr/bin/fm refuses until `sudo fm license` ("YOU HAVE NOT AGREED TO THE FOUNDATION MODELS CLI LEGAL NOTICE & TERMS"); applies to all users. — [source](https://ai.plainenglish.io/exploring-the-fm-cli-and-apple-foundation-model-in-macos-27-213f6481c7cc)
- Subcommands: respond, chat, schema, count-tokens, available, license, serve. `respond`: `-i`, `--text`, `--schema`, `-g`, `--no-stream`, `--save-transcript`, `--resume`, `--image`, `--tool` (barcode, ocr). — [source](https://medium.com/@michael.hannecke/apples-model-at-the-command-line-what-fm-serve-leaves-to-you-6a7a702d66c9)
- `fm serve`: 127.0.0.1:1976, /health, /v1/models, /v1/chat/completions, model id `system`. Recaps citing localhost:8000 and "apple-fm" were wrong. — [source](https://medium.com/macoclock/turn-your-mac-into-a-free-openai-compatible-server-with-one-command-18f255464a99)
- Not OpenAI drop-in: streams by default (send `"stream": false`); `tool_choice: "auto"` puts raw JSON in `content`, no `tool_calls`; `"required"` gives HTTP 500. — [source](https://medium.com/@michael.hannecke/apples-model-at-the-command-line-what-fm-serve-leaves-to-you-6a7a702d66c9)
- No auth (any bearer token returns 200); only 403 "Cross-site requests are not allowed"; Apple's help example binds 0.0.0.0. `--socket` is srwxr-xr-x, so use a 0700 directory. — [source](https://medium.com/@michael.hannecke/apples-model-at-the-command-line-what-fm-serve-leaves-to-you-6a7a702d66c9)
- Schema is shape only: injection followed 20/20; missing `age` became 30; 8/60 runs hit a context-size error; `--save-transcript` is 0644. — [source](https://medium.com/@michael.hannecke/apples-model-at-the-command-line-what-fm-serve-leaves-to-you-6a7a702d66c9)
- Python: `pip install apple-fm-sdk`, `import apple_fm_sdk as fm`; macOS 26.0+, matching Xcode, Python 3.10+. Wrong Xcode gives "model not available". — [source](https://apple.github.io/python-apple-fm-sdk/getting_started.html)

## LanguageModel protocol and Dynamic Profiles

- Executors share KV cache per equal Hashable configuration, get the full transcript each call and must diff it. `prewarm` may not run; load lazily. — [source](https://developer.apple.com/videos/play/wwdc2026/339/)
- Appending keeps the KV cache; removing entries, changing tools or editing instructions invalidates it. — [source](https://developer.apple.com/videos/play/wwdc2026/242/)
- `historyTransform` alters only what the model sees (example `history.suffix(20)`); hitting maximumResponseTokens stops without throwing; switching remote/on-device loses the cache. — [source](https://artemnovichkov.com/blog/building-a-custom-dynamic-profile-modifier-in-foundation-models)

## mlx-swift-lm adapter and AnyLanguageModel

- mlx-swift-lm 3.32.3 (2026-09-30) first includes MLXFoundationModels: `.upToNextMinor(from: "3.32.3")`, swift-huggingface 0.9.0, swift-transformers 1.3.0. Floors stay macOS 14; adapter is `@available(macOS 27.0)`. — [source](https://github.com/ml-explore/mlx-swift-lm)
- Use `#huggingFaceLanguageModel(configuration: LLMRegistry.qwen3_0_6b_4bit, capabilities: [.reasoning])`; default capability `.guidedGeneration`. On the 26 SDK the trait `FoundationModelsIntegration` compiles to an empty module. — [source](https://github.com/ml-explore/mlx-swift-lm/blob/main/Libraries/MLXFoundationModels/README.md)
- Guided generation masks logits (JSON Schema, EBNF, structural tags). First call per schema blocks hundreds of ms: avoid `@MainActor`, pre-warm a throwaway `GrammarConstraint`. — [source](https://github.com/ml-explore/mlx-swift-lm/blob/main/Libraries/MLXGuidedGeneration/README.md)
- Issues: #543 (no compile on Xcode 27 beta 5; final status unverified); #539 (27-SDK test bundle crashes on macOS 26 host). SwiftPM 9286: "exhausted attempts to resolve the dependencies graph"; add trait packages directly. — [source](https://github.com/ml-explore/mlx-swift-lm/issues/539)
- AnyLanguageModel 0.15.1 (2026-10-02) is maintained, not pre-27: Swift 6.3+, traits CoreML/MLX/Llama, Ollama at http://localhost:11434, `MLXLanguageModel(modelId:)`; MLX tests need `xcodebuild`. — [source](https://github.com/mattt/AnyLanguageModel)
- Rule: mlx-swift-lm `ChatSession` for macOS 14+; MLXFoundationModels on the 27 SDK; AnyLanguageModel for macOS 14-26 or llama.cpp/Ollama. — source: `asserted`

## Core AI

- `uv run coreai.llm.export <HF_ID> [--platform iOS]`, then `CoreAILanguageModel(resourcesAt:)` in `LanguageModelSession(model:)`. Start ~0.6B; Apple ships recipes only. Needs OS and Xcode 27+. — [source](https://github.com/apple/coreai-models)
- macOS export 4bit INT4 with dynamic KV; iOS needs `--max-context-length`. Engine (GPU/CPU/ANE) depends on export. AOT: `xcrun coreai-build compile`; builds with `.aimodel` fail without the Metal Toolchain; caches reset on OS update. — [source](https://developer.apple.com/documentation/coreai/managing-model-specialization-and-caching)
- Trap: same Qwen3-0.6B export ran 1,121 tok/s on macOS 26 but 484 on a macOS 27 beta; re-export does not fix. gpt-oss-20B export needed 128 GB. — [source](https://rockyshikoku.medium.com/7-llms-pre-converted-to-apples-core-ai-format-aimodel-now-on-hugging-face-0ad996e921e8)
- M4 Max, one author, Core AI vs MLX: gpt-oss-20b 78.1 vs 100.2; qwen3-8b 90.0 vs 94.1; qwen3-4b tie. Prefer MLX on Mac for HF, MoE, hybrids. — [source](https://rockyshikoku.medium.com/apple-core-ai-vs-mlx-which-is-faster-on-iphone-and-mac-for-the-same-model-8faf3c8784ff)

## Utilities package and prefix cache

- `ChatCompletionsLanguageModel(name:url:additionalHeaders:supportsGuidedGeneration:urlSessionConfiguration:)`; `v<digits>` path part gets `/chat/completions`, else `/v1/chat/completions`; README's `http://localhost/v1:8000` is malformed. Set `supportsGuidedGeneration: false` if `response_format` is ignored. — [source](https://github.com/apple/foundation-models-utilities)
- Always sends `stream_options`, `max_completion_tokens`; rejecting servers return HTTP 400. Seed throws `RequestError.invalidRequest`. — [source](https://github.com/apple/foundation-models-utilities)
- `summarizeHistory` counts entries, not tokens; pair with `rollingWindow`. Head rewrites defeat server prefix caches; prompt-based skills append only. — source: `asserted`
- llama-server `n_cache_reuse` shifts matching chunks, else logs "cache reuse is not supported". oMLX hashes chain from the head. Claude Code deleting old tool results forces re-prefill; `CLAUDE_AUTOCOMPACT_PCT_OVERRIDE=95` worsens it. — [source](https://github.com/jundot/omlx/issues/2333)

## Corrections and open questions

- Fixed: shipping fm 2.0.68.1.402 lists only `system` (no PCC); Python SDK is macOS 26.0+; Core AI is not ANE-only. — source: `asserted`
- Unverified: 3.32.3 on final Xcode 27.x; `fm serve` as ChatCompletions backend; Anthropic/Google packages shipped. — source: `asserted`

## Corrections and disagreements

- CONTRADICTS: on-device-local-llm-runtimes.md says fm CLI offers PCC; the shipping macOS 27.0 build (fm 2.0.68.1.402) tested lists only the `system` model while WWDC and a beta write-up show `--model pcc` and `quota-usage`. — [source](https://medium.com/@michael.hannecke/apples-model-at-the-command-line-what-fm-serve-leaves-to-you-6a7a702d66c9)
- CONTRADICTS: on-device-local-llm-runtimes.md says OS 27 on-device context is 8192; session 319 says the on-device model offers "4k" context vs 32K for PCC, and TN3193 says 4096 per session. — [source](https://developer.apple.com/videos/play/wwdc2026/319/)
- CONTRADICTS (within sources): articles show `MLXLanguageModel(modelID: "mlx-community/...")`, but the adapter README documents a macro or `configuration:`/`weightsLocation:`/`load:`; the session sample's `modelID:` is Apple's placeholder. — [source](https://medium.com/@nuthalapativarun/mlx-is-now-a-first-class-citizen-in-apples-ai-stack-run-any-hugging-face-model-through-foundation-9dfb8dad2191)
- CONTRADICTS: on-device-local-llm-runtimes.md lists the Python SDK under OS 27; the SDK requires macOS 26.0+ and has existed since 2026-02. — [source](https://apple.github.io/python-apple-fm-sdk/getting_started.html)
- CONTRADICTS (soft): on-device-local-llm-runtimes.md lists "rank-32 LoRA adapters that must be retrained per model version" without noting that no adapter toolkit exists for OS 27; the toolkit ended at 26.0.0. — [source](https://developer.apple.com/apple-intelligence/foundation-models-adapter/)
- Does CoreAILanguageModel run on the ANE or GPU? CONTRADICTS apple-foundation-models-framework.md line 96 ("CoreAILanguageModel (Neural Engine)"): Apple's article says it runs on GPU, CPU or ANE depending on export; macOS export is dynamic and lands on the GPU pipelined engine, iOS static export on the ANE. — source: `asserted`
- CONTRADICTS apple-foundation-models-framework.md (soft): that dossier calls AnyLanguageModel a "pre-27" package; v0.15.1 (2026-10-02) is still maintained and mirrors the 27 API (DynamicInstructions, Usage, reasoning entries) and wraps any 27 `LanguageModel` with `FoundationLanguageModel`. — [source](https://github.com/huggingface/AnyLanguageModel)
- CONTRADICTS a macOS 27 Medium article: its Package.swift uses `.upToNextMinor(from: "1.0.0")` and `.macOS(.v27)` for mlx-swift-lm, but the repo is at 3.32.3 with macOS 14 floors. — [source](https://medium.com/@nuthalapativarun/mlx-is-now-a-first-class-citizen-in-apples-ai-stack-run-any-hugging-face-model-through-foundation-9dfb8dad2191)
- CONTRADICTS the repo's SKILL.md (stale): it describes traits ChatCompletions, Skills and History and a `.model(...)` profile call; Package.swift has no traits and beta 3 removed `.model(...)`. — [source](https://raw.githubusercontent.com/apple/foundation-models-utilities/main/Package.swift)
- CONTRADICTS apple-foundation-models-framework.md on provenance only: a June 29 2026 article says Apple's WWDC session on the fm CLI shows fm respond, fm chat and fm schema and never mentions serve, so early fm serve walkthroughs were not backed by Apple's session. — [source](https://medium.com/macoclock/turn-your-mac-into-a-free-openai-compatible-server-with-one-command-18f255464a99)
- CONTRADICTS the early recaps: the widely shared claims of http://localhost:8000/v1/ and model id "apple-fm" for fm serve are not what the shipping build uses (127.0.0.1:1976, model system per the existing dossier). — [source](https://medium.com/macoclock/turn-your-mac-into-a-free-openai-compatible-server-with-one-command-18f255464a99)

## Concepts in this cluster

- Apple Foundation Models framework — source: `asserted`
- Core AI framework — source: `asserted`
- mlx-swift-lm and MLXFoundationModels adapter — source: `asserted`
- Apple foundation-models-utilities ChatCompletionsLanguageModel — source: `asserted`
- MLXGuidedGeneration and XGrammar constrained decoding on Apple Silicon — source: `asserted`
- Server-side prefix cache behavior under rolling-window transcripts — source: `asserted`
- Dynamic Profiles historyTransform — source: `asserted`
- fm serve and the fm CLI OpenAI-compatible endpoint — source: `asserted`
