On-Device & Local LLM Runtimes

Parent: LLM Models and APIs · researched 2026-06-03T22:51:20.748Z· 18 sources · 12 concepts · skill on-device-local-llm-runtimes

The local/on-device runtime + developer-experience layer: which runtime to

On-Device & Local LLM Runtimes

When to use / Skip

Why run locally

Ollama — the default on-ramp

llama.cpp — the engine under everything

LM Studio — GUI + SDK + MLX backend

Apple MLX / MLX-LM — Apple Silicon native

MLC-LLM — universal deploy via ML compilation

Long tail (one line each)

Browser & edge runtimes

Mobile runtimes

Hardware sizing & quant-level selection

Integration patterns

Anti-patterns & failure modes

2025-2026 frontier

Sources

Children

Frontier under this node: Apple Foundation Models framework & @Generable guided generation, Apple MLX & MLX-LM (unified-memory inference), Chrome Built-in AI / Gemini Nano Prompt API, GBNF grammars & local schema-constrained decoding, LM Studio (GUI + SDK + MLX backend), Local hardware sizing & quant-level selection (RAM/VRAM, KV cache, Q4/Q5/Q8), MLC-LLM & MLCEngine (universal/compiled deploy), Mobile on-device runtimes (MediaPipe to LiteRT-LM, ONNX Runtime GenAI / QNN), Ollama Modelfile & local model library, WebLLM & Transformers.js (WebGPU in-browser inference), Windows AI Foundry / Foundry Local & on-device NPU inference, llama.cpp / llama-server & the GGUF format

← the whole tree · 3D view· how to read this page