oMLX and Rapid-MLX runtimes
Parent: Mac local LLMs: oMLX, Rapid-MLX and related internals · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
oMLX engine layers: FastAPI HTTP layer, BatchedEngine / VLMBatchedEngine over mlx-lm BatchGenerator, an EnginePool that holds several models with LRU eviction, pinning and per-model TTL, and a per-model settings store in `~/.omlx/settings.json` (profiles, alias, chat-template kwargs, model-type o...
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- oMLX engine layers: FastAPI HTTP layer, BatchedEngine / VLMBatchedEngine over mlx-lm BatchGenerator, an EnginePool that holds several models with LRU eviction, pinning and per-model TTL, and a per-model settings store in `~/.omlx/settings.json` (profiles, alias, chat-template kwargs, model-type override). [source]
- oMLX model discovery: point `--model-dir` at a folder of MLX subdirectories (two-level `org/model` folders work). Types auto-detected: LLM (anything mlx-lm loads), VLM (via mlx-vlm), OCR (DeepSeek-OCR, DOTS-OCR, GLM-OCR), embedding (BERT, BGE-M3, ModernBERT), reranker (ModernBERT, XLM-RoBERTa). It can also reuse the Hugging Face hub cache and an existing LM Studio folder. [source]
- oMLX API surface: `/v1/chat/completions`, `/v1/completions`, `/v1/messages` (Anthropic), `/v1/embeddings`, `/v1/rerank`, `/v1/models`; the dashboard lives at `/admin`, with built-in chat at `/admin/chat`. Streaming usage stats, Anthropic adaptive thinking and image/video inputs are supported. Tool-call formats are auto-detected for Llama/Qwen/DeepSeek JSON, Qwen3.5 XML, Gemma, GLM, MiniMax, Mistral, Kimi K2, Longcat and IFM K2 Horizon; MCP tools need `pip install mcp` or the `[mcp]` extra. [source]
- oMLX admin dashboard: real-time monitoring, model download from Hugging Face, per-model settings applied without restart, one-click benchmark (PP and TG tok/s with partial prefix-hit testing), accuracy runs, one-click integration config for OpenClaw, OpenCode, Codex, Hermes Agent, Copilot, Pi and DeepSeek Harness, eight UI languages, vendored assets so it works offline. [source]
- oMLX menu-bar app: native Swift/SwiftUI (source `apps/omlx-mac`), signed and notarized DMG with in-app auto-update, auto-restart on crash, local usage history with hourly heatmap, and an `~/.omlx/bin/omlx` CLI shim that controls the app-managed server. [source]
- oMLX optional native custom kernels (GLM-5.2, MiniMax M3, Qwen3.5) ship precompiled only in the DMG; pip and Homebrew builds need full Xcode with the Metal toolchain, and without them those families fall back to much slower generic paths. [source]
- oMLX 0.6.0 added experimental distributed serving (tensor or pipeline parallelism over Ring or Thunderbolt RDMA/JACCL, with a Cluster dashboard) and heterogeneous Metal plus CUDA worker pools; it is disabled by default and documented as a source-build preview. [source]
- Rapid-MLX engine surface: `/v1/chat/completions`, `/v1/responses` (Codex), `/v1/messages`, `/v1/embeddings`, `/v1/audio/*`, `/v1/videos`, `/v1/images/generations`, `/v1/systemone`, and an authenticated `/v1/cua` computer-use API. CLI verbs mirror Ollama (run, pull, rm, ps, models) plus serve, chat, launch, agents, recipe, benchmark, doctor, upgrade, telemetry. [source]
- Rapid-MLX packaging: base install is text-only (~460 MB); vision, audio, image, video, embeddings and DFlash are pip extras (`[vision]`, `[audio]`, `[image]`, `[system-one]`, `[all]`). `rapid-mlx serve` picks the first free port in 8000-8009 when `--port` is omitted. [source]
- Rapid-MLX agent integration is tiered: five Tier-1 agents (Claude Code, Codex CLI, Hermes, Aider, DeepSeek Harness) are driven through a real multi-step bug-fix against a local 35B model before any release can tag; seven Tier-2 agents are wire-verified only. `rapid-mlx launch claude-code` and `rapid-mlx agents <name> --setup` write the client config. [source]
- Rapid-MLX accelerated profiles (0.15.4) pin one checkpoint plus its drafter or MTP head, admit one request at a time, reject tools, media and grammar constraints, and fail closed outside the qualified model and runtime revisions. [source]
- oMLX repository created 2026-02-13 ("initial commit"); announced 2026-03-04; 22.5k stars and 3,030 commits by 2026-10-04. [source]
- oMLX release train: v0.2.6 on 2026-03-08, 0.4.0rc2 on 2026-06-01, 0.6.2 about 2026-08-19, 0.7.0rc1 on 2026-09-24, 0.7.0 on 2026-09-30. 135 tags in about 7.5 months, with dev, rc and post releases between each minor. [source]
- oMLX 0.6.0 (about Aug 2026) added distributed serving, built-in web search, speech-to-text, community intelligence-benchmark publishing to omlx.ai, a prefill that yields GPU time to active decodes (claimed 1.6x to 43x better concurrent decode throughput), and Qwen3.8, Gemma 4 MTP, Ling 3.0 Flash and Jina Reranker v3.5 support. [source]
- oMLX 0.6.2 added a built-in ANE/GPU prefill split tuner and M5 NAX kernels; 0.7.0 rebuilt the memory guard (summary in disk-and-ssd-tiered-kv-cache-servers.md), added M5 tensor-unit prefill work, MoE expert SSD offload and a GPU keep-warm interval (`server.gpu_keep_warm_interval`, up to 5 minutes, claimed 1 to 1.5 s off next TTFT). [source]
- oMLX maintainer Jun Kim announced joining Hugging Face on 2026-09-22 to work on MLX; the post says oMLX stays Apache 2.0, Jun keeps leading it, and oMLX is expected to act as a testbed while upstreaming to mlx-lm and mlx-vlm. [source]
- Rapid-MLX: renamed from the vllm-mlx lineage on 2026-03-13; reached homebrew-core at 0.10.12 (no tap needed); 248 releases, 3.9k stars, 76 contributors by 2026-10-04. [source]
- Rapid-MLX recent cadence: 0.15.2 on 2026-09-25, 0.15.3 on 2026-09-30, 0.15.4 on 2026-10-02, 0.15.5 on 2026-10-03; each release is paired with a `rapid-mac-vX` Desktop tag. [source]
- Rapid-MLX headline claim moved from "4.2x faster than Ollama" (README as quoted in a 2026-09-19 review) to "3.0x Ollama aggregate decode at 8 streams, M2 Pro" (README on 2026-10-02) with the 1.6x whole-batch, 1.5x single-stream and "dense 12B no faster" caveats added. [source]
- Rapid-MLX 0.15.0 turned anonymous telemetry on by default (notice, not opt-in). [source]
- oMLX has 1,033 open and 1,109 closed issues; about 25 open PRs were opened in the final three days, many from a few contributors. [source]
- oMLX issue #4213 (0.7.0, M5 Max 128 GB, balanced tier, Qwen3.8-Flash-Next oQ4e-mtp, about 139k context): prefill rejected as "~78.05 GB peak ... dynamic ceiling is 77.82 GB" while the message itself says 33.73 GB was reclaimable; reporter did not see it in 0.7.0rc1. [source]
- oMLX issue #4228: every MiniMax-M3 chat completion fails in 0.7.0 with `'KVCache' object has no attribute 'keys_and_values'`. [source]
- oMLX issue #4224: on an M3 Ultra 512 GB, ten GPU lockups since 2026-08-27 all occurred while two LLM engines in one oMLX process were busy (`kIOGPUCommandBufferCallbackErrorHang`, about 1.3 per dual-engine hour; none in 35.9 hours with one engine busy); association only, no controlled test. oMLX gives each LLM engine its own thread and MLX stream since 0.4.0. [source]
- oMLX issue #4175: with `hot_cache_max_size 4GB` the cache-hit reconstruct step was 18x to 73x slower than SSD-only (98.9 ms vs 1.77 to 7.24 s, 82.5k-token prefix, M4 Max 36 GB, aggressive tier); suspected link to #1833. The hot tier's stated purpose (PR #58) is SSD write-lifespan protection through write-back batching, not faster reads. [source]
- oMLX issues #4209 and #4214: Lightning MTP and greedy output differ between runs on Qwen3.8 checkpoints since 0.7.0 changes, despite release notes claiming bit-identical verify. [source]
- oMLX issue #4199: the auto-updater cannot finish where the GitHub release CDN throttles per connection (single URLSession task, 3600 s timeout). [source]
- oMLX discussion #523 (unanswered): `brew install omlx` does not install the menu-bar app; the app comes only from the DMG. The Homebrew path is CLI plus brew services. [source]
- Rapid-MLX issue #4108 (0.15.5, M4 Pro 48 GB, qwen3.6-35b-4bit, Claude Code workload): Metal active memory climbs about 25 to 36 GB across 19 requests, evictions do not release it, and the server then returns HTTP 503 "max concurrent requests" with nothing running until restart; owner suspects a 0.15.5 regression from the new prefix-cache HIT path for Claude Code turns. [source]
- Rapid-MLX issue #4092 (owner-filed): PFlash defaults to `always` on some aliases and drops about 80% of the middle of no-tools prompts above roughly 11.5K tokens, silently, with no response signal; tool-carrying agent traffic is skipped so coding agents are unaffected, chat apps and RAG are not. [source]
- Rapid-MLX issue #4109: persisting the prefix cache on shutdown has no free-disk-space check and can fill the disk. [source]
- Rapid-MLX issues #4037 and #4038: the qwen3_coder_xml parser turns integer `type: number` arguments into floats (`10000` to `10000.0`) so Codex rejects the call, and a call to an undeclared tool streams raw `<tool_call>` markup into visible content. [source]
- Rapid-MLX issues #4096 to #4098: the 8 GB starter pick `lfm2.5-2.6b-4bit` fails to load in some HF-cache layouts and the recommendation file carries pre-0.14 speeds; 36 issues open versus 741 closed. [source]
- Rapid-MLX Computer Use (experimental, Desktop) needs Screen Recording and Accessibility permission and the project warns of incomplete tasks and looping planners. [source]
- Who is faster. Rapid-MLX project figures (3.0x Ollama at 8 streams on M2 Pro; 66.1 vs 51.2 vs 47.7 tok/s single-stream on M4 Pro 48 GB) vs oMLX project figures (Qwen3.5-122B-A10B 4-bit on M3 Ultra 512 GB: 56.6 tok/s single, 190.2 at 8x batch = 3.36x; Qwen3-Coder-Next 8-bit 4.14x at 8x). Each vendor measures on different hardware and models, so none of these numbers are comparable across the two sites. No independent third-party head-to-head of the two was found. [source]
- Tool-call parser count: Rapid-MLX README says 27 parser modules; a September 2026 review (danmackinlay) says 17. Likely version drift; unresolved. [source]
- oMLX menu-bar implementation: the project README says native Swift/SwiftUI; the Jimmy Song landscape entry (and the earlier oMLX README text it summarises) says PyObjC. The repository has `apps/omlx-mac`, so Swift is current and PyObjC likely older. [source]
- oMLX Python floor: the README says 3.11-3.13; omlx.ai says "Python 3.10+". Rapid-MLX README says 3.10+. [source]
- macOS floor: oMLX 15.0 (Sequoia) with separate DMGs for macOS 15 and macOS 26/27; Rapid-MLX Desktop says macOS 14 or later. [source]
- Is oMLX's SSD cache unique? oMLX site frames it as its differentiator; danmackinlay notes vllm-mlx has the same cold tier behind `--ssd-cache-dir`, off by default, and the Rapid-MLX README lists a disk-restored radix cache. Existing dossier disk-and-ssd-tiered-kv-cache-servers.md holds the measured differences. [source]
- Independent (non-vendor) benchmark of oMLX vs Rapid-MLX vs mlx-lm on the same Mac, model and settings: not found. [source]
- Whether the oMLX 0.7.0 memory-guard complaints (#4213) are a tier-calibration problem or a bug: unresolved, open. [source]
- Whether Hugging Face funding changes oMLX's release cadence or governance: only a statement of intent exists. [source]
- Rapid-MLX bus factor: owner raullenchai dominates issue authoring and release notes; the contributor list names `@claude` as a contributor, but commit-share data was not retrieved. [source]
- oMLX commit-share by author was not retrievable (contributors graph rendered empty). [source]
- oMLX is an Apache-2.0 LLM server for Apple Silicon with continuous batching, tiered KV cache and a macOS menu-bar app. [source]
- oMLX requires macOS 15.0 or newer, Python 3.11-3.13 and Apple Silicon (M1 to M5). [source]
- oMLX installs three ways: a signed DMG with in-app auto-update, `brew tap jundot/omlx https://github.com/jundot/omlx && brew install jundot/omlx/omlx`, or `pip install -e .` from source. [source]
- The oMLX DMG is built per OS family: `oMLX-0.7.0-macos26-27.dmg` and `oMLX-0.7.0-macos15-sequoia.dmg`. [source]
- The macOS app installs a `~/.omlx/bin/omlx` CLI shim so terminal commands and Apple Shortcuts can control the app-managed server. [source]
- `brew install omlx` does not install the menu-bar app; the app is DMG-only (discussion unanswered). [source]
- MCP support in oMLX is optional: `/opt/homebrew/opt/omlx/libexec/bin/pip install mcp` on Homebrew or `pip install -e ".[mcp]"`. [source]
- Optional oMLX native kernels for GLM-5.2, MiniMax M3 and Qwen3.5 are precompiled in the DMG; GLM-5.2 fused DSA prefill measured 845 tok/s with them vs about 29 without on an M3 Ultra. [source]
- Building the oMLX custom kernels needs full Xcode because Command Line Tools lack the `metal` utility. [source]
- oMLX API endpoints are `/v1/chat/completions`, `/v1/completions`, `/v1/messages`, `/v1/embeddings`, `/v1/rerank` and `/v1/models`. [source]
- oMLX serves LLM, VLM, OCR, embedding and reranker models from one model directory and auto-detects type. [source]
- oMLX models can be given an alias, and a saved profile can be exposed as its own model id `<model>:<profile>` that shares the base engine without extra memory or reload. [source]
- The oMLX admin dashboard has eight UI languages, vendored CDN dependencies for offline use, a Hugging Face model downloader, a one-click benchmark and one-click agent integration setup. [source]
- oMLX native custom kernels, OCR, native video input for Qwen3.5/3.6/3.8 (needs opencv-python-headless) and MiMo V2.6 audio input are documented in the README. [source]
- The oMLX menu-bar app is Swift/SwiftUI, not Electron, with usage history, heatmap and auto-restart. [source]
- oMLX acknowledges vllm-mlx v0.1.0 as its starting point and credits mlx-lm, mlx-vlm, mlx-embeddings, dflash-mlx, MTPLX, mlx-serve and venvstacks. [source]
- oMLX had 22.5k stars, 2k forks, 135 tags and 3,030 commits on 2026-10-04, with the repository created 2026-02-13. [source]
- oMLX v0.7.0 was published 2026-09-30, 0.7.0rc1 on 2026-09-24 and 0.7.0.dev4 on 2026-09-18. [source]
- oMLX release page lists 0.7.0, 0.7.0rc1, 0.7.0.dev1/2/4, 0.6.4, 0.6.3 with rc1-rc3, 0.6.2, 0.6.1, 0.6.0 with rc1 and dev1, 0.5.8 dev1-3, 0.5.7 and 0.4.x/0.2.x lines. [source]
- oMLX 0.6.0 introduced experimental distributed serving by tensor or pipeline parallelism; Qwen3.6-27B reached 28.6 tok/s across two Macs versus 16.1 on one. [source]
- oMLX 0.6.0 added heterogeneous Metal and CUDA model pools (Apple Silicon and NVIDIA workers in one logical pool). [source]
- oMLX distributed inference is a source-build preview that is off by default and enabled under Global Settings > Advanced > Distributed Inference; it uses SSH trust-on-first-use pairing and MLX's official launcher for Ring, JACCL and JACCL Ring. [source]
- oMLX 0.6.0 claims prefill now yields GPU time to active decodes, improving decode throughput during concurrent prefill 1.6x to 43x. [source]
- oMLX 0.6.2 added a built-in ANE/GPU split tuner that recommends GPU-only when the best ANE split is under 1% faster. [source]
- oMLX 0.7.0 added a per-model INT8-activation prefill mode for Q4/Q5 Qwen on M5 (off by default; 32K prefill 615.2 to 826.7 tok/s on M5 Max, +34.4%, decode 24.6 to 22.2) which activation quantization can change output. [source]
- oMLX 0.7.0 added experimental MoE expert SSD offload keeping 12.5% to 75% of experts resident for DeepSeek V4.1 Flash, Qwen3.8-Flash-Next, Gemma 4 MoE and OLMoE; MTP and DFlash must be disabled with it. [source]
- oMLX 0.7.0 added `server.gpu_keep_warm_interval` keeping the GPU out of idle power state up to 5 minutes, removing about 1 to 1.5 s from next TTFT in MiMo tests. [source]
- oMLX 0.7.0 lets `max_concurrent_requests` change live without unloading models. [source]
- Many oMLX 0.7.0 performance changes come from external contributors, notably @jonathan308 and @jerryfane; commits are also co-authored with Claude. [source]
- omlx.ai markets oMLX with 22,400+ stars, "5 seconds, not 90" TTFT, and M3 Ultra 512 GB tables (Qwen3.5-122B-A10B-4bit 56.6 tok/s single, 190.2 tok/s at 8x batch; Qwen3-Coder-Next-8bit 58.7 and 243.3; GLM-5-4bit 16.7 and 60.3). [source]
- The omlx.ai Qwen3.5-122B-A10B table shows prompt throughput of 768 tok/s at 1k and 765 at 32k context with peak memory 65.5 to 73 GB. [source]
- omlx.ai says oMLX reads the standard Hugging Face cache and an LM Studio folder, so models need not be re-downloaded. [source]
- Hugging Face announced on 2026-09-22 that oMLX creator Jun Kim joined it; oMLX stays Apache 2.0 and Jun keeps leading it; the stated goal is a funded project with oMLX as a testbed that upstreams to mlx-lm and mlx-vlm. [source]
- danmackinlay describes oMLX's pre-September-2026 caveats as "bus factor 1, MLX-only" and says it is still MLX-only. [source]
- oMLX has 1,033 open and 1,109 closed GitHub issues, with about 25 PRs opened in the last three days before 2026-10-04. [source]
- oMLX issue #4213 reports the 0.7.0 memory guard rejecting a prefill at ~78.05 GB peak against a 77.82 GB dynamic ceiling while 33.73 GB was reclaimable (M5 Max 128 GB, balanced tier, about 139k context). [source]
- oMLX issue #4228 reports MiniMax-M3 failing in 0.7.0 with `'KVCache' object has no attribute 'keys_and_values'`. [source]
- oMLX issue #4224 reports ten GPU hangs on an M3 Ultra only while two LLM engines were busy in one oMLX process, none in 35.9 hours with one engine busy. [source]
- oMLX issue #4175 measured the hot RAM cache making cache-hit reconstruct 18x to 73x slower than SSD-only on M4 Max 36 GB. [source]
- oMLX PR #58 stated the hot cache goal as SSD write-lifespan protection through write-back batching, not faster reads. [source]
- Rapid-MLX describes itself as an Apache-2.0 OpenAI- and Anthropic-compatible inference server and Mac app on MLX focused on reliable tool calling for coding agents. [source]
- Rapid-MLX installs by `brew install rapid-mlx` (homebrew-core bottle since 0.10.12), `curl -fsSL https://rapidmlx.com/install.sh | bash`, `uv tool install rapid-mlx@latest`, `pip install rapid-mlx` (Python 3.10+) or a signed Desktop DMG. [source]
- The Rapid-MLX guided installer detects RAM and picks lfm2.5-1b-4bit below 16 GB and qwen3.5-4b-4bit at 16 GB or above for the first chat. [source]
- Rapid-MLX tier picks by RAM: `rapid-mlx recipe` prints a smart and a fast model per tier, for example 18 GB gives qwen3.5-9b-4bit (about 36 tok/s) and qwen3.5-4b-4bit (about 61 tok/s). [source]
- Rapid-MLX Desktop is free, 134 MB, requires macOS 14+ and Apple Silicon, and includes local dictation with Whisper, Qwen3-ASR, SenseVoice or Parakeet. [source]
- The Rapid-MLX base install is text-only at about 460 MB; vision, audio, video, image, embeddings and DFlash are opt-in extras. [source]
- Rapid-MLX serves image generation (flux2-klein-4b, qwen-image-edit), video (Wan 2.1/2.2, CogVideoX-Fun, LTX-2.3) and 44 audio aliases behind OpenAI-style endpoints. [source]
- Rapid-MLX exposes `/v1/responses` for Codex CLI and `/v1/messages` for Claude Code, and `rapid-mlx launch claude-code` patches `~/.claude/settings.json`. [source]
- Rapid-MLX Cursor support is intentionally absent for localhost because Cursor routes BYOK requests through its own servers; only a public HTTPS tunnel with `RAPID_MLX_API_KEY` works. [source]
- Rapid-MLX Tier-1 agents (Claude Code, Codex CLI, Hermes, Aider, DeepSeek Harness) are gated by a release-blocking smoke test that fixes a real bug with a local 35B model. [source]
- Rapid-MLX telemetry is on by default since 0.15.0 (notice, not opt-in); disable with `rapid-mlx telemetry off`, `RAPID_MLX_TELEMETRY=0` or `DO_NOT_TRACK=1`; prompts, completions, paths and IPs are not collected. [source]
- Rapid-MLX mirrors Ollama verbs (run, pull, rm, ps, models) but `rapid-mlx serve <model>` serves one named model on port 8000 rather than a model-agnostic daemon. [source]
- Rapid-MLX's own migration guide states Ollama still wins on cross-platform support and model-library size. [source]
- Rapid-MLX had 3.9k stars, 424 forks, 248 releases, 76 contributors, 75.5% Python and 22.2% Swift on 2026-10-02. [source]
- Rapid-MLX shipped 0.15.2 on 2026-09-25, 0.15.3 on 2026-09-30, 0.15.4 on 2026-10-02 and 0.15.5 on 2026-10-03. [source]
- Rapid-MLX 0.15.3 added supervised experimental Computer Use in Desktop with an authenticated `/v1/cua` API, blocking credentials, payment and commerce actions. [source]
- Rapid-MLX 0.15.4 added narrowly qualified experimental profiles `qwen3.8-27b-tensorfold` (DFlash2 drafter) and `glm5.3-flash-tensorfold` (embedded MTP head, 256 GB Macs) that admit one request at a time and fail closed. [source]
- Rapid-MLX 0.15.5 added background input for Computer Use, one host sync per accepted MTP draft cycle, earlier checks and cancel-safe import for bring-your-own models, and a universal context-length override. [source]
- Rapid-MLX README large-model table, 256 GB M3 Ultra, 4-bit, median of three after clearing prefix cache: qwen3.8-27b TTFT 24.66 s, prefill 330.8 tok/s, decode 43.4; qwen3.8-flash-next TTFT 9.40 s, prefill 867.9, decode 23.0; glm5.3-flash TTFT 22.78 s, decode 27.8. [source]
- Rapid-MLX 0.13.4 verified MTP raised Qwen3.8-27B decode from 24.58 to 43.38 tok/s at 8K (1.77x) and 16.52 to 38.66 at 32K (2.34x) on the same Mac. [source]
- Rapid-MLX README headline was "4.2x faster than Ollama" in a review dated 2026-09-19 and "3.0x at 8 streams" with caveats by 2026-10-02. [source]
- The promptquorum review lists 194 text-model aliases, 3,773 stars and 416 forks on 2026-09-18 and did not independently verify any speed claim. [source]
- rapidmlx.com advertises 231 to 248 models, 6,222 tests and "0.08 s" cached TTFT. [source]
- Rapid-MLX issue #4108 reports a 0.15.5 wedge: Metal active memory not released after evictions on a Claude Code workload, then permanent HTTP 503 with an empty cache until restart. [source]
- Rapid-MLX issue #4092 reports PFlash `always` silently dropping about 80% of the middle of no-tools prompts above roughly 11.5K tokens. [source]
- Rapid-MLX had 36 open and 741 closed issues on 2026-10-04, many authored by the owner. [source]
- Rapid-MLX issues #4037 and #4038 report parser bugs: integer arguments turned into floats and undeclared-tool markup leaking into content. [source]
- danmackinlay judges that oMLX's SSD cache is less of a differentiator than first thought because vllm-mlx has a cold tier behind `--ssd-cache-dir`, off by default. [source]
- danmackinlay says Rapid-MLX forked away vllm-mlx's multi-model registry (185 lines vs 1195), and reports the multi-model logic in vllm-mlx has bugs (#627, #712). [source]
- waybarrios/vllm-mlx, the common ancestor, was still active on 2026-10-03 with 664 commits and 15 tags. [source]
- oMLX and Rapid-MLX are both listed by the Rapid-MLX README comparison table as continuous-batching MLX servers with Anthropic Messages support; LM Studio's MLX engine batches since 0.4.2 and mlx-lm.server gets no batching with a quantized KV cache. [source]
- Which to choose (by workload): oMLX for several models resident at once, a menu-bar plus web admin, persistent SSD cache across restarts, embeddings and rerankers on one endpoint, and mixed-precision oQ models; Rapid-MLX for one model served headless through CLI or Homebrew, Ollama-style verbs, wire-verified agent setups, and multimodal (image, video, audio) from one binary. [source]
- Which to choose (by risk): both ship several releases a week and carry fresh regressions in the latest version (oMLX 0.7.0 memory guard and MiniMax-M3 failures; Rapid-MLX 0.15.5 Metal wedge), so pin a version and hold the previous one for rollback. [source]
- Which to choose (by privacy): Rapid-MLX telemetry is default-on with a notice; oMLX's README documents local-only usage history and no telemetry, which this research did not find contradicted. [source]
- Maintenance status: oMLX has one named lead maintainer (now Hugging Face funded) plus a large external contributor base; Rapid-MLX has one dominant owner with 76 contributors; both have daily commits as of 2026-10-03. [source]
Children
- No children recorded.