osaurus vmlx-swift-lm runtime and vmlxctl
Parent: Mac local LLMs: oMLX, Rapid-MLX and related internals · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
vmlx-swift-lm adds, on top of upstream mlx-swift-lm: continuous batching (`BatchEngine`) with per-slot KV isolation and SSM-state merge for hybrid models; a three-tier cache; TurboQuant KV compression; speculative decoding; JANG mixed precision; and MoE or hybrid dispatch reduction.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- vmlx-swift-lm adds, on top of upstream mlx-swift-lm: continuous batching (`BatchEngine`) with per-slot KV isolation and SSM-state merge for hybrid models; a three-tier cache; TurboQuant KV compression; speculative decoding; JANG mixed precision; and MoE or hybrid dispatch reduction. [source]
- Cache tiers: L1 is an in-memory paged block pool (64-token blocks by default, SHA-256 chain hashing); L2 is SQLite plus safetensors keyed by prompt hash and survives process restarts; a third SSM-companion tier holds Mamba or similar state per layer and is disk-paged for hybrids when L2 is on. The cache key combines `modelKey`, a `cacheScopeSalt` (reasoning mode plus media) and, if set, the MoE top-k override, so requests at different reasoning modes never alias. [source]
- TurboQuant KV compression stores the cache at 3 key bits and 3 value bits in the documented example, claims 4.7 to 5.0 times smaller cache using a randomized Hadamard rotation, a Lloyd-Max codebook and (keys only) QJL residual correction, and skips non-KV layers (Mamba, rotating sliding-window caches, ZayaCCA). [source]
- Speculative decoding has three strategies selected by `GenerateParameters.draftStrategy`: classic autoregressive drafter, DFlash and DDTree. At temperature 0 the output is byte-identical to plain greedy decode; drafters load from `z-lab/<model>-DFlash` snapshots. [source]
- `VMLX_MOE_TOPK_OVERRIDE=<n>` lowers the routed-expert top-k at load time for supported MoE families, never raises it, leaves top-1 families (ZAYA) alone, and is folded into the cache key; a legacy `VMLINUX_` prefix is accepted. [source]
- Osaurus serves one local endpoint set on 127.0.0.1:1337: OpenAI `/v1/chat/completions`, Anthropic `/anthropic/v1/messages` and Ollama `/api/chat`. Its CLI is `osaurus serve --port 1337` (add `--expose` for the LAN), `osaurus ui`, `osaurus status`, `osaurus stop`; `osaurus mcp` is a stdio MCP bridge to the local server. [source]
- When an agent delegates to a different local model, Osaurus does a "single-residency handoff": it unloads the chat model, runs the job, then reloads, so two large models are never resident together; same-model children reuse the resident model, and an experimental RAM-safe coexistence mode can keep both when admitted. [source]
- Osaurus stores models in `~/MLXModels` (override `OSU_MODELS_DIR`) and ships two DMGs per release: a standard build of about 70 MB and a full build of about 4 GB with a bundled model (Raptor 0.6 4B JANG). Install is `brew install --cask osaurus`. [source]
- 2025-09: the Osaurus project was published by Dinoki Labs (repo then `dinoki-ai/osaurus`), requiring macOS 15.5+, with a default port of 8080 and OpenAI and Ollama endpoints only. [source]
- By 2026-10 the project lives at `osaurus-ai/osaurus` (3,898 commits, latest 2026-10-02), adds Anthropic and MCP endpoints, a sandbox, agents and memory, and serves on 1337. [source]
- 2026-05: the vmlx-swift-lm fork reached 715 commits; its last commit on the cached page is 2026-05-15, so the engine fork had been quiet for about 4.5 months by 2026-10-04 while the Osaurus app kept shipping. [source]
- Osaurus commit messages show it pins vmlx-swift-lm and "repins" it for fixes, so engine fixes reach users only through Osaurus releases. [source]
- Gemma 4 audio is not implemented: `Gemma4.prepare` throws `VLMError.processing` on audio input although Gemma 4 supports audio natively. [source]
- Speculative decoding needs trimmable caches and does not work after a rotating (sliding-window) cache wraps. [source]
- Raw Hugging Face checkpoints with fused `gate_up_proj` must be converted first; JANG and pre-converted mlx-community models load. [source]
- Known regressions recorded in its own table: Ling-2.6-flash MXFP4 decodes at 10.8 tok/s against 57.5 tok/s for its JANGTQ2 build; Hy3-preview is "functional, speed-open" at 14.9 tok/s; MiniMax M2.7 shows TurboQuant cross-slot drift at batch size 2, which gates its multi-slot production claim. [source]
- The README does not mention `vmlxctl`; its flags are documented only in the jangq runtime README, so treat `vmlxctl` as an unversioned internal tool. [source]
- The pip-installed vMLX (`vmlx serve`) is a different program from `vmlxctl`; instructions for one do not apply to the other. [source]
- Speed claims are vendor-measured and use unlike hardware. The fork's table shows Qwen 3.5-35B-A3B at 103 tok/s on an M4 Max 128 GB against Python `mlx_lm` at 94 tok/s, but the README states the Python baselines ran on an M3 Ultra 256 GB with about 1.5 times the memory bandwidth, and that the multi-turn Python decode (122 tok/s turn 1) is faster than the fork's (106). The README's own conclusion is that the fork trails Python by about 10% on long-context decode and leads on cold-start TTFT. Side by side: "+151% vs upstream Swift" and "-10% vs Python on long context". [source]
- The vmlx.net comparison table claims SSD prefix cache, native Anthropic API and a built-in GGUF-to-MLX converter as features LM Studio and Ollama lack, but publishes no benchmark: its TTFT table says "benchmark pending". The fork README publishes numbers but compares against omlx and LM Studio builds it ran itself. Neither is independent. [source]
- Does the current Osaurus release still pin a vmlx-swift-lm revision with Gemma 4 and JangPress support? The cached pages do not state the pinned revision. [source]
- Does `vmlxctl serve` expose the same endpoints as Osaurus, and is it used by Osaurus or only by the jangq scripts? Not documented in the pages read. [source]
- vmlx-swift-lm is "mlx-swift-lm (Osaurus fork)", maintained by Osaurus, a fork of ml-explore/mlx-swift-lm, and "the inference engine for Osaurus", with 715 commits and a latest commit dated 2026-05-15. [source]
- On an M4 Max 128 GB the fork's single-stream decode is Qwen 3.5-35B-A3B 103 tok/s (upstream Swift 41, Python mlx_lm 94 measured on M3 Ultra 256 GB), Gemma 4 26B-A4B 87 (upstream 27), Gemma 4 E2B 121 (upstream 120, Python 128). [source]
- The README attributes the gains to cutting graph-level `AsType` ops by 71 to 95% on MoE, MLA and hybrid-SSM families, and says dense models like Gemma 4 E2B were already near-optimal upstream. [source]
- In the Gemma 4 26B-A4B multi-turn table, decode on turn 1 is 98.2 tok/s for the fork, 71.6 for Python mlx_lm 0.31.2 and 77.7 for omlx 0.3.2. [source]
- On Llama 3.2 1B 4-bit all four compared runtimes land within 8% of each other, which the README uses to show the gains are specific to MoE, MLA and hybrid models. [source]
- Throughput scales near-linearly to about 6 to 8 concurrent slots for routed-MoE bundles, and the stress harness ran 199 of 199 mixed concurrent requests through the BatchEngine and cache stack. [source]
- The L1 cache uses 64-token blocks with SHA-256 chain hashing, the L2 cache is SQLite plus safetensors per prompt hash, and the cache key includes `modelKey`, `cacheScopeSalt` and the optional MoE top-k suffix. [source]
- TurboQuant compresses the KV cache 4.7 to 5.0 times with 3-bit keys and values in the example and skips MambaCache, RotatingKVCache and ZayaCCA. [source]
- `VMLX_MOE_TOPK_OVERRIDE` is lower-only, per-load, folded into the cache `modelKey`, and covers MiniMax, Hy3, NemotronH, Qwen3 through 3.6, BailingHybrid and Gemma 4. [source]
- Known limitations include no Gemma 4 audio encoder (`Gemma4.prepare` throws `VLMError.processing` on audio), speculative decoding incompatible with a wrapped RotatingKVCache, and raw HF checkpoints with fused `gate_up_proj` needing conversion. [source]
- Ling-2.6-flash MXFP4 decodes at 10.8 tok/s (labelled a known regression) against 57.5 tok/s for JANGTQ2, and MiniMax M2.7 shows a TurboQuant cross-slot drift at B=2. [source]
- The vmlx-swift-lm README page contains no mention of `vmlxctl`. [source]
- The jangq runtime README says `VMLXCTL` is the path to "the JangPress-aware vmlxctl" to be built from `osaurus-ai/vmlx-swift-lm`, at `.build/arm64-apple-macosx/release/vmlxctl`. [source]
- Osaurus is MIT licensed, installed with `brew install --cask osaurus`, offers a roughly 70 MB standard DMG and a roughly 4 GB full DMG with Raptor 0.6, and stores models in `~/MLXModels`. [source]
- Osaurus serves OpenAI, Anthropic and Ollama endpoints at `http://127.0.0.1:1337` and its CLI has `serve --port 1337`, `serve --expose`, `ui`, `status` and `stop`. [source]
- On a different-local-model delegation Osaurus unloads the chat model, runs the job, then reloads it (single-residency handoff), with an experimental RAM-safe coexistence mode. [source]
- Osaurus on macOS 26 and later runs agent sandboxes in a Linux VM and falls back to a `sandbox-exec` Seatbelt sandbox on earlier macOS. [source]
- A September 2025 article describes Osaurus by Dinoki Labs requiring macOS 15.5+, an Apple Silicon Mac, a default port of 8080 and OpenAI and Ollama endpoints. [source]
- vMLX at vmlx.net is a separate MIT project (`jjang-ai/vmlx`), installed with `pip install vmlx` and `vmlx serve <model>`, serving OpenAI, Anthropic and Ollama APIs on 127.0.0.1:8000 for macOS 14 and later on M1 through M5. [source]
- vmlx.net publishes no benchmark: its time-to-first-token section reads "benchmark pending" with numbers "deliberately absent". [source]
Children
- No children recorded.