<!-- llms-explorer concept facts · https://llms-explorer.com/tree/omlx-engine-serialization-a-b-omlx-serialize-eng/ · pack 2026-10-05 · ~779 tokens -->

# oMLX engine serialization A/B (OMLX_SERIALIZE_ENGINE_GPU) results

> 3 Oct 2026: issue 4224 is filed with the draft patch and an offer to open a PR.

Parent: [Mac local LLMs: GPU stability and kernel panics](https://llms-explorer.com/tree/mac-local-llms-gpu-stability-and-kernel-panics/) · 1 facets · 13 facts · page: https://llms-explorer.com/tree/omlx-engine-serialization-a-b-omlx-serialize-eng/

## Facts

- 3 Oct 2026: issue 4224 is filed with the draft patch and an offer to open a PR. — source: `asserted`
- 4 Oct 2026 (12:39 UTC fetch): the issue is still open with no comments. — source: `asserted`
- Without a result, the claim that serialization removes the Hang class stays a hypothesis. — source: `asserted`
- Two earlier oMLX issues describe other cross-thread and concurrency failures that serialization would not obviously address (see Claims). — source: `asserted`
- None. — source: `asserted`
- The existing dossiers already record issue 4224's A/B design and arithmetic, so no new design claim is added here. — source: `asserted`
- GitHub's REST record for issue 4224, fetched 4 Oct 2026 at 12:39 UTC, shows it open with 0 comments, created 2026-10-03T06:20:37Z and last updated 2026-10-03T06:56:31Z, and its comments endpoint returns an empty list. — [source](https://api.github.com/repos/jundot/omlx/issues/4224)
- Issue 4224's closing list still names "the result of the pure-MLX reproduction once it has run" and "the serialized-engines A/B described in (a), in a maintenance window" as items the reporter can provide later. — [source](https://github.com/jundot/omlx/issues/4224)
- A search of the oMLX repository for "serialize engine gpu" returns 38 issues and pull requests; its top hit is issue 4224, and the five pull requests among the first 20 results (PRs 757, 3991, 3982, 3990, 2756) are unrelated to cross-engine serialization. — [source](https://api.github.com/search/issues?q=repo:jundot/omlx+serialize+engine+gpu&per_page=20)
- oMLX issue 2624 (13 Aug 2026, oMLX 0.5.4, DeepSeek-V4-Flash-0731-MXFP4 at 155.9 GB) records the engine accepting HTTP requests for 6 hours but never producing a body, with every thread idle and no `mlx::core::eval_impl` in 21 stack captures, and recovery after a clean `launchctl bootout`. — [source](https://github.com/jundot/omlx/issues/2624)
- oMLX issue 1558 (30 May 2026) reports a fatal SIGABRT on the background SSD cache-saver thread, "There is no Stream(gpu, 3) in current thread", because MLX streams are thread-local and lazy KV arrays scheduled under the generation stream were evaluated from a ThreadPoolExecutor worker. — [source](https://github.com/jundot/omlx/issues/1558)
- oMLX issue 1970 (22 Jun 2026, linked to PR 3721) reports that when memory fits one model, a concurrent request for another model is rejected with a 507 or blocked 20-37 s because `EnginePool.get_engine()` holds an exclusive asyncio lock through the whole load sequence, stalling the already-loaded model too. — [source](https://github.com/jundot/omlx/issues/1970)
- No cached page contains a result for the `OMLX_SERIALIZE_ENGINE_GPU` A/B. — source: `asserted`
