MLX batched decode scheduler crashes
Parent: Mac local LLMs: MLX kernels, numerics and internals · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Issue 883 was reproduced on a single-stream agentic client (growing context), not on concurrent batching, so it should not be filed under batched-decode crashes.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Issue 883 was reproduced on a single-stream agentic client (growing context), not on concurrent batching, so it should not be filed under batched-decode crashes. [source]
- No public issue was found that reproduces a hard crash specific to BatchGenerator with a hybrid (GatedDeltaNet or Mamba) cache under concurrency. Reddit threads (r/LocalLLaMA 1rq22mq, LM Studio MLX reboot) could not be fetched (site unsupported by the fetcher). [source]
- Issue 883 (opened 2026-02-11): mlx-lm 0.30.6 on a Mac Studio M3 Ultra 96 GB with Qwen3-Coder-30B-A3B 8-bit panicked with "completeMemory() prepare count underflow" @IOGPUMemory.cpp:550 at about 58k tokens of agentic context, with 80.14 GB wired, 0.01 GB free and memoryPressure false. [source]
- The issue 883 reporter's memory breakdown at crash: weights about 32 GB, KV about 10-20 GB, MoE activations about 3-5 GB, Metal buffer cache about 2-4 GB, MLX process resident 83.23 GB. [source]
- The reporter proposed a server `--memory-limit` flag and a lower default wired limit than the roughly 75% of RAM that `mlx_lm.server` wires at startup. [source]
- Hannecke's workaround is a wrapper that calls `mx.metal.set_memory_limit(48 * 1024**3)` before `mlx_lm.server.main()` so MLX raises a Python exception instead of driving the GPU driver into a panic; he says Ollama never panicked because it fixes context size at startup. [source]
- Hannecke's article corrected its own model description: Gemma 4 26B is the MoE variant (gemma-4-26B-A4B, about 3.8B active), the dense one is 31B; the panic diagnosis was unchanged. [source]
Children
- No children recorded.