<!-- llms-explorer concept facts · https://llms-explorer.com/tree/mlx-batched-decode-scheduler-crashes/ · pack 2026-10-05 · ~623 tokens -->

# MLX batched decode scheduler crashes

> Issue 883 was reproduced on a single-stream agentic client (growing context), not on concurrent batching, so it should not be filed under batched-decode crashes.

Parent: [Mac local LLMs: MLX kernels, numerics and internals](https://llms-explorer.com/tree/mac-local-llms-mlx-kernels-numerics-and-internals/) · 1 facets · 7 facts · page: https://llms-explorer.com/tree/mlx-batched-decode-scheduler-crashes/

## Facts

- Issue 883 was reproduced on a single-stream agentic client (growing context), not on concurrent batching, so it should not be filed under batched-decode crashes. — source: `asserted`
- No public issue was found that reproduces a hard crash specific to BatchGenerator with a hybrid (GatedDeltaNet or Mamba) cache under concurrency. Reddit threads (r/LocalLLaMA 1rq22mq, LM Studio MLX reboot) could not be fetched (site unsupported by the fetcher). — source: `asserted`
- Issue 883 (opened 2026-02-11): mlx-lm 0.30.6 on a Mac Studio M3 Ultra 96 GB with Qwen3-Coder-30B-A3B 8-bit panicked with "completeMemory() prepare count underflow" @IOGPUMemory.cpp:550 at about 58k tokens of agentic context, with 80.14 GB wired, 0.01 GB free and memoryPressure false. — [source](https://github.com/ml-explore/mlx-lm/issues/883)
- The issue 883 reporter's memory breakdown at crash: weights about 32 GB, KV about 10-20 GB, MoE activations about 3-5 GB, Metal buffer cache about 2-4 GB, MLX process resident 83.23 GB. — [source](https://github.com/ml-explore/mlx-lm/issues/883)
- The reporter proposed a server `--memory-limit` flag and a lower default wired limit than the roughly 75% of RAM that `mlx_lm.server` wires at startup. — [source](https://github.com/ml-explore/mlx-lm/issues/883)
- Hannecke's workaround is a wrapper that calls `mx.metal.set_memory_limit(48 * 1024**3)` before `mlx_lm.server.main()` so MLX raises a Python exception instead of driving the GPU driver into a panic; he says Ollama never panicked because it fixes context size at startup. — [source](https://medium.com/@michael.hannecke/how-my-local-coding-agent-crashed-my-mac-and-what-i-learned-about-mlx-memory-management-e0cbad01553c)
- Hannecke's article corrected its own model description: Gemma 4 26B is the MoE variant (gemma-4-26B-A4B, about 3.8B active), the dense one is 31B; the panic diagnosis was unchanged. — [source](https://medium.com/@michael.hannecke/how-my-local-coding-agent-crashed-my-mac-and-what-i-learned-about-mlx-memory-management-e0cbad01553c)
