MLX Metal queue contention between two processes on one GPU
Parent: Mac local LLMs: GPU stability and kernel panics · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
MLX calls `newCommandQueue()` once per stream encoder and attaches the residency sets to that queue, so processes never share a queue
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- MLX calls `newCommandQueue()` once per stream encoder and attaches the residency sets to that queue, so processes never share a queue [source]
- MLX command buffers use unretained references; buffers are kept alive by MLX through completion handlers [source]
- Default per-command-buffer limits are 20 ops/40 MB (phone), 40/40 (base and Pro), 50/50 (Max and Ultra), overridable by environment variables [source]
Children
- No children recorded.