<!-- llms-explorer concept facts · https://llms-explorer.com/tree/mlx-metal-queue-contention-between-two-processes/ · pack 2026-10-05 · ~282 tokens -->

# MLX Metal queue contention between two processes on one GPU

> MLX calls `newCommandQueue()` once per stream encoder and attaches the residency sets to that queue, so processes never share a queue

Parent: [Mac local LLMs: GPU stability and kernel panics](https://llms-explorer.com/tree/mac-local-llms-gpu-stability-and-kernel-panics/) · 1 facets · 3 facts · page: https://llms-explorer.com/tree/mlx-metal-queue-contention-between-two-processes/

## Facts

- MLX calls `newCommandQueue()` once per stream encoder and attaches the residency sets to that queue, so processes never share a queue — [source](https://raw.githubusercontent.com/ml-explore/mlx/main/mlx/backend/metal/device.cpp)
- MLX command buffers use unretained references; buffers are kept alive by MLX through completion handlers — [source](https://raw.githubusercontent.com/ml-explore/mlx/main/mlx/backend/metal/eval.cpp)
- Default per-command-buffer limits are 20 ops/40 MB (phone), 40/40 (base and Pro), 50/50 (Max and Ultra), overridable by environment variables — [source](https://raw.githubusercontent.com/ml-explore/mlx/main/mlx/backend/metal/device.cpp)
