oMLX remaining ad-hoc mx.clear_cache sites needing stream analysis
Parent: Mac local LLMs: oMLX, Rapid-MLX and related internals · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
On main, `engine/base.py` `_finish_activity` still runs `(mx.synchronize(), mx.clear_cache())` on `get_mlx_executor()` after every non-streaming request, with a comment that it always clears because gating on `_active_count == 0` grew the pool without bound (issue 684).
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- On main, `engine/base.py` `_finish_activity` still runs `(mx.synchronize(), mx.clear_cache())` on `get_mlx_executor()` after every non-streaming request, with a comment that it always clears because gating on `_active_count == 0` grew the pool without bound (issue 684). [source]
- The same bare pair remains in `engine/reranker.py`, `engine/tts.py`, `engine/stt.py` and `engine/sts.py` (one site each, via `get_mlx_executor()`), `engine/dflash.py` (two sites, in model eviction and its settle loop) and `admin/oq_manager.py` (two sites inside retry loops of 3). [source]
- `engine_pool.py` has about 18 sites using the same `(mx.synchronize(), mx.clear_cache())` lambda on the MLX executor for load, unload, emergency and settle paths; none passes a stream or takes `_mx_buffer_access_lock`. [source]
- `engine/vlm.py` still calls bare `mx.synchronize(); mx.clear_cache()` after vision encoding and in the `finally` of the diffusion stream, on the calling thread. [source]
- `engine/embedding.py` calls `mx.clear_cache()` after `mx.synchronize()` only when `fairness.should_clear_cache()` is true, inside the batch worker's `finally`. [source]
- `oq.py` has about 37 `mx.clear_cache()` mentions: most follow `mx.eval` and `del` in per-tensor or per-layer quantization loops, many with a preceding `mx.synchronize()`, and one set replaces `mx.clear_cache` with a no-op in a test or fake-array harness. [source]
- In `patches/glm_moe_dsa/generate_patch.py` the decode-side clear every 512 steps is converted to `_sync_and_clear_cache(self._stream)`, but the patched prefill loop still calls bare `mx.clear_cache()` after `mx.eval` twice (lines near 164 and 172). [source]
- `speculative/vlm_mtp.py` now uses `_sync_and_clear_cache(_vlm_generation_stream)` at three sites, matching PR 2413. [source]
- `scheduler.py` routes every clear through `Scheduler._clear_cache`, which calls `_sync_and_clear_cache(self._stream)`, and the deferred, periodic, per-chunk, abort and pressure-reclaim clears all use it. [source]
- Scheduler call sites include a clear every 1024 decode tokens and a deferred clear after request end (`_schedule_deferred_metal_clear`), whose comment says an immediate clear after completion races. [source]
- Scheduler wraps its error-recovery clear in try/except because `mx.synchronize()` or `mx.clear_cache()` can raise a C++ exception that would SIGABRT (issue 435). [source]
Corrections and disagreements
- CONTRADICTS: omlx-metal-cache-clears-wait-for-the-owning-stre.md open question implies the sites may have been converted; on main only the `vlm_mtp.py` and GLM decode sites are, and the engine, engine_pool, dflash, vlm and oq sites are not. [source]
Children
- No children recorded.