oMLX store-cache admission gate and 60 s admission stall error
Parent: Mac local LLMs: oMLX, Rapid-MLX and related internals · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
`_StoreCacheGate` is guarded by a lock and exposes `note_submitted`, `note_done`, `set_cap`, `cap`, `in_flight` and `has_capacity`; `has_capacity` is `in_flight < cap` and the cap is at least 1.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- `_StoreCacheGate` is guarded by a lock and exposes `note_submitted`, `note_done`, `set_cap`, `cap`, `in_flight` and `has_capacity`; `has_capacity` is `in_flight < cap` and the cap is at least 1. [source]
- The store-cache executor is a `ThreadPoolExecutor` with `max_workers=1` and thread name prefix `omlx-store-cache`, so stores are serialized and never race on the paged SSD index. [source]
- `ProcessMemoryEnforcer` adjusts the gate cap on every poll through `adjust_store_cache_cap`: one step up toward `max_num_seqs` at ok pressure, one step down to a floor of 1 at soft or hard pressure. [source]
- `note_submitted()` is called before `executor.submit`, `note_done()` is called if the submit raises, and otherwise `note_done()` runs in `_drain_pending_async_removes` after the request cache references are released, not when the write finishes. [source]
- `_schedule_waiting` defers admission when the gate has no capacity or when `len(_pending_async_removes) >= gate.cap`, and logs `Admission deferred: store-cache pipeline full (in_flight=%d pending_cleanups=%d cap=%d)` at debug level. [source]
- If the gate is full and process memory is already at or above the limit (or admission is paused), the scheduler uses `_memory_admission_stall_output("store_cache_backpressure")`, whose error code is `memory_admission_stalled`, instead of the store-cache timer. [source]
- `_MEMORY_ADMISSION_STALL_TIMEOUT_S` and `_STORE_CACHE_ADMISSION_STALL_TIMEOUT_S` are both 60.0 seconds. [source]
- The stalled request's output has `finish_reason="error"`, `error_code="store_cache_admission_stalled"` and `error_metadata` with `request_id`, `reason`, `stalled_seconds`, `store_cache_in_flight`, `pending_store_cleanups` and `store_cache_cap`. [source]
- The error text reads `Request could not be admitted because store-cache cleanup stayed full for <n>s (<reason>). The previous response cache is still being persisted; retry after the cache writer drains or reduce cache/write pressure.` [source]
- The scheduler logs `Store-cache admission stalled for %s: %s (in_flight=%d pending_cleanups=%d cap=%d)` at warning level when it fails a request this way. [source]
- The timer state is `_store_cache_admission_blocked_request_id` and `_store_cache_admission_blocked_since`; both reset when the blocked request is admitted, a different request becomes head, or the waiting queue empties. [source]
- After a drain completes, the scheduler schedules another deferred Metal clear, because the earlier deferred clear may have fired while the async store worker still owned the extracted KV buffers. [source]
Children
- No children recorded.