<!-- llms-explorer concept facts · https://llms-explorer.com/tree/omlx-paged-ssd-cache-auto-limit-policy/ · pack 2026-10-05 · ~1448 tokens -->

# oMLX paged SSD cache auto limit policy

> `CacheSettings.ssd_cache_max_size` defaults to `"auto"` with the comment `"auto" reserves half of available cache space`.

Parent: [Mac local LLMs: Prompt cache and persistent KV](https://llms-explorer.com/tree/mac-local-llms-prompt-cache-and-persistent-kv/) · 1 facets · 21 facts · page: https://llms-explorer.com/tree/omlx-paged-ssd-cache-auto-limit-policy/

## Facts

- `CacheSettings.ssd_cache_max_size` defaults to `"auto"` with the comment `"auto" reserves half of available cache space`. — [source](https://raw.githubusercontent.com/jundot/omlx/main/omlx/settings.py)
- `get_auto_ssd_cache_size` returns `(free_bytes + cache_bytes) // 2`, where `cache_bytes` sums `*.safetensors` files under the directories `0123456789abcdef` and `_gdn_sidecars`, skipping symlinks. — [source](https://raw.githubusercontent.com/jundot/omlx/main/omlx/settings.py)
- `get_ssd_cache_max_size_bytes` returns `get_auto_ssd_cache_size(cache_dir)` for `"auto"` and `parse_size(...)` otherwise. — [source](https://raw.githubusercontent.com/jundot/omlx/main/omlx/settings.py)
- The environment variable `OMLX_SSD_CACHE_MAX_SIZE` overrides the setting, and validation accepts `"auto"` or a positive parsed size. — [source](https://raw.githubusercontent.com/jundot/omlx/main/omlx/settings.py)
- `CacheSettings.from_dict` replaces a hot-cache size of `"auto"` with `"0"`, and settings validation says hot-cache `'auto' is not supported`. — [source](https://raw.githubusercontent.com/jundot/omlx/main/omlx/settings.py)
- In v0.6.4, `get_ssd_cache_max_size_bytes` returned `int(get_ssd_capacity(cache_dir) * 0.1)` for `"auto"`, with the docstring "10% of SSD if auto". — [source](https://raw.githubusercontent.com/jundot/omlx/v0.6.4/omlx/settings.py)
- `get_ssd_capacity` returns `shutil.disk_usage(...).total` for the nearest existing path and assumes 500 GB if detection fails. — [source](https://raw.githubusercontent.com/jundot/omlx/main/omlx/settings.py)
- In v0.6.4, `_get_effective_max_size` returned `min(self._max_size, int((tracked_size + disk_free) * 0.99))` using a 30-second disk-usage cache. — [source](https://raw.githubusercontent.com/jundot/omlx/v0.6.4/omlx/cache/paged_ssd_cache.py)
- On main, `_get_effective_max_size` returns `disk_available // 2` in auto mode and `min(self._max_size, int(disk_available * 0.99))` otherwise, where `disk_available` is the cache size sampled with the free-space reading plus that free space. — [source](https://raw.githubusercontent.com/jundot/omlx/main/omlx/cache/paged_ssd_cache.py)
- The docstring says sampling both sizes together "keeps writes from increasing the cached budget"; the disk-usage snapshot is reused for 30 seconds. — [source](https://raw.githubusercontent.com/jundot/omlx/main/omlx/cache/paged_ssd_cache.py)
- `PagedSSDCacheManager.__init__` takes `auto_size: bool = False` documented as "Use 50% of the sum of free space and existing SSD cache", and logs `auto_size=` at startup. — [source](https://raw.githubusercontent.com/jundot/omlx/main/omlx/cache/paged_ssd_cache.py)
- `_enforce_size_limit_for_new_block` sets `target_size = effective_max - estimated_new_size`, falls back to 90% of the limit if negative, caps inline eviction at `_MAX_INLINE_UNLINKS_PER_SAVE = 32` unless `unbounded`, and warns about disk pressure only when `not self._auto_size`. — [source](https://raw.githubusercontent.com/jundot/omlx/main/omlx/cache/paged_ssd_cache.py)
- `_evict_tracked_until_size` walks the compatible, incompatible and GDN sidecar indexes in one LRU order, breaking ties in that fixed order. — [source](https://raw.githubusercontent.com/jundot/omlx/main/omlx/cache/paged_ssd_cache.py)
- On ENOSPC or EDQUOT during a write, the manager logs "SSD cache disk full" and sets `_disk_usage_cache = None` so the next save recomputes free space. — [source](https://raw.githubusercontent.com/jundot/omlx/main/omlx/cache/paged_ssd_cache.py)
- Commit 0faa19c changes the macOS app and README text from "auto = 10% of SSD capacity" to "50% of the sum of free space and existing SSD cache", adds `paged_ssd_cache_auto_size` to the scheduler config and `ssd_cache_auto_size_bytes` to the settings-defaults response. — [source](https://github.com/jundot/omlx/commit/0faa19cdcc38079bdc6aaa3e0a99a4239392502c)
- The README note added by 0faa19c says the default `auto` uses 50% of the sum of free disk space and existing SSD cache files, including GDN sidecars, refreshed during use and not shrinking because the cache grows or the server restarts. — [source](https://github.com/jundot/omlx/commit/0faa19cdcc38079bdc6aaa3e0a99a4239392502c)
- Issue 3829 was opened on 2026-09-22 by a user on an M2 Max 32 GB with 10 GB free, oMLX 0.6.4 from Homebrew, macOS 15, with the default `"ssd_cache_max_size": "auto"`. — [source](https://github.com/jundot/omlx/issues/3829)
- The maintainer's formula is `Cache limit = (free disk space + existing SSD cache) x 50%`, and the maintainer measured 3.22 GiB added by four distinct 8,192-token requests with 128-token outputs, about 53% in GDN sidecars, on Qwen3.8-27B 4-bit with memory detection set to 32 GB. — [source](https://github.com/jundot/omlx/issues/3829)
- The reporter corrected the figure: 6.5 GB was from 8 requests over two `omlx serve` sessions, so 0.81 GB per request, matching the maintainer's rate. — [source](https://github.com/jundot/omlx/issues/3829)
- The 0.7.0rc1 release notes list "Safer automatic SSD-cache sizing", basing automatic capacity on available disk space and including offline GDN sidecars in cache maintenance, crediting issues 3829 and 3883. — [source](https://github.com/jundot/omlx/releases)
- On a Mac with 10 GB free and an empty cache, the new `auto` limit is 5 GB, where v0.6.4 allowed up to about 99% of free space. — source: `asserted`
