<!-- llms-explorer concept facts · https://llms-explorer.com/tree/ssd-write-budgets-and-wear-for-kv-session-stores/ · pack 2026-10-05 · ~1511 tokens -->

# SSD write budgets and wear for KV session stores

> MTPLX's default write budget of 128G per rolling hour is 3.07 TB per day and about 1.12 PB per year if the budget is saturated for every hour (128 x 24 x 365 = 1,121,280 GB, taking "G" as decimal GB).

Parent: [Mac local LLMs: Prompt cache and persistent KV](https://llms-explorer.com/tree/mac-local-llms-prompt-cache-and-persistent-kv/) · 1 facets · 26 facts · page: https://llms-explorer.com/tree/ssd-write-budgets-and-wear-for-kv-session-stores/

## Facts

- MTPLX's default write budget of 128G per rolling hour is 3.07 TB per day and about 1.12 PB per year if the budget is saturated for every hour (128 x 24 x 365 = 1,121,280 GB, taking "G" as decimal GB). — source: `asserted`
- Against the existing derived rated life of about 2.7-3.1 PB for a 1 TB Apple SSD, a saturated 128G per hour budget would use the whole rating in about 2.4 to 2.8 years on a 1 TB machine. A 4 TB machine would scale the rating up roughly fourfold if endurance scales with capacity, which no Apple source states. — source: `asserted`
- The budget bounds the worst case; typical use writes far less. No source reports actual bytes written per hour from MTPLX, oMLX or ds4 on a real session. — source: `asserted`
- oMLX's cold tier writes KV blocks to SSD in safetensors format when the in-memory hot cache fills, as a "write-back" tier. — [source](https://github.com/jundot/omlx)
- oMLX's default SSD cache limit is `auto`: 50% of the sum of free disk space and existing SSD-cache files (including GDN sidecars), refreshed during use, and it "does not shrink simply because the cache grows or the server restarts". That is a size limit with no per-hour write limit. — [source](https://github.com/jundot/omlx)
- A shared-storage KV benchmark on Lustre shows the I/O shape: during prefill the workload is writes-only with near-zero reads; during decode reads dominate (11-18 MB average read size, 13-27 GB/s peaks) and writes are background fsync flushes. — [source](https://xinnor.io/blog/kv-cache-storage-is-the-new-ai-inference-bottleneck-how-lustre-ro-pcc-and-xiraid-make-shared-storage-competitive-with-local-nvme-for-long-context-llm-serving/)
- Wear therefore accrues in prefill: every long prompt that is stored writes its whole KV state once, and the stored copy is read only if a later prefix matches. Write volume scales with the number of distinct long prefixes, not with generated tokens. — source: `asserted`
- No dated history of KV-store write budgets was found beyond the MTPLX docs already recorded. — source: `asserted`
- A cap on cache bytes does not cap writes: an evict-and-rewrite loop at full capacity writes continuously. ds4, oMLX `auto` and MTPLX's backlog all bound resident or queued bytes; only MTPLX also bounds bytes per hour. — source: `asserted`
- oMLX's `auto` limit grows with free disk, so on a large empty SSD the cache can reach half the free space (hundreds of GB on a 1 TB drive with 700 GB free) before it evicts. This raises the volume that can be written before reuse, which matters most for agent sessions that rarely repeat a prefix. — [source](https://github.com/jundot/omlx)
- Apple SSD wear counters (Data Units Written) include factory and setup writes and disagree between tools (existing dossier), so a before-and-after counter delta around a session needs the same tool on both reads. — source: `asserted`
- Restoring from SSD is read-only and adds no wear; the wear question concerns saves only. — source: `asserted`
- Is the hourly budget protective? MTPLX's authors present it as a wear control. Computed over a year it still permits about 1.1 PB, close to a third of a 1 TB drive's derived rating per year, so it limits burst abuse rather than lifetime wear. The author's stated intent is not documented beyond "for SSD wear". — source: `asserted`
- Bytes written per session for any Mac KV store (MTPLX, oMLX, ds4, llama.cpp slot saves): none published. A one-day `smartctl -a` Data Units Written delta with the store on and off would settle it. — source: `asserted`
- Write amplification of safetensors block files on Apple's controller. — source: `asserted`
- Whether endurance scales with capacity on Apple's soldered NAND. — source: `asserted`
- The MTPLX default `MTPLX_SSD_WRITE_BUDGET_PER_HOUR` of 128G allows about 3.07 TB per day and about 1.12 PB per year if saturated. — source: `asserted`
- At that saturated rate a 1 TB Mac with the derived 2.7-3.1 PB rating would reach it in about 2.4 to 2.8 years. — source: `asserted`
- oMLX's cold tier stores KV blocks as safetensors files on SSD with write-back from a hot in-memory tier. — [source](https://github.com/jundot/omlx)
- oMLX's default SSD cache limit `auto` is 50% of free disk plus existing cache files and is refreshed during use. — [source](https://github.com/jundot/omlx)
- oMLX documents a size limit for its SSD cache and no per-hour write limit. — [source](https://github.com/jundot/omlx)
- In a Lustre KV-offload benchmark, prefill is writes-only with near-zero reads, and decode is read-dominated with writes limited to background fsync flushes. — [source](https://xinnor.io/blog/kv-cache-storage-is-the-new-ai-inference-bottleneck-how-lustre-ro-pcc-and-xiraid-make-shared-storage-competitive-with-local-nvme-for-long-context-llm-serving/)
- The same benchmark reports decode read sizes averaging 11-18 MB per I/O and read bandwidth peaking at 13-27 GB/s. — [source](https://xinnor.io/blog/kv-cache-storage-is-the-new-ai-inference-bottleneck-how-lustre-ro-pcc-and-xiraid-make-shared-storage-competitive-with-local-nvme-for-long-context-llm-serving/)
- KV-store wear is driven by the number of distinct long prefixes saved, because each save writes the whole prefix state once. — source: `asserted`
- A cache size cap does not bound lifetime writes; only a rolling write budget does, and only MTPLX among the stores read documents one. — source: `asserted`
- No source publishes measured bytes written per hour or per session for any Mac KV session store. — source: `asserted`
