SSD write budgets and wear for KV session stores
Parent: Mac local LLMs: Prompt cache and persistent KV · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
MTPLX's default write budget of 128G per rolling hour is 3.07 TB per day and about 1.12 PB per year if the budget is saturated for every hour (128 x 24 x 365 = 1,121,280 GB, taking "G" as decimal GB).
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- MTPLX's default write budget of 128G per rolling hour is 3.07 TB per day and about 1.12 PB per year if the budget is saturated for every hour (128 x 24 x 365 = 1,121,280 GB, taking "G" as decimal GB). [source]
- Against the existing derived rated life of about 2.7-3.1 PB for a 1 TB Apple SSD, a saturated 128G per hour budget would use the whole rating in about 2.4 to 2.8 years on a 1 TB machine. A 4 TB machine would scale the rating up roughly fourfold if endurance scales with capacity, which no Apple source states. [source]
- The budget bounds the worst case; typical use writes far less. No source reports actual bytes written per hour from MTPLX, oMLX or ds4 on a real session. [source]
- oMLX's cold tier writes KV blocks to SSD in safetensors format when the in-memory hot cache fills, as a "write-back" tier. [source]
- oMLX's default SSD cache limit is `auto`: 50% of the sum of free disk space and existing SSD-cache files (including GDN sidecars), refreshed during use, and it "does not shrink simply because the cache grows or the server restarts". That is a size limit with no per-hour write limit. [source]
- A shared-storage KV benchmark on Lustre shows the I/O shape: during prefill the workload is writes-only with near-zero reads; during decode reads dominate (11-18 MB average read size, 13-27 GB/s peaks) and writes are background fsync flushes. [source]
- Wear therefore accrues in prefill: every long prompt that is stored writes its whole KV state once, and the stored copy is read only if a later prefix matches. Write volume scales with the number of distinct long prefixes, not with generated tokens. [source]
- No dated history of KV-store write budgets was found beyond the MTPLX docs already recorded. [source]
- A cap on cache bytes does not cap writes: an evict-and-rewrite loop at full capacity writes continuously. ds4, oMLX `auto` and MTPLX's backlog all bound resident or queued bytes; only MTPLX also bounds bytes per hour. [source]
- oMLX's `auto` limit grows with free disk, so on a large empty SSD the cache can reach half the free space (hundreds of GB on a 1 TB drive with 700 GB free) before it evicts. This raises the volume that can be written before reuse, which matters most for agent sessions that rarely repeat a prefix. [source]
- Apple SSD wear counters (Data Units Written) include factory and setup writes and disagree between tools (existing dossier), so a before-and-after counter delta around a session needs the same tool on both reads. [source]
- Restoring from SSD is read-only and adds no wear; the wear question concerns saves only. [source]
- Is the hourly budget protective? MTPLX's authors present it as a wear control. Computed over a year it still permits about 1.1 PB, close to a third of a 1 TB drive's derived rating per year, so it limits burst abuse rather than lifetime wear. The author's stated intent is not documented beyond "for SSD wear". [source]
- Bytes written per session for any Mac KV store (MTPLX, oMLX, ds4, llama.cpp slot saves): none published. A one-day `smartctl -a` Data Units Written delta with the store on and off would settle it. [source]
- Write amplification of safetensors block files on Apple's controller. [source]
- Whether endurance scales with capacity on Apple's soldered NAND. [source]
- The MTPLX default `MTPLX_SSD_WRITE_BUDGET_PER_HOUR` of 128G allows about 3.07 TB per day and about 1.12 PB per year if saturated. [source]
- At that saturated rate a 1 TB Mac with the derived 2.7-3.1 PB rating would reach it in about 2.4 to 2.8 years. [source]
- oMLX's cold tier stores KV blocks as safetensors files on SSD with write-back from a hot in-memory tier. [source]
- oMLX's default SSD cache limit `auto` is 50% of free disk plus existing cache files and is refreshed during use. [source]
- oMLX documents a size limit for its SSD cache and no per-hour write limit. [source]
- In a Lustre KV-offload benchmark, prefill is writes-only with near-zero reads, and decode is read-dominated with writes limited to background fsync flushes. [source]
- The same benchmark reports decode read sizes averaging 11-18 MB per I/O and read bandwidth peaking at 13-27 GB/s. [source]
- KV-store wear is driven by the number of distinct long prefixes saved, because each save writes the whole prefix state once. [source]
- A cache size cap does not bound lifetime writes; only a rolling write budget does, and only MTPLX among the stores read documents one. [source]
- No source publishes measured bytes written per hour or per session for any Mac KV session store. [source]
Children
- No children recorded.