<!-- llms-explorer concept facts · https://llms-explorer.com/tree/ssd-endurance-and-wear-under-local-llm-inference/ · pack 2026-10-05 · ~2195 tokens -->

# SSD endurance and wear under local LLM inference

> One NVMe "data unit" is 512,000 bytes: 303,071,863 units is printed as 155 TB, which is 303,071,863 x 512,000 B.

Parent: [Mac local LLMs: Memory and wired limits](https://llms-explorer.com/tree/mac-local-llms-memory-and-wired-limits/) · 1 facets · 39 facts · page: https://llms-explorer.com/tree/ssd-endurance-and-wear-under-local-llm-inference/

## Facts

- One NVMe "data unit" is 512,000 bytes: 303,071,863 units is printed as 155 TB, which is 303,071,863 x 512,000 B. — source: `asserted`
- Read counters do not age the drive the way write counters do; on the M1 reports large read totals (for example 161 TB) sit beside write totals of the same order, which is consistent with swap pages being written and then read back. — source: `asserted`
- A 1 TB Apple SSD ages by about 1% of Percentage Used per roughly 27-35 TB written (derived from the reported pairs below), implying a rated life near 2.7-3.1 PB written. — source: `asserted`
- Streaming experts from the weight file (Flash-MoE, ds4, SwiftLM) reads file-backed pages; evicting clean file-backed pages needs no write, so steady-state expert streaming should add essentially no wear, unlike OS swap of anonymous memory. — source: `asserted`
- Swap of an oversized eagerly loaded model is the write-heavy case: the cached Ollama default-load report of 4.3 million swapouts on a 16 GB Mac is the extreme example (already in moe-expert-offload-to-ssd-on-macos.md); each swapout is a write. — source: `asserted`
- Disk-persisted KV caches are the other writer. ds4's default budget is 4096 MiB, it stores a continued checkpoint every 10,240 tokens of a growing conversation, evicts by deleting files, and rewrites a 48-byte header on every hit; write volume therefore grows with total conversation length times checkpoint size, bounded by the budget only in resident bytes, not in lifetime writes. — source: `asserted`
- JangPress generates a "prestack" overlay of about 150 GB per Kimi variant on first load into `JANGPRESS_PRESTACK_CACHE_DIR`; that is a one-time 150 GB write per bundle. — source: `asserted`
- 2021-02-23: Hacker News item 26244093 (691 points) on a Linus Tech Tips thread reporting M1 Macs with very high SSD writes in the first weeks, attributed to aggressive swap. — source: `asserted`
- 2021-07-07: a Medium analysis ("may be overblown") of the 8 GB M1 MacBook Air SSD concern, concluding the user would replace the machine before the SSD failed; the text beyond that is behind a paywall. — source: `asserted`
- 2022-09-29: Ask HN 33024863 ("Has the Apple Silicon excessive disk read/write issue been fixed?") collects dozens of smartctl dumps; one reply says it has "mostly been fixed" and that individual apps may still trigger abusive writes. — source: `asserted`
- Reported smartctl numbers on Apple Silicon were disputed as possibly misreported at the time: commenters questioned whether IOKit-backed smartctl values were correct, and one MacRumors user found DriveDX and smartmontools printing different TB totals (6.1 TB vs 6.7 TB) for the same block count. — source: `asserted`
- A new MacBook Pro showed 110 GB written out of the box in one MacRumors report, so the counter includes factory and setup writes. — source: `asserted`
- Apple does not publish a TBW rating, and older Macs exposed a wear count that M1-era tooling could not read at first. — source: `asserted`
- Consumer NVMe retail specs can be used as a rough scale only: the compute-market guide lists a Samsung 990 Pro 4 TB at 2,400 TBW, unrelated to Apple's soldered NAND. — source: `asserted`
- Heavy swap thrash can burn drive life quickly: one M1 report counted 155 TB written in 2.5 months (about 2 TB per day). — source: `asserted`
- Is M1 swap wear a real danger? floatingatoll (HN 26244093): 1% used in 2 months extrapolates to about 16.5 years before read-only, so "what is the problem"; tyingq and the first LTT reports: 3% of a 2 TB drive in two months scales badly to a 256 GB drive (up to 24% in two months by the poster's estimate). tomxor doubted 25 MB/s sustained writes; yarcob answered that RAM-heavy swapping at 3000 MB/s could reach 155 TB in 14 hours. Neither side measured LLM workloads. — source: `asserted`
- Did Apple fix it? One HN 33024863 reply says mostly fixed; the same thread's own data shows M1 laptops at 15-28% Percentage Used after 22 months to 5 months of heavy use. These are not reconciled. — source: `asserted`
- No measured Data Units Written delta for a session of local inference (resident, streamed or swapped) on any Mac was found. — source: `asserted`
- No source states whether macOS's compressed memory spills to swap faster under MoE page-cache pressure than under dense models. — source: `asserted`
- No source measures write amplification for ds4/llama.cpp KV checkpoint saves on Apple SSDs. — source: `asserted`
- A 1 TB M1 MacBook Air two and a half months old showed Percentage Used 5%, Data Units Read 315,158,527 (161 TB) and Data Units Written 303,071,863 (155 TB). — [source](https://news.ycombinator.com/item?id=26244093)
- A 2 TB drive in a heavy-swap Linux laptop showed Percentage Used 1% at 18.4 TB written after more than a year, as a non-Apple reference. — [source](https://news.ycombinator.com/item?id=26244093)
- A 2 TB Mac Studio M1 Max (April 2022) showed Percentage Used 0% with 25.5 TB read and 15.1 TB written after 428 power-on hours. — [source](https://news.ycombinator.com/item?id=33024863)
- An M1 MacBook Pro (December 2020) showed Percentage Used 24% with 1.31 PB read and 732 TB written after about 22 months of heavy use. — [source](https://news.ycombinator.com/item?id=33024863)
- A 2020 M1 Air 16 GB/1 TB showed Percentage Used 28% at 769 TB written and another M1 Air 16 GB/1 TB showed 15% at 438 TB written after about 5 months. — [source](https://news.ycombinator.com/item?id=33024863)
- A 2020 M1 Air 8 GB/256 GB showed Percentage Used 4% at 63.0 TB written and an M1 Mac mini showed 5% at 35.9 TB written. — [source](https://news.ycombinator.com/item?id=33024863)
- A commenter wrote that a 1 TB drive at 5 GB per day would last at least 10 years and that SSD makers should publish TBW. — [source](https://news.ycombinator.com/item?id=33024863)
- A MacRumors M1 owner measured 5.07 TB written in two weeks (about 370 GB per day) on a 256 GB machine, and another reported about 2 TB per week. — [source](https://forums.macrumors.com/threads/smartmontool-reporting-abnormally-high-ssd-write-activities.2273896/)
- On that MacRumors thread, DriveDX reported 6.1 TB written while smartmontools reported 6.7 TB for the same block count. — [source](https://forums.macrumors.com/threads/smartmontool-reporting-abnormally-high-ssd-write-activities.2273896/)
- A new MacBook Pro showed 110 GB written out of the box in a MacRumors report. — [source](https://forums.macrumors.com/threads/smartmontool-reporting-abnormally-high-ssd-write-activities.2273896/)
- compute-market advises that fast NVMe matters for model loading and swapping but not for inference itself, and lists the Samsung 990 Pro 4 TB at 7,450 MB/s read and 2,400 TBW. — [source](https://www.compute-market.com/blog/best-nvme-ssd-local-ai-llm-2026)
- The HILOS near-storage inference paper (ASPLOS 2026) describes the KV cache workload as write-once, read-many with endurance limited by total write volume, and estimates a 3.84 TB SmartSSD at 7.008 PBW. — [source](https://arxiv.org/html/2502.09921v2)
- JangPress's prestack overlay is about 150 GB per Kimi variant and its cache directory can be moved with `JANGPRESS_PRESTACK_CACHE_DIR`. — [source](https://raw.githubusercontent.com/jjang-ai/jangq/main/scripts/jangpress/README.md)
- ds4's KV cache store writes checkpoints via a temp file and rename, rewrites a 48-byte header on every hit, and defaults to a 4096 MiB budget. — [source](https://raw.githubusercontent.com/antirez/ds4/main/ds4_kvstore.c)
- ThinkDifferent describes swap on a model larger than unified memory as a swap death spiral and advises keeping the model file under roughly 60-65% of RAM, with no wear figure. — [source](https://www.thinkdifferent.blog/blog/the-7-ai-mistakes-mac-power-users-keep-making/)
- Sitepoint reports throughput dropping 10x or more when a model plus context exceeds unified memory and macOS swaps, with no wear figure. — [source](https://www.sitepoint.com/local-llms-apple-silicon-mac-2026/)
- No source found measures SSD wear caused by local LLM inference itself. — source: `asserted`
- Streaming experts from a read-only weight file adds essentially no wear because clean file-backed pages are dropped, not written. — source: `asserted`
- The wear risk in a local LLM setup comes from swap, KV checkpoint writes and downloads, not from reading weights. — source: `asserted`
