SSD endurance and wear under local LLM inference
Parent: Mac local LLMs: Memory and wired limits · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
One NVMe "data unit" is 512,000 bytes: 303,071,863 units is printed as 155 TB, which is 303,071,863 x 512,000 B.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- One NVMe "data unit" is 512,000 bytes: 303,071,863 units is printed as 155 TB, which is 303,071,863 x 512,000 B. [source]
- Read counters do not age the drive the way write counters do; on the M1 reports large read totals (for example 161 TB) sit beside write totals of the same order, which is consistent with swap pages being written and then read back. [source]
- A 1 TB Apple SSD ages by about 1% of Percentage Used per roughly 27-35 TB written (derived from the reported pairs below), implying a rated life near 2.7-3.1 PB written. [source]
- Streaming experts from the weight file (Flash-MoE, ds4, SwiftLM) reads file-backed pages; evicting clean file-backed pages needs no write, so steady-state expert streaming should add essentially no wear, unlike OS swap of anonymous memory. [source]
- Swap of an oversized eagerly loaded model is the write-heavy case: the cached Ollama default-load report of 4.3 million swapouts on a 16 GB Mac is the extreme example (already in moe-expert-offload-to-ssd-on-macos.md); each swapout is a write. [source]
- Disk-persisted KV caches are the other writer. ds4's default budget is 4096 MiB, it stores a continued checkpoint every 10,240 tokens of a growing conversation, evicts by deleting files, and rewrites a 48-byte header on every hit; write volume therefore grows with total conversation length times checkpoint size, bounded by the budget only in resident bytes, not in lifetime writes. [source]
- JangPress generates a "prestack" overlay of about 150 GB per Kimi variant on first load into `JANGPRESS_PRESTACK_CACHE_DIR`; that is a one-time 150 GB write per bundle. [source]
- 2021-02-23: Hacker News item 26244093 (691 points) on a Linus Tech Tips thread reporting M1 Macs with very high SSD writes in the first weeks, attributed to aggressive swap. [source]
- 2021-07-07: a Medium analysis ("may be overblown") of the 8 GB M1 MacBook Air SSD concern, concluding the user would replace the machine before the SSD failed; the text beyond that is behind a paywall. [source]
- 2022-09-29: Ask HN 33024863 ("Has the Apple Silicon excessive disk read/write issue been fixed?") collects dozens of smartctl dumps; one reply says it has "mostly been fixed" and that individual apps may still trigger abusive writes. [source]
- Reported smartctl numbers on Apple Silicon were disputed as possibly misreported at the time: commenters questioned whether IOKit-backed smartctl values were correct, and one MacRumors user found DriveDX and smartmontools printing different TB totals (6.1 TB vs 6.7 TB) for the same block count. [source]
- A new MacBook Pro showed 110 GB written out of the box in one MacRumors report, so the counter includes factory and setup writes. [source]
- Apple does not publish a TBW rating, and older Macs exposed a wear count that M1-era tooling could not read at first. [source]
- Consumer NVMe retail specs can be used as a rough scale only: the compute-market guide lists a Samsung 990 Pro 4 TB at 2,400 TBW, unrelated to Apple's soldered NAND. [source]
- Heavy swap thrash can burn drive life quickly: one M1 report counted 155 TB written in 2.5 months (about 2 TB per day). [source]
- Is M1 swap wear a real danger? floatingatoll (HN 26244093): 1% used in 2 months extrapolates to about 16.5 years before read-only, so "what is the problem"; tyingq and the first LTT reports: 3% of a 2 TB drive in two months scales badly to a 256 GB drive (up to 24% in two months by the poster's estimate). tomxor doubted 25 MB/s sustained writes; yarcob answered that RAM-heavy swapping at 3000 MB/s could reach 155 TB in 14 hours. Neither side measured LLM workloads. [source]
- Did Apple fix it? One HN 33024863 reply says mostly fixed; the same thread's own data shows M1 laptops at 15-28% Percentage Used after 22 months to 5 months of heavy use. These are not reconciled. [source]
- No measured Data Units Written delta for a session of local inference (resident, streamed or swapped) on any Mac was found. [source]
- No source states whether macOS's compressed memory spills to swap faster under MoE page-cache pressure than under dense models. [source]
- No source measures write amplification for ds4/llama.cpp KV checkpoint saves on Apple SSDs. [source]
- A 1 TB M1 MacBook Air two and a half months old showed Percentage Used 5%, Data Units Read 315,158,527 (161 TB) and Data Units Written 303,071,863 (155 TB). [source]
- A 2 TB drive in a heavy-swap Linux laptop showed Percentage Used 1% at 18.4 TB written after more than a year, as a non-Apple reference. [source]
- A 2 TB Mac Studio M1 Max (April 2022) showed Percentage Used 0% with 25.5 TB read and 15.1 TB written after 428 power-on hours. [source]
- An M1 MacBook Pro (December 2020) showed Percentage Used 24% with 1.31 PB read and 732 TB written after about 22 months of heavy use. [source]
- A 2020 M1 Air 16 GB/1 TB showed Percentage Used 28% at 769 TB written and another M1 Air 16 GB/1 TB showed 15% at 438 TB written after about 5 months. [source]
- A 2020 M1 Air 8 GB/256 GB showed Percentage Used 4% at 63.0 TB written and an M1 Mac mini showed 5% at 35.9 TB written. [source]
- A commenter wrote that a 1 TB drive at 5 GB per day would last at least 10 years and that SSD makers should publish TBW. [source]
- A MacRumors M1 owner measured 5.07 TB written in two weeks (about 370 GB per day) on a 256 GB machine, and another reported about 2 TB per week. [source]
- On that MacRumors thread, DriveDX reported 6.1 TB written while smartmontools reported 6.7 TB for the same block count. [source]
- A new MacBook Pro showed 110 GB written out of the box in a MacRumors report. [source]
- compute-market advises that fast NVMe matters for model loading and swapping but not for inference itself, and lists the Samsung 990 Pro 4 TB at 7,450 MB/s read and 2,400 TBW. [source]
- The HILOS near-storage inference paper (ASPLOS 2026) describes the KV cache workload as write-once, read-many with endurance limited by total write volume, and estimates a 3.84 TB SmartSSD at 7.008 PBW. [source]
- JangPress's prestack overlay is about 150 GB per Kimi variant and its cache directory can be moved with `JANGPRESS_PRESTACK_CACHE_DIR`. [source]
- ds4's KV cache store writes checkpoints via a temp file and rename, rewrites a 48-byte header on every hit, and defaults to a 4096 MiB budget. [source]
- ThinkDifferent describes swap on a model larger than unified memory as a swap death spiral and advises keeping the model file under roughly 60-65% of RAM, with no wear figure. [source]
- Sitepoint reports throughput dropping 10x or more when a model plus context exceeds unified memory and macOS swaps, with no wear figure. [source]
- No source found measures SSD wear caused by local LLM inference itself. [source]
- Streaming experts from a read-only weight file adds essentially no wear because clean file-backed pages are dropped, not written. [source]
- The wear risk in a local LLM setup comes from swap, KV checkpoint writes and downloads, not from reading weights. [source]
Children
- No children recorded.