Zipfian hot-row cache tier for n-gram tables (DRAM over NVMe)
Parent: Mac local LLMs: MoE streaming and offload · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
ds4 opens its Engram table with F_NOCACHE and F_RDAHEAD 0 on Apple platforms and documents the table descriptor as "never an mmap or Metal model view".
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- ds4 opens its Engram table with F_NOCACHE and F_RDAHEAD 0 on Apple platforms and documents the table descriptor as "never an mmap or Metal model view". [source]
- ds4's batch reader carries the comment "Fixed concurrency hides random-read latency without caching the table." [source]
- `ds4_engram_read_batch` processes at most 2,048 tokens per round, builds row requests, sorts them by row id with qsort, and gives consecutive duplicates a memcpy of the previous row instead of a second read. [source]
- ds4 uses 16 readers (`ENGRAM_READERS`) through `dispatch_apply_f` on the user-initiated global queue on Apple platforms, with a parallel threshold of 8 rows; on other platforms the threshold is 256 rows. [source]
- ds4's header states temporary storage for a batch read is bounded to 384 KiB independent of table and prefix size. [source]
- ds4's Qwen n-gram reader sorts requests by row id and reads at most 4,096 rows per round with up to 16 readers on Apple platforms. [source]
- No cache, LRU, frequency or hot-row structure for Engram or n-gram rows appears in ds4's `ds4.c`, `ds4_engram.c` or `ds4_engram.h` at the 2026-10-04 snapshot. [source]
- The paper says natural-language N-grams follow a Zipfian distribution and that this motivates a multi-level cache with hot embeddings in HBM or host DRAM and the long tail on NVMe. [source]
- llama.cpp PR 27794 reads lazily marked tensors through the mmap with prefetch skipped and relies on the page cache, with no row-level cache. [source]
- A downstream fork's measurement says the lazy-read win depends on the table not fitting in page cache, and that where it fits the benefit is far smaller. [source]
- With ds4's 264-byte rows, a 1 GiB cache of stored rows holds about 4.07 million rows, about 0.53% of the roughly 768 million V4.1 rows. [source]
- Decoded rows are 256 FP32 values (1,024 bytes) against 264 stored bytes, so caching stored bytes is about 3.9 times denser. [source]
- A frequency cache would have to live above ds4's F_NOCACHE descriptor, since F_NOCACHE reads never enter the unified buffer cache. [source]
- No source found measures the Zipf exponent, hit rate or row-reuse distance for any shipped Engram or n-gram table. [source]
Children
- No children recorded.