<!-- llms-explorer concept facts · https://llms-explorer.com/tree/zipfian-hot-row-cache-tier-for-n-gram-tables-dra/ · pack 2026-10-05 · ~832 tokens -->

# Zipfian hot-row cache tier for n-gram tables (DRAM over NVMe)

> ds4 opens its Engram table with F_NOCACHE and F_RDAHEAD 0 on Apple platforms and documents the table descriptor as "never an mmap or Metal model view".

Parent: [Mac local LLMs: MoE streaming and offload](https://llms-explorer.com/tree/mac-local-llms-moe-streaming-and-offload/) · 1 facets · 14 facts · page: https://llms-explorer.com/tree/zipfian-hot-row-cache-tier-for-n-gram-tables-dra/

## Facts

- ds4 opens its Engram table with F_NOCACHE and F_RDAHEAD 0 on Apple platforms and documents the table descriptor as "never an mmap or Metal model view". — [source](https://raw.githubusercontent.com/antirez/ds4/main/ds4_engram.c)
- ds4's batch reader carries the comment "Fixed concurrency hides random-read latency without caching the table." — [source](https://raw.githubusercontent.com/antirez/ds4/main/ds4_engram.c)
- `ds4_engram_read_batch` processes at most 2,048 tokens per round, builds row requests, sorts them by row id with qsort, and gives consecutive duplicates a memcpy of the previous row instead of a second read. — [source](https://raw.githubusercontent.com/antirez/ds4/main/ds4_engram.c)
- ds4 uses 16 readers (`ENGRAM_READERS`) through `dispatch_apply_f` on the user-initiated global queue on Apple platforms, with a parallel threshold of 8 rows; on other platforms the threshold is 256 rows. — [source](https://raw.githubusercontent.com/antirez/ds4/main/ds4_engram.c)
- ds4's header states temporary storage for a batch read is bounded to 384 KiB independent of table and prefix size. — [source](https://raw.githubusercontent.com/antirez/ds4/main/ds4_engram.h)
- ds4's Qwen n-gram reader sorts requests by row id and reads at most 4,096 rows per round with up to 16 readers on Apple platforms. — [source](https://raw.githubusercontent.com/antirez/ds4/main/ds4.c)
- No cache, LRU, frequency or hot-row structure for Engram or n-gram rows appears in ds4's `ds4.c`, `ds4_engram.c` or `ds4_engram.h` at the 2026-10-04 snapshot. — [source](https://raw.githubusercontent.com/antirez/ds4/main/ds4.c)
- The paper says natural-language N-grams follow a Zipfian distribution and that this motivates a multi-level cache with hot embeddings in HBM or host DRAM and the long tail on NVMe. — [source](https://arxiv.org/html/2601.07372v1)
- llama.cpp PR 27794 reads lazily marked tensors through the mmap with prefetch skipped and relies on the page cache, with no row-level cache. — [source](https://github.com/ggml-org/llama.cpp/pull/27794)
- A downstream fork's measurement says the lazy-read win depends on the table not fitting in page cache, and that where it fits the benefit is far smaller. — [source](https://github.com/ggml-org/llama.cpp/pull/27794)
- With ds4's 264-byte rows, a 1 GiB cache of stored rows holds about 4.07 million rows, about 0.53% of the roughly 768 million V4.1 rows. — source: `asserted`
- Decoded rows are 256 FP32 values (1,024 bytes) against 264 stored bytes, so caching stored bytes is about 3.9 times denser. — source: `asserted`
- A frequency cache would have to live above ds4's F_NOCACHE descriptor, since F_NOCACHE reads never enter the unified buffer cache. — source: `asserted`
- No source found measures the Zipf exponent, hit rate or row-reuse distance for any shipped Engram or n-gram table. — source: `asserted`
