<!-- llms-explorer concept facts · https://llms-explorer.com/tree/engram-and-n-gram-embedding-tables-read-from-dis/ · pack 2026-10-05 · ~1055 tokens -->

# Engram and n-gram embedding tables read from disk per token

> Engram (arXiv 2601.07372) retrieves static embeddings for suffix n-grams by deterministic hashing, using tokenizer compression and K hash heads per n-gram order, then fuses them with the hidden state through context-aware gating.

Parent: [Mac local LLMs: MoE streaming and offload](https://llms-explorer.com/tree/mac-local-llms-moe-streaming-and-offload/) · 1 facets · 17 facts · page: https://llms-explorer.com/tree/engram-and-n-gram-embedding-tables-read-from-dis/

## Facts

- Engram (arXiv 2601.07372) retrieves static embeddings for suffix n-grams by deterministic hashing, using tokenizer compression and K hash heads per n-gram order, then fuses them with the hidden state through context-aware gating. — [source](https://arxiv.org/pdf/2601.07372)
- The published 27B configs use n-gram orders 2 and 3, 8 heads, d_mem 1280 and Engram layers 2 and 15. — [source](https://arxiv.org/pdf/2601.07372)
- The hash is described as a lightweight multiplicative-XOR function into prime-sized per-head tables. — [source](https://arxiv.org/pdf/2601.07372)
- Row indices are fixed once the token sequence is known and can be computed before the corresponding layer executes, so reads can be prefetched and overlapped with earlier blocks. — [source](https://arxiv.org/pdf/2601.07372)
- The paper measured a 100B-parameter Engram table in host DRAM on one NVIDIA H800 with nano-vLLM: 4B dense baseline 9,031.62 tok/s versus 8,858.28 tok/s with offload, 8B dense 6,315.52 versus 6,140.02, a worst-case penalty of 2.8%. — [source](https://arxiv.org/pdf/2601.07372)
- The paper's benchmark forced all retrievals over PCIe from host memory and used no locality cache, and the authors call it a conservative baseline. — [source](https://arxiv.org/pdf/2601.07372)
- The paper describes a multi-level cache hierarchy (HBM or DRAM for hot rows, NVMe SSD for the long tail) as a design option motivated by Zipfian n-gram frequency, and reports no NVMe measurement. — [source](https://arxiv.org/pdf/2601.07372)
- The paper states the effective communication volume per step scales with the number of activated slots rather than the total table size. — [source](https://arxiv.org/pdf/2601.07372)
- The paper's ablation says modeling quality favors an early Engram layer while system latency hiding favors a later one, so placement must satisfy both. — [source](https://arxiv.org/pdf/2601.07372)
- The official deepseek-ai/Engram repository contains the paper, figures and one standalone demo script (engram_demo_v1.py) and no inference or offload runtime. — [source](https://github.com/deepseek-ai/Engram)
- In XNU, F_NOCACHE sets VNOCACHE_DATA, which makes the cluster read path set IO_NOCACHE, which disables read-ahead and stops the read from taking a page reference. — [source](https://raw.githubusercontent.com/apple-oss-distributions/xnu/main/bsd/vfs/vfs_cluster.c)
- With IO_NOCACHE a user-space read becomes a candidate for the direct I/O path (cluster_io_type picks IO_DIRECT when the iovec is long enough), and in the buffered copy path the pages are aborted with UPL_ABORT_DUMP_PAGES instead of being committed to the cache. — [source](https://raw.githubusercontent.com/apple-oss-distributions/xnu/main/bsd/vfs/vfs_cluster.c)
- ds4's separate F_NOCACHE plus F_RDAHEAD 0 descriptor for Engram rows therefore matches what the kernel does: no read-ahead and no cache residency for single-row reads. — source: `asserted`
- Engram row reads on a Mac cannot hit the page cache after the first read when F_NOCACHE is set, so every row read is an SSD read, unlike the warm-cache expert reads in Flash-MoE. — source: `asserted`
- F_NOCACHE does not apply to file pages faulted through mmap, so a runtime that maps the Engram table gets page-cache residency and reclaim behavior instead. — source: `asserted`
- On a Mac the paper's prefetch window (one transformer block of compute) must cover an NVMe random-read latency instead of a PCIe-from-DRAM copy, so the 2.8% overhead figure does not transfer. — source: `asserted`
- No source found measures Engram or n-gram table reads on an Apple SSD in rows per second, bytes per token or latency per row. — source: `asserted`
