Engram and n-gram embedding tables read from disk per token
Parent: Mac local LLMs: MoE streaming and offload · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Engram (arXiv 2601.07372) retrieves static embeddings for suffix n-grams by deterministic hashing, using tokenizer compression and K hash heads per n-gram order, then fuses them with the hidden state through context-aware gating.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Engram (arXiv 2601.07372) retrieves static embeddings for suffix n-grams by deterministic hashing, using tokenizer compression and K hash heads per n-gram order, then fuses them with the hidden state through context-aware gating. [source]
- The published 27B configs use n-gram orders 2 and 3, 8 heads, d_mem 1280 and Engram layers 2 and 15. [source]
- The hash is described as a lightweight multiplicative-XOR function into prime-sized per-head tables. [source]
- Row indices are fixed once the token sequence is known and can be computed before the corresponding layer executes, so reads can be prefetched and overlapped with earlier blocks. [source]
- The paper measured a 100B-parameter Engram table in host DRAM on one NVIDIA H800 with nano-vLLM: 4B dense baseline 9,031.62 tok/s versus 8,858.28 tok/s with offload, 8B dense 6,315.52 versus 6,140.02, a worst-case penalty of 2.8%. [source]
- The paper's benchmark forced all retrievals over PCIe from host memory and used no locality cache, and the authors call it a conservative baseline. [source]
- The paper describes a multi-level cache hierarchy (HBM or DRAM for hot rows, NVMe SSD for the long tail) as a design option motivated by Zipfian n-gram frequency, and reports no NVMe measurement. [source]
- The paper states the effective communication volume per step scales with the number of activated slots rather than the total table size. [source]
- The paper's ablation says modeling quality favors an early Engram layer while system latency hiding favors a later one, so placement must satisfy both. [source]
- The official deepseek-ai/Engram repository contains the paper, figures and one standalone demo script (engram_demo_v1.py) and no inference or offload runtime. [source]
- In XNU, F_NOCACHE sets VNOCACHE_DATA, which makes the cluster read path set IO_NOCACHE, which disables read-ahead and stops the read from taking a page reference. [source]
- With IO_NOCACHE a user-space read becomes a candidate for the direct I/O path (cluster_io_type picks IO_DIRECT when the iovec is long enough), and in the buffered copy path the pages are aborted with UPL_ABORT_DUMP_PAGES instead of being committed to the cache. [source]
- ds4's separate F_NOCACHE plus F_RDAHEAD 0 descriptor for Engram rows therefore matches what the kernel does: no read-ahead and no cache residency for single-row reads. [source]
- Engram row reads on a Mac cannot hit the page cache after the first read when F_NOCACHE is set, so every row read is an SSD read, unlike the warm-cache expert reads in Flash-MoE. [source]
- F_NOCACHE does not apply to file pages faulted through mmap, so a runtime that maps the Engram table gets page-cache residency and reclaim behavior instead. [source]
- On a Mac the paper's prefetch window (one transformer block of compute) must cover an NVMe random-read latency instead of a PCIe-from-DRAM copy, so the 2.8% overhead figure does not transfer. [source]
- No source found measures Engram or n-gram table reads on an Apple SSD in rows per second, bytes per token or latency per row. [source]
Children
- No children recorded.