Engram row prefetch overlapped with preceding transformer blocks in Mac runtimes
Parent: Mac local LLMs: MoE streaming and offload · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
ds4 V4.1 prefill starts a reader pthread per Engram table that reads rows through `ds4_engram_read_batch` in 2,048-token chunks and publishes a completed-row counter with release semantics.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- ds4 V4.1 prefill starts a reader pthread per Engram table that reads rows through `ds4_engram_read_batch` in 2,048-token chunks and publishes a completed-row counter with release semantics. [source]
- The prefill consumer waits on that counter, polling with a 1 ms nanosleep, so a chunk's Engram layer can start before later chunks have been read. [source]
- ds4 enables the Engram overlap only when the prompt has at least 1,024 tokens, and offers environment switches to disable the prefetch, the pipelined wait, or the batched read. [source]
- ds4 V4.1 Engram layers are indices 1 and 14; table 0's prefetch starts before the layer loop and table 1's starts at layer 2, with a code comment saying it fills the buffer for layer 14 while the intervening encoder layers run. [source]
- Both prefetch starts write to the same `engram_prefetch` GPU-visible buffer. [source]
- ds4's decode step (`ds41_graph_step`) reads both tables' rows synchronously with `ds4_engram_read_batch` before it begins the GPU command buffer on Apple platforms. [source]
- ds4's Qwen n-gram path reads each 256-token chunk's rows synchronously, sorted by row id, with up to 16 readers via `dispatch_apply_f`, before the forward pass. [source]
- ds4's Engram constants are 2 layers, n-gram size 4, 8 heads, 24 columns per token per layer, 256 dimensions and 264 bytes per row. [source]
- `ds4_engram_hash` builds, for each of n-gram orders 2 to 4, 8 head ids per token per layer from a multiplicative hash modulo per-head primes, so 24 row ids per token per layer. [source]
- The V4.1 GGUF metadata in ds4 names Engram row counts 384,006,168 and 384,016,682, the encoding `e4m3_e8m0_32_row264` and a compressed vocabulary of 99,092. [source]
- One ds4 decode token issues 2 tables times 24 rows = 48 random 264-byte reads (12,672 bytes) before deduplication. [source]
- The two V4.1 tables total about 769 million rows times 264 bytes, roughly 189 GiB, which matches the 189 GiB of Engram tables in ds4's model documentation. [source]
- The paper's prefetch window is the compute of the blocks before the Engram layer, and the paper places the module at specific layers to use that buffer. [source]
- llama.cpp PR 27794 has a reviewer remark that qwen4 gives n-grams to the second layer, not the first, which can hide the lazy read. [source]
- A downstream fork's note on PR 27794 states upstream's own commit measured MADV_RANDOM without batched row prefetch at 94.4 s, an untouched mapping at 36.7 s, and MADV_RANDOM plus prefetch at 34.1 s. [source]
- The same fork measured its lazy-read default at pp512 352 tok/s versus 99 tok/s with lazy reading off, and tg128 33.4 versus 26.1 tok/s, on an AMD gfx1151 machine with a 95 GiB table and 62 GB RAM; the numbers are not Apple Silicon. [source]
- No source found overlaps Engram or n-gram reads with preceding blocks during decode on a Mac, and none reports a per-token stall for it. [source]
Children
- No children recorded.