Server-side prefix cache behavior under rolling-window transcripts

Parent: Mac local LLMs: Apple Foundation Models and Core AI · Published reference · snapshot 2026-10-05

↓ Facts as markdownall context files

llama-server, base reuse: `n_past` is the token-level common prefix of the slot's stored tokens and the new prompt. Everything after the first differing token is recomputed unless chunk reuse applies.

These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.

Facts

Children

← the whole tree · 3D view· how to read this page