LCKV and YOCO cross-layer KV sharing variants
Parent: Mac local LLMs: KV cache sizing and quantization · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Gemma 4's scheme (last N layers reuse K/V of the last non-shared layer of the same attention type) resembles YOCO's bottom-target sharing but is per attention type and keeps sliding and global caches separate.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Gemma 4's scheme (last N layers reuse K/V of the last non-shared layer of the same attention type) resembles YOCO's bottom-target sharing but is per attention type and keeps sliding and global caches separate. [source]
- YOCO stacks a cross-decoder on a self-decoder; the self-decoder encodes a global KV cache that the cross-decoder reuses through cross-attention, and prefill can exit early. [source]
- In YOCO the first L/2 layers form the self-decoder, whose output at layer L/2 produces the shared K and V; the remaining layers are the cross-decoder. [source]
- YOCO's self-decoder uses constant-memory efficient attention (sliding-window attention or gated retention), so total cache is O(N + C L). [source]
- YOCO reports about 80x less KV cache memory for 65B models, a 71.8x faster 1M-token prefill and 2.87x faster 32K prefill; a 512K prefill drops from 180 seconds to under 6. [source]
- YOCO's paper gives 512K tokens of a 65B GQA model with 8-bit KV as about 86 GB of GPU memory. [source]
- LCKV pairs queries of all layers with the KVs of the top layer only and discards K and V weights for the other layers. [source]
- LCKV masks the attention diagonal so a token does not attend to itself, using zero vectors as dummy KVs for the first token. [source]
- LCKV keeps standard attention in w warmup layers, half at the top and half at the bottom, and the sandwich placement beat all-bottom and all-top. [source]
- LCKV trains with an approximate parallel scheme of m iterations (default m = 7, b = 2), and pre-training took about 3 times as long as TinyLlama on the same data. [source]
- LCKV prompt encoding needs m + b iterations, which the paper calls negligible against the tokens generated. [source]
- LCKV reports up to 26x higher throughput and up to 32x larger batch sizes than standard transformers for 1B to 30B models, and combines with StreamingLLM. [source]
- The systematic study defines kv(i) as the layer whose KVs pair with layer i's queries, and partitions layers as pizza, sandwich or lasagna with target at bottom, top or middle. [source]
- In the study's framework sandwich-top is LCKV, pizza-bottom is YOCO and lasagna-bottom is CLA. [source]
- The study finds top and middle configurations lose throughput on long prompts (512+1024) because of iterative prompt encoding, while bottom configurations stay well above the baseline. [source]
- The study finds that when only half the layers rely on other layers' KVs most configurations match the baseline, and with more sharing the bottom-target configurations degrade most. [source]
- The study notes that a layer reusing another layer's KVs needs no W_K or W_V, so the number of KV layers sets those parameter counts. [source]
Children
- No children recorded.