Mac local LLMs: KV cache sizing and quantization

Parent: Running LLM models locally on a Mac · Published reference · snapshot 2026-10-05

↓ Facts as markdownall context files

KV/token = attn layers x 2 x KV heads x head_dim x bytes. Hybrids: count only full-attention layers; sliding layers add a bounded ring, DeltaNet a constant state.

These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.

Sizing (count only layers that hold KV)

Runtime knobs and defaults

Failure modes and fixes

KV quantization decisions

Tool calls and agents

Corrections to earlier claims

Open

Corrections and disagreements

Concepts in this cluster

Children

← the whole tree · 3D view· how to read this page