Prompt-token counting parity across runtimes

Parent: Mac local LLMs: Benchmarking and comparisons · Published reference · snapshot 2026-10-05

↓ Facts as markdownall context files

llama-server OpenAI-style responses carry a `usage` object with `prompt_tokens` and `prompt_tokens_details.cached_tokens`, and a separate `timings` object with `cache_n` (prompt tokens reused from cache) and `prompt_n` (prompt tokens being processed).

These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.

Facts

Children

← the whole tree · 3D view· how to read this page