Time-to-first-token measurement artifacts
Parent: Mac local LLMs: Benchmarking and comparisons · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Start edge: process launch and model load (cold), generate() call (warm in-process), or HTTP send (server).
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Start edge: process launch and model load (cold), generate() call (warm in-process), or HTTP send (server). [source]
- End edge: first decoded token's logits, first non-empty content delta over SSE, or first chunk of any kind. [source]
- A TTFT includes the whole prompt prefill plus one decode step, so it scales with uncached prompt tokens, not with total prompt tokens when a prompt cache hits. [source]
- Prefill speed derived from TTFT inherits both edges: a runtime without a prefill timer yields prefill = prompt tokens / (TTFT minus one decode step). [source]
- In the 2026-07-28/29 Gemma-4-E2B re-capture the apple-silicon-llm-bench harness fixed a LiteRT prefill-counter bug so that one instrument yields the cross-arm prefill column on both iPhone and Mac. [source]
- Decode-capped runs can lose prefill counters: the app path for LiteRT discarded per-turn counters whenever decode was capped, and the deep-context task always caps, so promptTokenCount was never reported; the fix records the prefill counters at the end of the prefill turn (before decode) and falls back to wall clock for decode only. [source]
- Fast runtimes make TTFT a fixed-overhead measurement: at short-chat prompts the M4 Max llama.cpp and mlx-swift TTFTs are 21-90 ms, dominated by launch overhead and one decode step, not prefill compute. [source]
- Stream chunks are not tokens: the bench counts streamed pieces as approximately tokens for some runtimes, and the Apple Foundation Model arm estimates tokens as utf8 bytes divided by 4 (+-20%), so any tok/s or prefill figure derived from them inherits that error. [source]
- Non-streaming Ollama timings fold model loading into the total: load_duration was 58% of total_duration in the documentation example. [source]
- A harness that counts only `content` deltas starts the TTFT clock late for a thinking model that streams its reasoning in a separate field, so the number reflects "time to visible answer", not "time to first token". [source]
- Which TTFT to headline. The bench defines TTFT as generate() call to first decoded token and reports it in integer milliseconds; llm-benchpacks defines it as first non-empty content delta over HTTP. On the same runtime these differ by HTTP and server overhead plus any reasoning preamble. Neither is averaged; compare only within one harness. [source]
- Whether llama-server's reported prompt_ms and the client-side first-content-delta time agree for the same request and build; no page read measures both. [source]
- How large the first-content-delta vs first-token gap is for thinking models on a Mac. [source]
- The bench defines TTFT as the time of the first decoded token minus the time generate was called, includes prefill in it, and reports it in integer milliseconds. [source]
- For runtimes with no separate prefill timer the bench infers prefill time as TTFT minus 1 divided by decode tok/s, that is one decoded token's worth of time. [source]
- The bench times with Date.timeIntervalSinceReferenceDate on the generating thread, and rejects DispatchTime (it wants wall time across thread hops) and Process.systemUptime (it pauses during sleep, which it wants to capture as an anomaly). [source]
- The bench samples thermalState every 1 s, resident memory every 100 ms during decode and once 200 ms after generation, and records inter-token latency p50, p95 and p99 from gaps between consecutive chunk events. [source]
- MediaPipeRuntime discarded LiteRT's per-turn prefill counters whenever decode was capped, so the deep-context task never reported promptTokenCount on the app path. [source]
- After the fix a capped run reports promptTok=1106 (equal to the Mac count) and capped-versus-EOS prefill agrees within 3.5%. [source]
- On the unified instrument LiteRT leads the cross-arm prefill column at 3,513 tok/s on iPhone and 7,803 tok/s on Mac, and the native Mac run reports 8,101. [source]
- On iPhone the native LiteRT prefill median is 3,889 tok/s (n=7, range 2,827 to 3,927) at context 2048. [source]
- M4 Max short-chat (128-token budget, n=3) TTFT in ms and decode tok/s: llama.cpp Qwen 2.5 0.5B 22 ms and 297.1, Gemma 4 E2B 41 ms and 119.2, Gemma 4 E4B 62 ms and 80.5. [source]
- M4 Max short-chat mlx-swift TTFT and decode: Qwen 2.5 0.5B 21 ms and 531.1, Gemma 4 E2B 68 ms and 185.4, Gemma 4 E4B 90 ms and 113.5, so llama.cpp has the lower TTFT and MLX the higher decode on the same models. [source]
- M4 Max short-chat CoreML/ANE TTFT ranges from 171 ms (Qwen 2.5 0.5B) to 665 ms (Qwen 3.5 2B) with decode of 181.2 and 35.0 tok/s. [source]
- The Apple Foundation Model arm measured TTFT 269 ms and 85.2 tok/s with tokens estimated as utf8 count divided by 4, which the README treats as +-20%. [source]
- The Gemma-4 per-layer-embedding S=1 unbatched prefill is called "the honest cost" of the patched Core AI engine's TTFT in the bench's own footnote. [source]
- Ollama's usage metrics are total_duration, load_duration, prompt_eval_count, prompt_eval_cached_count, prompt_eval_duration, eval_count and eval_duration, all in nanoseconds. [source]
- Ollama's prompt_eval_duration covers only the uncached prompt tokens, and prompt_eval_cached_count reports how many prompt tokens were read from cache. [source]
- For streaming endpoints Ollama puts the usage fields in the final chunk where done is true. [source]
- In Ollama's documented example load_duration is 101,397,084 ns of a 174,560,334 ns total, with 8 of 11 prompt tokens cached and 18 output tokens in 52,479,709 ns. [source]
- llm-benchpacks measures streaming TTFT from the first non-empty content delta, and with `--openai-stream-usage omit` TTFT is preserved but usage-derived token counts and rates may be null. [source]
- llm-benchpacks prints prefill_tps only where prompt tokens and cached prompt tokens agree across runtimes, so a prefill rate computed from a TTFT with a cache hit is suppressed. [source]
- llm-benchpacks' Ollama-native rows reported decode timing but no cached-prompt fields in the 2026-05-05 sweep, so cache parity could not be established for them. [source]
- A harness that counts only content deltas will report a thinking model's TTFT as time to the first visible answer token. [source]
Children
- No children recorded.