Ollama count_tokens 404 stall with Claude Code
Parent: Mac local LLMs: Chat templates, reasoning and tool calling · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
The reporter's logs show four 404s in about two seconds (one `count_tokens?beta=true` and three `/v1/messages?beta=true`), not only on `count_tokens`.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- The reporter's logs show four 404s in about two seconds (one `count_tokens?beta=true` and three `/v1/messages?beta=true`), not only on `count_tokens`. [source]
- After the 404s the log shows `/v1/messages` returning 500 after 1m22s, then 10.5 s, 20.7 s and 41.7 s, after which the server was unresponsive and needed `docker restart ollama`. [source]
- The reporter ran Ollama 0.15.2 in Docker behind Traefik with HTTP/1.1 forced and a 1 ms flush interval, and says `curl` against `/v1/messages` with `stream: true` returned correct SSE every time right after a restart. [source]
- The reporter reproduced with both `llama3.2:1b` and `qwen3-coder:30b`, and says Claude Code triggers the problem "consistently every time it connects". [source]
- A second user saw similar behaviour on Ollama 0.15.2 started with `OLLAMA_CONTEXT_LENGTH=64000 ollama serve` on macOS under Homebrew. [source]
- A collaborator replied that Ollama does not support Claude telemetry and suggested `DISABLE_TELEMETRY=1`, `DISABLE_ERROR_REPORTING=1` and `CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1` to stop the 404s. [source]
- The reporter set all three variables on both the Ollama container and the Claude Code machine (with `ollama launch claude --model qwen3-coder:30b`) and still saw no reply to "hi". [source]
- Inferred: since the 404 is documented as harmless, the practical test is to run Claude Code against Ollama with a reverse proxy that answers `count_tokens` with a 200 and a rough `input_tokens`, and see whether the hang persists. [source]
- Inferred: an Ollama host that hangs after unsupported-route 404s is better diagnosed by checking the runner log for a model-load failure than by blaming `count_tokens`. [source]
Corrections and disagreements
- This CONTRADICTS the causal framing in ollama-anthropic-adapter-tool-use-block-fidelity.md ("count_tokens 404 ... stalls"): Claude Code's protocol page lists `/v1/messages/count_tokens` as optional and says a gateway lacking it gets a character-based estimate, so `/context` shows approximate counts, with no error. [source]
Children
- No children recorded.