<!-- llms-explorer concept facts · https://llms-explorer.com/tree/ollama-count-tokens-404-stall-with-claude-code/ · pack 2026-10-05 · ~745 tokens -->

# Ollama count_tokens 404 stall with Claude Code

> The reporter's logs show four 404s in about two seconds (one `count_tokens?beta=true` and three `/v1/messages?beta=true`), not only on `count_tokens`.

Parent: [Mac local LLMs: Chat templates, reasoning and tool calling](https://llms-explorer.com/tree/mac-local-llms-chat-templates-reasoning-and-tool-calling/) · 2 facets · 10 facts · page: https://llms-explorer.com/tree/ollama-count-tokens-404-stall-with-claude-code/

## Facts

- The reporter's logs show four 404s in about two seconds (one `count_tokens?beta=true` and three `/v1/messages?beta=true`), not only on `count_tokens`. — [source](https://github.com/ollama/ollama/issues/13949)
- After the 404s the log shows `/v1/messages` returning 500 after 1m22s, then 10.5 s, 20.7 s and 41.7 s, after which the server was unresponsive and needed `docker restart ollama`. — [source](https://github.com/ollama/ollama/issues/13949)
- The reporter ran Ollama 0.15.2 in Docker behind Traefik with HTTP/1.1 forced and a 1 ms flush interval, and says `curl` against `/v1/messages` with `stream: true` returned correct SSE every time right after a restart. — [source](https://github.com/ollama/ollama/issues/13949)
- The reporter reproduced with both `llama3.2:1b` and `qwen3-coder:30b`, and says Claude Code triggers the problem "consistently every time it connects". — [source](https://github.com/ollama/ollama/issues/13949)
- A second user saw similar behaviour on Ollama 0.15.2 started with `OLLAMA_CONTEXT_LENGTH=64000 ollama serve` on macOS under Homebrew. — [source](https://github.com/ollama/ollama/issues/13949)
- A collaborator replied that Ollama does not support Claude telemetry and suggested `DISABLE_TELEMETRY=1`, `DISABLE_ERROR_REPORTING=1` and `CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1` to stop the 404s. — [source](https://github.com/ollama/ollama/issues/13949)
- The reporter set all three variables on both the Ollama container and the Claude Code machine (with `ollama launch claude --model qwen3-coder:30b`) and still saw no reply to "hi". — [source](https://github.com/ollama/ollama/issues/13949)
- Inferred: since the 404 is documented as harmless, the practical test is to run Claude Code against Ollama with a reverse proxy that answers `count_tokens` with a 200 and a rough `input_tokens`, and see whether the hang persists. — source: `asserted`
- Inferred: an Ollama host that hangs after unsupported-route 404s is better diagnosed by checking the runner log for a model-load failure than by blaming `count_tokens`. — source: `asserted`

## Corrections and disagreements

- This CONTRADICTS the causal framing in ollama-anthropic-adapter-tool-use-block-fidelity.md ("count_tokens 404 ... stalls"): Claude Code's protocol page lists `/v1/messages/count_tokens` as optional and says a gateway lacking it gets a character-based estimate, so `/context` shows approximate counts, with no error. — [source](https://code.claude.com/docs/en/llm-gateway-protocol)
