<!-- llms-explorer concept facts · https://llms-explorer.com/tree/task-success-versus-tool-count-and-prompt-size-f/ · pack 2026-10-05 · ~411 tokens -->

# Task-success versus tool count and prompt size for small local models

> Anthropic's tool-writing guide says too many or overlapping tools distract agents and recommends consolidating chained steps into one tool.

Parent: [Mac local LLMs: Agent clients, context and compaction](https://llms-explorer.com/tree/mac-local-llms-agent-clients-context-and-compaction/) · 1 facets · 6 facts · page: https://llms-explorer.com/tree/task-success-versus-tool-count-and-prompt-size-f/

## Facts

- Anthropic's tool-writing guide says too many or overlapping tools distract agents and recommends consolidating chained steps into one tool. — source: `asserted`
- The GitButler local-model gauntlet reports pass rates (Qwen3.5 35B 118/132, Qwen3-Coder 30B 114/132, GLM-4.7-Flash 99/132) on 22 problems, but it does not vary tool count or prompt size, so it is not evidence for this concept. — source: `asserted`
- A controlled run varying tools (5, 20, 90) and listing size on one local model at fixed task set; not found in any source. — source: `asserted`
- Anthropic's guide to writing tools states that too many tools or overlapping tools can distract agents from efficient strategies — [source](https://www.anthropic.com/engineering/writing-tools-for-agents)
- GitButler's local benchmark gives Qwen3.5 35B 118/132, Qwen3-Coder 30B 114/132 and GLM-4.7-Flash 99/132 passes on a 22-problem set and does not vary tool count or prompt size — [source](https://blog.gitbutler.com/local-llm-gauntlet)
- No source read measures local-model task success against tool count or listing size; the effect for 9B to 35B models is unmeasured here — source: `asserted`
