Task-success versus tool count and prompt size for small local models
Parent: Mac local LLMs: Agent clients, context and compaction · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Anthropic's tool-writing guide says too many or overlapping tools distract agents and recommends consolidating chained steps into one tool.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Anthropic's tool-writing guide says too many or overlapping tools distract agents and recommends consolidating chained steps into one tool. [source]
- The GitButler local-model gauntlet reports pass rates (Qwen3.5 35B 118/132, Qwen3-Coder 30B 114/132, GLM-4.7-Flash 99/132) on 22 problems, but it does not vary tool count or prompt size, so it is not evidence for this concept. [source]
- A controlled run varying tools (5, 20, 90) and listing size on one local model at fixed task set; not found in any source. [source]
- Anthropic's guide to writing tools states that too many tools or overlapping tools can distract agents from efficient strategies [source]
- GitButler's local benchmark gives Qwen3.5 35B 118/132, Qwen3-Coder 30B 114/132 and GLM-4.7-Flash 99/132 passes on a 22-problem set and does not vary tool count or prompt size [source]
- No source read measures local-model task success against tool count or listing size; the effect for 9B to 35B models is unmeasured here [source]
Children
- No children recorded.