<!-- llms-explorer concept facts · https://llms-explorer.com/tree/tool-calling-benchmarks-for-local-models/ · pack 2026-10-05 · ~1482 tokens -->

# Tool-calling benchmarks for local models

> BFCL V1 (2024) introduced AST matching plus executable checks over simple, multiple, parallel and parallel-multiple function calls in several languages. V2 (2024-08-19) added enterprise and OSS-contributed "live" data. V3 (2024-09-19) added multi-turn and multi-step. V4 adds holistic agentic eval...

Parent: [Mac local LLMs: Chat templates, reasoning and tool calling](https://llms-explorer.com/tree/mac-local-llms-chat-templates-reasoning-and-tool-calling/) · 1 facets · 25 facts · page: https://llms-explorer.com/tree/tool-calling-benchmarks-for-local-models/

## Facts

- BFCL V1 (2024) introduced AST matching plus executable checks over simple, multiple, parallel and parallel-multiple function calls in several languages. V2 (2024-08-19) added enterprise and OSS-contributed "live" data. V3 (2024-09-19) added multi-turn and multi-step. V4 adds holistic agentic evaluation: web search, memory (key-value, vector, recursive summarization) and format sensitivity. — source: `asserted`
- The BFCL V4 table also carries hallucination columns (relevance, irrelevance), live and non-live single-turn columns, cost, and latency (mean, SD, P95). — source: `asserted`
- Docker model-test scores F1 over three dimensions (tool invocation, tool selection, parameter accuracy), accepts several valid tool sequences per case, and runs an agent loop capped at 5 rounds. — source: `asserted`
- BFCL leaderboard page last updated 2026-04-12 and is headed "From Tool Use to Agentic Evaluation of Large Language Models". — source: `asserted`
- Docker's study (2025-06-30) predates Qwen3.5/3.6 and is the only Mac-run (M4 Max 128 GB) tool-calling comparison found across 21 models. — source: `asserted`
- BFCL's overall score mixes agentic columns: a model can lead single-turn and still rank low overall. Qwen3-8B shows about 87-89 on non-live single turn but 12-13.5 on web search; xLAM-2-8b-fc-r shows 6.5 on web search yet 70 on multi-turn. — source: `asserted`
- Scores for the same model appear in two rows (function-calling mode versus prompt mode), differing by several points; the rows can be mistaken for duplicates. — source: `asserted`
- Docker found high BFCL-listed specialist models (xLAM-2-8b-fc-r, watt-tool-8B) failing a simple chat-cart task with eager invocation, wrong tool selection and bad arguments, and ranked them last (llama-xlam 0.570, watt-tool 0.484). — source: `asserted`
- Specialist versus general: BFCL ranks xLAM-2-32b-fc-r (overall 54.66) above Qwen3-32B (48.71 and 46.78); Docker ranked llama-xlam:8B (0.570) far below qwen3:8B (0.919-0.933). Different task shapes (BFCL breadth versus a five-tool shopping loop) and neither averaged. — source: `asserted`
- Bigger is not better in Docker's table: llama3.3:70B Q4_K_M scored 0.607, below qwen2.5:7B (0.753). — source: `asserted`
- No benchmark found that runs BFCL or similar through llama-server, mlx-lm or Ollama on a Mac with the chat template and parser the agent will actually use. — source: `asserted`
- Whether BFCL V4 multi-turn scores on a local endpoint depend on the harness's own prompt-mode handler rather than the server's tool parser. — source: `asserted`
- BFCL V3 (released 2024-09-19) added multi-turn and multi-step function-calling evaluation. — [source](https://gorilla.cs.berkeley.edu/blogs/8_berkeley_function_calling_leaderboard.html)
- BFCL V2 (released 2024-08-19) added enterprise-contributed "live" data. — [source](https://gorilla.cs.berkeley.edu/blogs/8_berkeley_function_calling_leaderboard.html)
- The BFCL V4 leaderboard groups columns as Agentic (web search, memory), Multi Turn, Single Turn (non-live, live), Hallucination Measurement and Format Sensitivity, plus cost and latency. — [source](https://gorilla.cs.berkeley.edu/leaderboard.html)
- BFCL V4 agentic memory is split into KV, Vector and Recursive Summarization backends. — [source](https://gorilla.cs.berkeley.edu/leaderboard.html)
- The BFCL V4 leaderboard was last updated 2026-04-12; rank 1 was Claude-Opus-4-5-20251101 at 77.47 overall. — [source](https://gorilla.cs.berkeley.edu/leaderboard.html)
- On that leaderboard Qwen3-8B scores 42.57 and 40.43 overall in its two listed rows, Qwen3-14B 41.03 and 37.77, Qwen3-32B 48.71 and 46.78, Qwen3-30B-A3B-Instruct-2507 41.39 and 36.70, Qwen3-4B-Instruct-2507 35.68 and 35.52. — [source](https://gorilla.cs.berkeley.edu/leaderboard.html)
- On that leaderboard xLAM-2-8b-fc-r scores 46.68 overall with 6.5 on web search and 70 on multi-turn. — [source](https://gorilla.cs.berkeley.edu/leaderboard.html)
- Docker's model-test measured 21 models over 3,570 test cases in 210 batch runs on a MacBook Pro M4 Max with 128 GB. — [source](https://www.docker.com/blog/local-llm-tool-calling-a-practical-evaluation/)
- Docker's F1 ranking: gpt-4 0.974, qwen3:14B-Q4_K_M 0.971, qwen3:14B-Q6_K 0.943, qwen3:8B-F16 0.933, qwen3:8B-Q4_K_M 0.919, llama3.1:8B-F16 0.835, llama3.3:70B-Q4_K_M 0.607, llama-xlam:8B-Q4_K_M 0.570, watt-tool:8B-Q4_K_M 0.484. — [source](https://www.docker.com/blog/local-llm-tool-calling-a-practical-evaluation/)
- Docker's model-test caps its simulated agent loop at 5 rounds and counts a case failed if the model still wants to continue. — [source](https://www.docker.com/blog/local-llm-tool-calling-a-practical-evaluation/)
- Docker's model-test accepts multiple valid tool-call variants per case, for example direct add or search then add. — [source](https://www.docker.com/blog/local-llm-tool-calling-a-practical-evaluation/)
- Docker concluded that even the Qwen3 8B variant beat every other local model it tested and that higher accuracy meant higher latency. — [source](https://www.docker.com/blog/local-llm-tool-calling-a-practical-evaluation/)
- MCP-Bench reports overall scores of 0.678 for qwen3-235b-a22b-2507, 0.627 for qwen3-30b-a3b-instruct-2507 and 0.428 for llama-3-1-8b-instruct. — [source](https://github.com/Accenture/mcp-bench)
