BFCL and tool-calling evals for quantized local models
Parent: Mac local LLMs: Quantization evaluation · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
BFCL's `bfcl generate` can start vllm or sglang itself (`--backend`, default vllm) or reuse a running server with `--skip-server-setup`; the endpoint defaults to localhost:1053 and is set by LOCAL_SERVER_ENDPOINT and LOCAL_SERVER_PORT in `.env`.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- BFCL's `bfcl generate` can start vllm or sglang itself (`--backend`, default vllm) or reuse a running server with `--skip-server-setup`; the endpoint defaults to localhost:1053 and is set by LOCAL_SERVER_ENDPOINT and LOCAL_SERVER_PORT in `.env`. [source]
- For a remote or existing OpenAI-compatible server, `REMOTE_OPENAI_BASE_URL`, `REMOTE_OPENAI_API_KEY` and an optional local tokenizer path are read from `.env`; `bfcl evaluate` then scores saved responses. [source]
- The harness's own backends are CUDA-oriented (sglang needs SM 80+ GPUs); a Mac run must go through `--skip-server-setup` against llama-server, mlx-lm or Ollama. [source]
- 2026: quantization claims moved from perplexity to task benchmarks; the bartowski study (already recorded) used BFCL, and a Neo-agent case study ran BFCL on Qwen3.6-27B at BF16, Q8_0 and Q4_K_M. [source]
- Sample size: the Neo study used 400 BFCL samples and 200 samples for HumanEval and HellaSwag; its own text attributes Q8_0 beating Q4_K_M on HellaSwag-like orderings to possible variance. [source]
- Wall time: BF16 averaged 37 s per BFCL sample, over four hours for 400, so long evals need checkpoint-and-resume (the study resumed at samples 139 and 152). [source]
- Summary columns count unevaluated BFCL categories as zero, so a partial local run under-reports the aggregate unless categories are compared one by one. [source]
- Does 4-bit hurt tool calls? Docker (qwen3:8B, F16 0.933 vs Q4_K_M 0.919; qwen3:14B Q6_K 0.943 vs Q4_K_M 0.971) saw "no significant difference". The Neo study saw BFCL 63.25 (BF16) vs 63.00 vs 63.00 but code generation fell 56.10 to 50.61 (HumanEval). bartowski's Q2_K result (already recorded) shows a 28-point BFCL loss at 2-bit without imatrix. These agree if loss only appears below roughly 4 bits or in generation tasks, but no source tests that boundary. [source]
- Docker's Q4_K_M beating Q6_K at 14B contradicts monotone quality-with-bits; with 3,570 tests spread across 21 models, per-cell sample size is small and the gap is likely noise. [source]
- BFCL multi-turn and agentic categories on a 4-bit local model: no source reports them. [source]
- Whether the Neo BFCL score used the model's native tool parser or a prompt-mode handler. [source]
- A three-way run of Qwen3.6-27B scored BFCL 63.25% at BF16, 63.00% at Q8_0 and 63.00% at Q4_K_M over 400 samples. [source]
- The same run scored HumanEval 56.10% (BF16), 52.44% (Q8_0), 50.61% (Q4_K_M) and HellaSwag 86%, 85%, 84%. [source]
- In that run the BF16 variant averaged 37 seconds per BFCL sample and the files were about 54 GB (BF16), 29 GB (Q8_0) and 17 GB (Q4_K_M). [source]
- Docker saw "no significant difference" in tool-calling behavior between quantized and non-quantized variants in all pairs it tried. [source]
- Docker F1: qwen3:8B-F16 0.933 versus Q4_K_M 0.919; qwen3:14B Q6_K 0.943 versus Q4_K_M 0.971; llama3.1:8B F16 0.835 versus Q4_K_M 0.793. [source]
- The BFCL harness's local backends are vllm (default) and sglang, and sglang supports only GPUs with SM 80 or newer. [source]
- `bfcl generate ... --skip-server-setup` reuses an already running OpenAI-compatible server at localhost port 1053 by default. [source]
- BFCL's web_search category calls SerpAPI and needs a key in `.env`; its summary columns treat unevaluated categories as 0. [source]
- BFCL's `--enable-lora`, `--lora-modules` and `--max-lora-rank` flags work only with the vllm backend. [source]
- The Oleksandr Iegorov write-up found xLAM-2-3b-fc-r (BF16) emitted correct tool JSON but looped on the same call in an agent, and moved to an 8B W4A16 GPTQ quant (about 5.4 GB versus 16 GB) so the agent could hold tool schemas and planning state. [source]
- No source found measures BFCL on an MLX or llama.cpp Metal quantization with Metal-specific kernels; all quantization-versus-BFCL evidence is CUDA or unspecified hardware. [source]
Children
- No children recorded.