Tool-call repair proxies and self-healing parsers
Parent: Mac local LLMs: Chat templates, reasoning and tool calling · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Harness repair loop (llmconfigurator): on a parse failure, re-prompt with the raw output plus the schema and ask only for corrected JSON; do not count it against the loop's iteration cap; stop after two failed repairs.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Harness repair loop (llmconfigurator): on a parse failure, re-prompt with the raw output plus the schema and ask only for corrected JSON; do not count it against the loop's iteration cap; stop after two failed repairs. [source]
- Validation repair: check arguments semantically (path exists and is inside the root, enum membership, required fields) and return the error as a tool result. [source]
- Unsloth Studio markets "self-healing tool calling" and an API endpoint that Claude Code and Codex can use; the pages give accuracy-gain percentages but no mechanism. [source]
- 2026: self-healing moved from community proxies to a packaged local-server feature (Unsloth Studio, llama.cpp-based, macOS supported). [source]
- A well-formed call with a hallucinated path executes; a parse error does not, so parse-only repair can raise risk. [source]
- Lowering temperature to 0.0-0.2 on action turns is advised; Unsloth's own guide gives 0.7 for GLM 4.7 and 0.15 for Devstral, so the advice is model-specific. [source]
- A malformed-call rate above about one in twenty is called a configuration problem (template, constraint, tool count), not bad luck. [source]
- Repair versus prevention: llmconfigurator ranks native template tool calling first, grammar constraint second, tool-count reduction third, and repair last; Unsloth sells repair as the headline feature. Neither side measures the other. [source]
- Unsloth quotes "50%" in one place and "30% to 80% more accurate" in another for the same feature, with no benchmark named. [source]
- No source defines what Unsloth's self-healing actually changes (re-parse, constrained retry, or template fix). [source]
- No independent measurement of repair-proxy recovery rates on Mac-served models. [source]
- Unsloth Studio advertises "self-healing tool calling" for local GGUF and safetensor models, powered by llama.cpp and available on macOS without a GPU. [source]
- Unsloth says its Studio "auto-fixes malformed or broken tool-calls by 50%". [source]
- Unsloth's studio page claims tool calls across all models are "30% to 80% more accurate", and its overview page says "50% more accurate". [source]
- Unsloth Studio can act as an API endpoint so Claude Code and Codex use local Qwen and Gemma models with its self-healing tool calling. [source]
- The Unsloth tool-calling guide lists Devstral's suggested temperature as 0.15 and GLM 4.7's as 0.7 with top_p 1.0. [source]
- llmconfigurator advises re-prompting a failed tool call with the raw output and schema, outside the loop's iteration cap, and capping repair at two attempts. [source]
- llmconfigurator advises validating tool arguments (path exists, enum members, required fields) because a valid call with a hallucinated path executes. [source]
- llmconfigurator treats a malformed-call rate above about 1 in 20 as a configuration problem and tells readers to log every malformed call. [source]
- llmconfigurator says a chat template with no tool tokens leaves a model on the prompted-JSON path, and that `ollama show --modelfile` reveals the template. [source]
- llmconfigurator lists three tool-call mechanisms: native template tool calling, prompted JSON, and edit formats (Aider's diff/search-replace). [source]
- A llama-server parser capturing an extra newline in a reasoning block pushed Step 3.7 Flash into "Actually..." self-corrections and loops in long agentic sessions (Hacker News account of a GitHub issue). [source]
- Vito Rallo reports that Ollama's Anthropic adapter returned text that looked like tool calls instead of `tool_use` blocks for qwen3.5:9b, qwen3.5:35b-a3b and glm-4.7-flash, and that vllm-mlx returned real blocks. [source]
Corrections and disagreements
- CONTRADICTS local-llm-server-as-a-coding-agent-backend-on-mac.md (Ollama >= 0.14 serves /v1/messages natively): the Rallo article (2026-04-04) says Ollama tool_use blocks do not survive translation; it is one reporter and a single date, so treat it as an unresolved conflict. [source]
Children
- No children recorded.