<!-- llms-explorer concept facts · https://llms-explorer.com/tree/tool-call-repair-proxies-and-self-healing-parser/ · pack 2026-10-05 · ~1390 tokens -->

# Tool-call repair proxies and self-healing parsers

> Harness repair loop (llmconfigurator): on a parse failure, re-prompt with the raw output plus the schema and ask only for corrected JSON; do not count it against the loop's iteration cap; stop after two failed repairs.

Parent: [Mac local LLMs: Chat templates, reasoning and tool calling](https://llms-explorer.com/tree/mac-local-llms-chat-templates-reasoning-and-tool-calling/) · 2 facets · 24 facts · page: https://llms-explorer.com/tree/tool-call-repair-proxies-and-self-healing-parser/

## Facts

- Harness repair loop (llmconfigurator): on a parse failure, re-prompt with the raw output plus the schema and ask only for corrected JSON; do not count it against the loop's iteration cap; stop after two failed repairs. — source: `asserted`
- Validation repair: check arguments semantically (path exists and is inside the root, enum membership, required fields) and return the error as a tool result. — source: `asserted`
- Unsloth Studio markets "self-healing tool calling" and an API endpoint that Claude Code and Codex can use; the pages give accuracy-gain percentages but no mechanism. — source: `asserted`
- 2026: self-healing moved from community proxies to a packaged local-server feature (Unsloth Studio, llama.cpp-based, macOS supported). — source: `asserted`
- A well-formed call with a hallucinated path executes; a parse error does not, so parse-only repair can raise risk. — source: `asserted`
- Lowering temperature to 0.0-0.2 on action turns is advised; Unsloth's own guide gives 0.7 for GLM 4.7 and 0.15 for Devstral, so the advice is model-specific. — source: `asserted`
- A malformed-call rate above about one in twenty is called a configuration problem (template, constraint, tool count), not bad luck. — source: `asserted`
- Repair versus prevention: llmconfigurator ranks native template tool calling first, grammar constraint second, tool-count reduction third, and repair last; Unsloth sells repair as the headline feature. Neither side measures the other. — source: `asserted`
- Unsloth quotes "50%" in one place and "30% to 80% more accurate" in another for the same feature, with no benchmark named. — source: `asserted`
- No source defines what Unsloth's self-healing actually changes (re-parse, constrained retry, or template fix). — source: `asserted`
- No independent measurement of repair-proxy recovery rates on Mac-served models. — source: `asserted`
- Unsloth Studio advertises "self-healing tool calling" for local GGUF and safetensor models, powered by llama.cpp and available on macOS without a GPU. — [source](https://unsloth.ai/docs/new/studio/chat)
- Unsloth says its Studio "auto-fixes malformed or broken tool-calls by 50%". — [source](https://unsloth.ai/docs/new/studio/chat)
- Unsloth's studio page claims tool calls across all models are "30% to 80% more accurate", and its overview page says "50% more accurate". — [source](https://unsloth.ai/docs/new/studio)
- Unsloth Studio can act as an API endpoint so Claude Code and Codex use local Qwen and Gemma models with its self-healing tool calling. — [source](https://unsloth.ai/docs/new/studio)
- The Unsloth tool-calling guide lists Devstral's suggested temperature as 0.15 and GLM 4.7's as 0.7 with top_p 1.0. — [source](https://unsloth.ai/docs/basics/tool-calling-guide-for-local-llms)
- llmconfigurator advises re-prompting a failed tool call with the raw output and schema, outside the loop's iteration cap, and capping repair at two attempts. — [source](https://llmconfigurator.com/en/guides/coding-agents/tool-calling-local-models)
- llmconfigurator advises validating tool arguments (path exists, enum members, required fields) because a valid call with a hallucinated path executes. — [source](https://llmconfigurator.com/en/guides/coding-agents/tool-calling-local-models)
- llmconfigurator treats a malformed-call rate above about 1 in 20 as a configuration problem and tells readers to log every malformed call. — [source](https://llmconfigurator.com/en/guides/coding-agents/tool-calling-local-models)
- llmconfigurator says a chat template with no tool tokens leaves a model on the prompted-JSON path, and that `ollama show --modelfile` reveals the template. — [source](https://llmconfigurator.com/en/guides/coding-agents/tool-calling-local-models)
- llmconfigurator lists three tool-call mechanisms: native template tool calling, prompted JSON, and edit formats (Aider's diff/search-replace). — [source](https://llmconfigurator.com/en/guides/coding-agents/tool-calling-local-models)
- A llama-server parser capturing an extra newline in a reasoning block pushed Step 3.7 Flash into "Actually..." self-corrections and loops in long agentic sessions (Hacker News account of a GitHub issue). — [source](https://news.ycombinator.com/item?id=49402232)
- Vito Rallo reports that Ollama's Anthropic adapter returned text that looked like tool calls instead of `tool_use` blocks for qwen3.5:9b, qwen3.5:35b-a3b and glm-4.7-flash, and that vllm-mlx returned real blocks. — [source](https://medium.com/@vito.rallo/running-claude-code-with-local-llms-all-lies-until-now-3e9a0084dfe1)

## Corrections and disagreements

- CONTRADICTS local-llm-server-as-a-coding-agent-backend-on-mac.md (Ollama >= 0.14 serves /v1/messages natively): the Rallo article (2026-04-04) says Ollama tool_use blocks do not survive translation; it is one reporter and a single date, so treat it as an unresolved conflict. — [source](https://medium.com/@vito.rallo/running-claude-code-with-local-llms-all-lies-until-now-3e9a0084dfe1)
