Every Token-Saving Strategy in My Stack, With Sources and Numbers

Published

Project article; see sources and editorial standards.

Compiled 2026-09-27 from four read-only research passes over my agent-memory corpus (~/.llms, auto-memory, .remember/ logs), the Claude Code harness config in ~/.claude, the LLM-infrastructure repos in ~/dev, and the application repos plus the local semantic hub. Numbers carry a source (path:line, commit or PR) where one exists; figures from the 2026-09-27 trim come from that day’s commands, PRs and routing probes. The paths point into my own, mostly private, repos and memory files: they show where each number came from, not something a reader can open. Where a number is a plugin’s own marketing claim, or an estimate rather than a measurement, the text says so.


Abstract

I collected the token-saving techniques recorded in my stack and linked them to code, configuration or run notes. This is a September 27 snapshot: current-looking settings and counts below refer to that date. Later checks can confirm an implementation without reproducing a historical saving. The private replay inputs, routing-probe outputs and several historical binary snapshots are not published, so their measured claims remain author-recorded results. The result is 57 sections, several of which hold more than one technique. They fall into six groups:

  1. Always-on harness overhead: what Claude Code loads before I type anything.
  2. Skill-listing engineering: folding hundreds of standalone skills into hubs, so referenced content stays reachable with fewer listing entries when the loader does not independently discover nested entrypoints.
  3. Avoiding the call: caching, replay, coalescing, hash gates, fast paths.
  4. Moving work to cheaper or local models: Batches API, Ollama, scout models, skill distillation.
  5. Shrinking what enters context: dedup funnels, retrieval, llms-small digests, byte caps.
  6. Workflow discipline: slim subagents, fan-out hygiene, when not to delegate.

The largest recorded savings came from reducing what loaded into context.


How this list was built


Part 1 — Always-on harness overhead

Everything in this part is overhead the harness adds to every session, every request or every reply. Prompt caching makes a stable prefix cheap after the first turn. It still fills the context window and pushes compaction earlier, and output tokens are never cached.

1.1 The skill-listing budget: skillListingBudgetFraction and skillListingMaxDescChars

1.2 What the listing actually shows: about 250 characters

With skillListingMaxDescChars: 300 set, the listing the model receives ends each description at about 250 characters with an ellipsis; I read that directly from the listing text on 2026-09-27. The same day, trigger words I had just added to a hub (a router skill; see section 1.3) sat at character 700, well past that point, and a routing probe never saw them.

Rule: put the routing vocabulary right after TRIGGER:, inside the first ~250 characters. After I moved the firecrawl hub’s workflow list to the front, a probe that had previously been routed to a synced plugin started reaching firecrawl/references/firecrawl-company-directories.md.

1.3 skillOverrides: "off" versus "name-only"

1.4 Description caps for hubs: 1000 and 1536

1.5 Plugins: enabled by allow-list

1.8 Capped hook injections

1.9 Output style

1.10 Subagent model routing

1.11 Prompt caching as the floor


Part 2 — Skill-listing engineering

2.1 Hub-and-spoke folding

2.2 Auto-tiering, like S3 Intelligent-Tiering for skills

2.3 Deduplicating skill directories

2.4 Untangling the config repo from the skills repo

This is the least obvious item on the list.

2.5 Compact skill indexes for other agents

2.6 Routing probes as the test

Link checks do not prove routing. Each fold is verified by spawning a fresh Haiku subagent with a plain-language question and no hints, and checking which references/*.md it actually opens. Gate: 80% or more.

2.7 A routing index instead of reading every SKILL.md

2.8 A ceiling on SKILL.md bodies

skill-optimizer Pass J flags any SKILL.md body over 10,000 tokens and moves a section to references/<name>.md, leaving a one-paragraph pointer. Nothing is deleted; it loads only when routing sends the agent there. One audit found 3 oversized skills, for example telemetry-pipeline at about 15.7k tokens with 2 sections moved out (skills/SKILLS-OPTIMIZATION-GUIDE.md:383; SKO_OFFLINE_REPORT.md:18-22).

2.9 Deterministic checks before LLM audits

skill_optimizer_offline.py runs regex, YAML and character-count checks across the whole library at zero token cost and reserves the LLM /sko loop for skills that need content judgment. Across 141 skills it fixed 7 real defects and avoided about $308 of one-time /sko runs (about $2.18 per skill at Opus rates) (skills/SKO_OFFLINE_REPORT.md:1-38).


Part 3 — Avoiding the call entirely

3.1 Exact-match response cache: llm-cache-proxy

3.2 Partial-key cache tiers

3.3 Request coalescing

Identical requests that are in flight at the same time share one upstream fetch, and the extra responses carry x-cache: HIT-COALESCED. The live fidelity test confirms that N concurrent identical calls make exactly 1 upstream call (README.md:316-325).

3.4 Skip unchanged inputs: mtime plus sha256 gates

3.5 Fast paths that skip the crawl

3.6 Gate the paid call

net-dns-monitor diagnoses network incidents offline first and calls a model only when that fails:

3.7 Poll less, trigger once, fetch nothing live


Part 4 — Moving work to cheaper or local models

4.1 Batches API for bulk extraction

llm_extractor.py sends one Haiku extraction per session through client.messages.batches.create (“50% cost, async”):

Together with semantic dedup, the memory pyramid records a claimed 10–25× compression, with no before-and-after token counts behind it. 157 sessions became about 2.4k records; records are atomic facts, so their count rises even as the text shrinks (.remember/archive.md:4; llm_extractor.py:7,43-45,171-227).

4.2 An extraction chain with a local-only option

4.3 Local embeddings everywhere

No repo in the infrastructure set calls a paid embedding API:

4.4 A weighted LAN embedding pool

4.5 Local models for grounding and coding

4.6 Scout models and skill distillation

4.7 The hub answers with a local model

hub_ask federates search across the global hub and answers with a weighted pool of local Ollama hosts (default qwen3:8b), never a paid API. It sets think=False because the reasoning trace otherwise dominates latency and output length. --retrieve-only skips generation and returns the fused, reranked hits; if the model is down, ask() returns the hits with an empty answer instead of failing (.global-ai-hub/scripts/semantic_ops/llm.py:1-23; ask.py:1-6).

4.8 Cheapest model first, escalate on hard cases


Part 5 — Shrinking what enters context

5.1 The funnel before generative extraction

5.2 Retrieval instead of reading

5.3 Distill a source only after it earns it

/dr counts touches per host. On the 3rd or 4th touch it runs /distill-offline in corpus mode, and from then on it reads the distilled list. After a re-crawl, --diff re-distills only the changes. A host touched once or twice “never earns back the distill pass’s own cost.” The break-even point is labeled “computed, not measured” (mdb-context-hub/memory.md:742,751).

5.4 The four-file llms family

5.5 Hard byte caps

Cap Where Value
llms-small ceiling (“Cursor-stability ceiling”) export_llms.py SMALL_MAX_CHARS 200,000 chars (~50k tokens)
Index section split INDEX_SPLIT_BYTES / PART_PAGES 10,000 B, 60 pages per part
Related-concept links /lca (concept abstractor) --max-related 20
Per-skill context pulled by /sync-skills commands/sync-skills.md:25 50,000 chars
Oversized memory.md distill_context.py tail read last 20,000 B
Per-file text before hashing or embedding hub_lib.max_content_chars 4,000 chars
SDK output llmsx-js DEFAULT_MAX_TOKENS 4,096
Hub answer context ask.py CONTEXT_CHARS / PER_HIT_CHARS 6,000 chars total, 700 per hit, top 8 of 50 reranked
Codebase search results hub_search_codebase n_results 5 by default
Keyword index vs embedding, per file keyword_max_chars / max_content_chars 200,000 vs 4,000 chars
Log evidence sent to a model net-dns-monitor MAX_OUTBOUND_LOG_LINES last 200 lines × 300 chars
Output per incident diagnosis net-dns-monitor max_tokens 512
Output per mdb-tam call type maxTokens 120 (one sentence), 512 (dedup JSON), 1,024 (live recommender), 4,096 (reports)
Prompt attachments mdb-tam normalizePromptAttachmentText 30,000 chars each, 8 max, 120,000 total
MCP result counts mdb-tam / mdb-case-assistant limit clamps 100, 50, 500, 1,000 by tool
Diff under review prompt-deep-optimizer worked example most recent 1,500 lines
External fact checks story-context skill 5 per run

5.6 Prefilter and fold before embedding

5.7 Retrieval prefiltering in /lca

5.8 Structured claims instead of stitched prose (/dr v3)

5.9 Memory pyramid and routed memory

5.10 A token-count header so agents can budget before fetching

The llms-explorer site sends X-Markdown-Tokens on every Markdown twin, computed with the same len//4 estimator the manifests use. The header had to move from Cloudflare _headers into a Pages Function: it needed 103 rules, and _headers allows 100 (site/functions/_middleware.ts).

5.11 Designing prompts for the cache

5.12 An overflow fixed by stripping and batching

About 3,000 roadmap items in one ranking prompt produced a 1.78M-token JSON envelope against a 1M-token window. The fix removed each item’s ~600-character signed preview URL, which “adds zero ranking signal”, kept it in a local id-to-url map, and reattached it after scoring. It then ranked in batches of 200 items, 3 at a time, each “comfortably under 100k tokens” (mdb-tam commit dfe9879).

5.13 Bounding what an app pulls into context

5.14 Tell agents what a file is before they open it


Part 6 — Workflow discipline

6.1 Slim headless subagents

6.2 Workers write files, not replies

In COG-second-brain, a worker whose output would be 2K tokens or more writes it to /tmp/{task-slug}-{context}.md and returns only a status and the path. Parallel workers never see each other’s raw output, only digested context (COG-second-brain/CLAUDE.md:79-100).

6.3 Fan-out hygiene: never pay twice

6.4 When not to fan out

In mdb-case-assistant, 6 of 10 delegated subagents failed (connection closed or a 600 s stall), mostly after finishing analysis but before writing output. The default flipped to working inline, because delegation’s expected cost (wasted turns and context) had passed the cost of doing the work directly (work-inline-not-subagents.md:15-38).

6.5 Search before reading

6.6 Model effort and prompting

A benchmark earlier on this blog reported a larger score difference between visible-working prompt regimes than between neutral model tiers. The prompt-regime difference came from one multi-step arithmetic task. Each cell contained one response, so the result does not establish where more model spend or longer reasoning improves other work.

6.7 Bound the reviewers

6.8 Count real tokens

skills-tui reads the cost and token counts that claude --output-format json reports and falls back to a chars/4 estimate only when the CLI reports nothing (skills_tui/core/cost.py:1-39). The proxy README makes the same point from the other side: a cached reply still carries usage numbers, so only the proxy’s own ledger shows what was saved.


What measured out as not worth it (or backfired)

Idea What happened Source
Response-cache proxy for interactive Claude Code ≤0.31% of tokens saveable; would also switch Max-plan OAuth to API billing replay simulation; proxy-a.mjs:273-274
skillListingBudgetFraction: 0.04 ~66k fewer listing tokens than at 0.12 (mostly cached), but hid 321 of 480 skills, cut alphabetically reference_skill_listing_budget_fraction.md
name-only on router hubs Questions for skills behind those hubs were misrouted (probe failed twice) token-trim, 2026-09-27
Trigger words appended at the end of a description Invisible past the ~250-character listing prefix token-trim, 2026-09-27
Routing probes on a tiered tree Two test reads promoted a spoke to hot, adding it back to the listing project-token-trim-2026-09-27.md
Explanatory output style plus caveman One added output that the other tried to remove token-trim, 2026-09-27
Delegating everything 6 of 10 subagents died mid-task; switched to inline work-inline-not-subagents.md
0-byte prompt_cost_hook.py Still spawns a process every prompt settings.json:786-793

The checklist I would apply to a new setup

  1. Cap each skill description’s length before lowering the listing budget fraction. Put trigger words in the first ~250 characters.
  2. Fold sibling skills into hubs. Never make a hub name-only. Test routing with blind probes and keep their test traffic distinguishable from real access history.
  3. Enable plugins by allow-list. Disconnect connectors you don’t use; their tool names and instruction blocks cost tokens every session.
  4. Keep global CLAUDE.md short. Mandatory per-reply footers are uncached output.
  5. Cap every hook injection by count, score and characters.
  6. Give subagents a cheaper default model, and run fan-outs headless without the skill listing when they don’t need skills.
  7. Retrieve, don’t read: an FTS5 keyword index for exact tokens, vectors for meaning, llms-small before llms-full.
  8. Dedup mechanically before generative extraction, and have models emit JSON while scripts render prose.
  9. Push bulk extraction to the Batches API or a local model.
  10. Put stable instructions first and volatile input last, so prompt caching can hold the prefix.
  11. Gate paid calls behind cheap offline checks, cap max_tokens per call type, and choose SDK retries deliberately when repeated processing is unacceptable.
  12. Cache responses only where requests actually repeat (evals, CI, reruns), and measure your hit rate on real traffic before deploying a cache.

Lessons