# LLMS-Explorer — facts

> Source-anchored units extracted from the docs: 517 across 53 pages. Each line ends in the page URL and anchor it came from.

<!-- generated 2026-09-02 by site/tools/build_llms.py from the .md twin of every page in this site's build output -->

## The attribute rubric
<https://llms-explorer.com/reference/attributes/>

- [parameter] I1: Attribute=Exactly one H1 naming the site/product (not a page); Applies=index, facts, family; Measure=deterministic; Bar=1 H1; title = product/site; Miss=High — https://llms-explorer.com/reference/attributes/#1-identity-and-shape
- [parameter] I2: Attribute=Blockquote summary immediately after H1, 1–3 sentences, self-contained; Applies=index, family; Measure=deterministic + judgment; Bar=present; says what the thing is and who it is for; Miss=Medium — https://llms-explorer.com/reference/attributes/#1-identity-and-shape
- [parameter] I3: Attribute=Free-form info before the first H2 (how to read this file, versions, languages); Applies=index, family; Measure=judgment; Bar=only if it changes how a reader should use the links; Miss=Low — https://llms-explorer.com/reference/attributes/#1-identity-and-shape
- [parameter] I4: Attribute=Sections are H2 only; each is a link list; no prose after the first H2 except list notes; Applies=index, family; Measure=deterministic; Bar=no H3+, no stray paragraphs; Miss=Medium — https://llms-explorer.com/reference/attributes/#1-identity-and-shape
- [parameter] I5: Attribute=Link entries match `- [name](url)` + optional `: notes`; Applies=index, family; Measure=deterministic; Bar=100% of list items; Miss=High if <90%, else Medium — https://llms-explorer.com/reference/attributes/#1-identity-and-shape
- [parameter] I6: Attribute=Kind is unambiguous from the first 20 lines (index vs full vs facts) — a full file is never served as an index; Applies=all; Measure=deterministic; Bar=grammar detected with one candidate; Miss=High — https://llms-explorer.com/reference/attributes/#1-identity-and-shape
- [parameter] N1: Attribute=Two hops: index → page (or family → index → page); no index links a bare directory of more indexes; Applies=index, family; Measure=deterministic (link targets); Bar=≤2 hops to any page; Miss=High — https://llms-explorer.com/reference/attributes/#2-navigation
- [parameter] N2: Attribute=Section design mirrors how users ask (task/topic groups), not the URL tree or an alphabet; Applies=index; Measure=judgment; Bar=≥80% of sections are task/topic named; Miss=Medium — https://llms-explorer.com/reference/attributes/#2-navigation
- [parameter] N3: Attribute=Ordering by expected query frequency: quickstart/auth/reference/errors first; the first 20% of links should answer 80% of questions; Applies=index; Measure=judgment + agent test; Bar=hot pages in the first section; Miss=Medium — https://llms-explorer.com/reference/attributes/#2-navigation
- [parameter] N4: Attribute=`## Optional` holds only skippable material (changelog, legal, old posts, appendices); it is the last section; Applies=index; Measure=deterministic + judgment; Bar=last; no reference/pricing inside; Miss=Medium — https://llms-explorer.com/reference/attributes/#2-navigation
- [parameter] N5: Attribute=Every page the source publishes that a reader would need is reachable (coverage); Applies=index; Measure=deterministic vs source page list; Bar=≥95% of `reference`+`guide` pages linked; Miss=High if <80% — https://llms-explorer.com/reference/attributes/#2-navigation
- [parameter] N6: Attribute=No dead ends: each link resolves (200, markdown or `.md` twin), no redirect to an HTML app shell; Applies=index, family; Measure=deterministic (`--check-links`); Bar=0 dead links; Miss=High — https://llms-explorer.com/reference/attributes/#2-navigation
- [parameter] N7: Attribute=Cross-cutting material (errors, auth, glossary) linked once, not once per section; Applies=index, family; Measure=judgment; Bar=no duplicate targets; Miss=Low — https://llms-explorer.com/reference/attributes/#2-navigation
- [parameter] D1: Attribute=Every link carries a description; Applies=index, family; Measure=deterministic; Bar=100%; Miss=Medium (High if <60%) — https://llms-explorer.com/reference/attributes/#3-descriptions
- [parameter] D2: Attribute=Description says what the reader FINDS there, with the exact tokens (flags, env vars, error strings) — not a restated title; Applies=index; Measure=judgment; Bar="Authentication docs." fails; "API key creation, OAuth scopes, token rotation. Required before any call." passes; Miss=Medium — https://llms-explorer.com/reference/attributes/#3-descriptions
- [parameter] D3: Attribute=Length 10–25 words; no trailing ellipsis from truncation; Applies=index; Measure=deterministic; Bar=95% within band; Miss=Low — https://llms-explorer.com/reference/attributes/#3-descriptions
- [parameter] D4: Attribute=No duplicate descriptions across links; Applies=index; Measure=deterministic; Bar=0 duplicates; Miss=Medium — https://llms-explorer.com/reference/attributes/#3-descriptions
- [parameter] D5: Attribute=Descriptions are extractive or verified — model-written ones audited against the page; Applies=index; Measure=judgment (sampled); Bar=sample of 10: 0 hallucinated claims; Miss=High — https://llms-explorer.com/reference/attributes/#3-descriptions
- [parameter] D6: Attribute=Family lines carry counts (pages, ~tokens) so a consumer can budget; Applies=family; Measure=deterministic; Bar=100% of product links; Miss=Medium — https://llms-explorer.com/reference/attributes/#3-descriptions
- [parameter] C1: Attribute=One declared page grammar, stated in a header comment; every page block parses; Applies=full; Measure=deterministic (`split_llms_full`); Bar=blocks parsed = blocks present; Miss=High — https://llms-explorer.com/reference/attributes/#4-content-fidelity-full-and-facts
- [parameter] C2: Attribute=Every page has a title and a resolvable source URL; Applies=full; Measure=deterministic; Bar=100%; Miss=High — https://llms-explorer.com/reference/attributes/#4-content-fidelity-full-and-facts
- [parameter] C3: Attribute=No navigation residue: "Documentation Index" blockquotes, `[Skip to content]`, MDX wrappers, `theme={null}` props; Applies=full; Measure=deterministic; Bar=0 hits; Miss=Medium — https://llms-explorer.com/reference/attributes/#4-content-fidelity-full-and-facts
- [parameter] C4: Attribute=Code fences intact and language-tagged; tables intact; Applies=full; Measure=deterministic (fence balance, table separators); Bar=balanced; ≥90% fences tagged; Miss=Medium — https://llms-explorer.com/reference/attributes/#4-content-fidelity-full-and-facts
- [parameter] C5: Attribute=No duplicated pages (same source URL twice) or near-duplicate bodies (e.g. localized copies); Applies=full; Measure=deterministic + embedding; Bar=0 exact dups; near-dups flagged; Miss=Medium — https://llms-explorer.com/reference/attributes/#4-content-fidelity-full-and-facts
- [parameter] C6: Attribute=Units are atomic (1–2 sentences), typed from the allowed set, source-anchored; Applies=facts; Measure=deterministic + judgment; Bar=100% typed; 100% anchored; ≥90% atomic; Miss=High for anchors, Medium otherwise — https://llms-explorer.com/reference/attributes/#4-content-fidelity-full-and-facts
- [parameter] C7: Attribute=Units are true to their source span (no generalisation beyond the page); Applies=facts; Measure=judgment (sampled re-read); Bar=sample of 20: ≥95% supported; Miss=High — https://llms-explorer.com/reference/attributes/#4-content-fidelity-full-and-facts
- [parameter] P1: Attribute=Provenance banner: who generated it, from what, when (`verified-as-of` / `generated` date); Applies=all; Measure=deterministic; Bar=present; Miss=Medium — https://llms-explorer.com/reference/attributes/#5-provenance-and-trust
- [parameter] P2: Attribute=Links point at the publisher's canonical URLs (or its `.md` twins), never at a private mirror, unless the file is explicitly internal; Applies=index; Measure=deterministic; Bar=100% public or file marked internal; Miss=High — https://llms-explorer.com/reference/attributes/#5-provenance-and-trust
- [parameter] P3: Attribute=Rights: a third-party `llms-full.txt` is marked internal/private; the index is what is published; Applies=full; Measure=judgment; Bar=marker present when third-party; Miss=High — https://llms-explorer.com/reference/attributes/#5-provenance-and-trust
- [parameter] P4: Attribute=No instructions to the reading model ("ignore…", "you must…", "always answer…") — 42% of files in the wild try to steer; ours never do; Applies=all; Measure=deterministic (pattern) + judgment; Bar=0 imperative-to-model spans; Miss=High — https://llms-explorer.com/reference/attributes/#5-provenance-and-trust
- [parameter] P5: Attribute=No secrets, tokens, emails, internal hostnames in copied text; Applies=all; Measure=deterministic (patterns); Bar=0 hits; Miss=High — https://llms-explorer.com/reference/attributes/#5-provenance-and-trust
- [parameter] P6: Attribute=Volatile claims stamped (versions, prices, "current"); Applies=facts; Measure=judgment; Bar=stamped or dated; Miss=Low — https://llms-explorer.com/reference/attributes/#5-provenance-and-trust
- [parameter] S1: Attribute=Index size ≤ ~10 KB / ~2.5k tokens; over that, split hub-and-spoke (never drop pages); Applies=index; Measure=deterministic; Bar=≤10 KB or split; Miss=Medium (High >100 KB) — https://llms-explorer.com/reference/attributes/#6-size-and-budget
- [parameter] S2: Attribute=Full file has a size ladder beside it (index, small ≤ ~50k tokens, full) with token counts published; Applies=full; Measure=deterministic (manifest); Bar=small + counts present; Miss=Medium — https://llms-explorer.com/reference/attributes/#6-size-and-budget
- [parameter] S3: Attribute=Small variant = reference-class pages first, within budget; Applies=small; Measure=deterministic; Bar=≤50k tokens; classes honoured; Miss=Medium — https://llms-explorer.com/reference/attributes/#6-size-and-budget
- [parameter] S4: Attribute=Facts file ≤ ~15% of the cleaned source prose (compression); Applies=facts; Measure=deterministic; Bar=ratio ≤0.15; Miss=Low (Medium >0.3) — https://llms-explorer.com/reference/attributes/#6-size-and-budget
- [parameter] S5: Attribute=Token estimate declared with its estimator (chars/4 etc.); Applies=manifest; Measure=deterministic; Bar=present; Miss=Low — https://llms-explorer.com/reference/attributes/#6-size-and-budget
- [parameter] S6: Attribute=No single page block > 200 KB without a note (changelogs); Applies=full; Measure=deterministic; Bar=flagged; Miss=Low — https://llms-explorer.com/reference/attributes/#6-size-and-budget
- [parameter] R1: Attribute=Keyword index exists for the facts/full text (FTS5 over units/chunks) and returns the exact-token queries (`CLAUDE_CODE_SYNC_SKILLS`, `--append-system-prompt`); Applies=facts, full; Measure=measured; Bar=10/10 exact-token probes hit; Miss=High — https://llms-explorer.com/reference/attributes/#7-retrieval-readiness
- [parameter] R2: Attribute=Vector index exists (`<key>__facts` collection) and the facts layer answers the golden questions better than raw; Applies=facts; Measure=measured (`query --layer`); Bar=golden score ≥ raw score; Miss=Medium — https://llms-explorer.com/reference/attributes/#7-retrieval-readiness
- [parameter] R3: Attribute=Anchors are stable (`#slug` of the heading) so a hit can be opened at the span; Applies=facts, full; Measure=deterministic; Bar=100% anchors resolve to a heading; Miss=Medium — https://llms-explorer.com/reference/attributes/#7-retrieval-readiness
- [parameter] R4: Attribute=Unit text carries the exact tokens in `keywords` so BM25 can find them; Applies=facts; Measure=deterministic; Bar=≥80% of units with a code/flag/env token have it in keywords; Miss=Medium — https://llms-explorer.com/reference/attributes/#7-retrieval-readiness
- [parameter] R5: Attribute=Agent test: an agent given ONLY the index answers N seeded questions by following ≤2 links; Applies=index; Measure=live agent test; Bar=≥8/10; Miss=High if <6/10 — https://llms-explorer.com/reference/attributes/#7-retrieval-readiness
- [parameter] R6: Attribute=Facts test: an agent given ONLY the facts file answers the same questions without opening pages; Applies=facts; Measure=live agent test; Bar=≥7/10; Miss=Medium — https://llms-explorer.com/reference/attributes/#7-retrieval-readiness
- [parameter] R7: Attribute=Every page in the index has ≥1 unit in the facts file (no silent gaps); Applies=index+facts; Measure=deterministic; Bar=≥95% pages covered; Miss=Medium — https://llms-explorer.com/reference/attributes/#7-retrieval-readiness
- [parameter] F1: Attribute=Family file links indexes, never pages; Applies=family; Measure=deterministic; Bar=100% targets are `llms.txt` files; Miss=High — https://llms-explorer.com/reference/attributes/#8-family-nesting
- [parameter] F2: Attribute=Each product line carries page + token counts and, where present, a facts link; Applies=family; Measure=deterministic; Bar=100%; Miss=Medium — https://llms-explorer.com/reference/attributes/#8-family-nesting
- [parameter] F3: Attribute=Shared material (errors, auth, glossary) appears once, in the family file; Applies=family; Measure=judgment; Bar=no duplication into products; Miss=Low — https://llms-explorer.com/reference/attributes/#8-family-nesting
- [parameter] F4: Attribute=The most-specific rule holds: a product's own index is authoritative for its pages; the family never restates them; Applies=family; Measure=judgment; Bar=no page links; Miss=Medium — https://llms-explorer.com/reference/attributes/#8-family-nesting
- [parameter] F5: Attribute=Family membership matches the concept tree / hub taxonomy it claims to represent; Applies=family; Measure=deterministic vs tree; Bar=100% of tree children present; Miss=Medium — https://llms-explorer.com/reference/attributes/#8-family-nesting
- [parameter] F6: Attribute=Root → family → product is discoverable by `Link: rel=describedby` from any file; Applies=family; Measure=deterministic (headers); Bar=header present; Miss=Low — https://llms-explorer.com/reference/attributes/#8-family-nesting
- [parameter] H1: Attribute=UTF-8, LF, no tabs in list lines, no trailing whitespace, single trailing newline; Applies=all; Measure=deterministic; Bar=clean; Miss=Hygiene (Low) — https://llms-explorer.com/reference/attributes/#9-hygiene-and-serving
- [parameter] H2: Attribute=`Content-Type: text/markdown; charset=utf-8` (or `text/plain`), HTTP 200, no redirect, no auth on the path; Applies=served; Measure=deterministic (HEAD); Bar=pass; Miss=High — https://llms-explorer.com/reference/attributes/#9-hygiene-and-serving
- [parameter] H3: Attribute=`Link: rel=describedby` on files; `rel=alternate type=text/markdown` on HTML pages; Applies=served; Measure=deterministic; Bar=present; Miss=Low — https://llms-explorer.com/reference/attributes/#9-hygiene-and-serving
- [parameter] H4: Attribute=`X-Markdown-Tokens` (or manifest tokens) available before fetch; Applies=served; Measure=deterministic; Bar=present; Miss=Low — https://llms-explorer.com/reference/attributes/#9-hygiene-and-serving
- [parameter] H5: Attribute=Regenerated by the build, not hand-maintained; a `generated` stamp newer than the source; Applies=all; Measure=deterministic (mtime/stamp); Bar=stamp ≥ source mtime; Miss=Medium — https://llms-explorer.com/reference/attributes/#9-hygiene-and-serving
- [parameter] H6: Attribute=Validator-clean on the community validators' strict rules where they do not contradict the spec; Applies=index; Measure=deterministic; Bar=0 High; Miss=Low — https://llms-explorer.com/reference/attributes/#9-hygiene-and-serving
- [parameter] H7: Attribute=Lighthouse agentic audit would not flag it (no 5xx on fetch); Applies=served; Measure=deterministic; Bar=200; Miss=Medium — https://llms-explorer.com/reference/attributes/#9-hygiene-and-serving
- [parameter] H8: Attribute=`manifest.json` present and consistent with the files (bytes, tokens, pages, units); Applies=export dir; Measure=deterministic; Bar=consistent; Miss=Medium — https://llms-explorer.com/reference/attributes/#9-hygiene-and-serving
- [parameter] Purpose: index (`llms.txt`)=orientation + navigation; full (`llms-full.txt`)=whole text in one fetch; facts (`llms-facts.txt`)=the checkable claims, each anchored — https://llms-explorer.com/reference/attributes/#10-the-three-kinds-side-by-side
- [parameter] Reader: index (`llms.txt`)=an agent deciding where to look; full (`llms-full.txt`)=a big-context agent or an indexer; facts (`llms-facts.txt`)=a retriever answering a question — https://llms-explorer.com/reference/attributes/#10-the-three-kinds-side-by-side
- [parameter] Unit: index (`llms.txt`)=link + description; full (`llms-full.txt`)=page block; facts (`llms-facts.txt`)=typed unit with source + anchor — https://llms-explorer.com/reference/attributes/#10-the-three-kinds-side-by-side
- [parameter] Size: index (`llms.txt`)=≤10 KB; full (`llms-full.txt`)=unbounded (ladder beside it); facts (`llms-facts.txt`)=≤15% of prose — https://llms-explorer.com/reference/attributes/#10-the-three-kinds-side-by-side
- [parameter] Judged mostly on: index (`llms.txt`)=N*, D*; full (`llms-full.txt`)=C1–C5, S*; facts (`llms-facts.txt`)=C6–C7, R*, P4 — https://llms-explorer.com/reference/attributes/#10-the-three-kinds-side-by-side
- [parameter] Tested by: index (`llms.txt`)=agent test (R5); full (`llms-full.txt`)=grammar round-trip (C1); facts (`llms-facts.txt`)=keyword + vector probes (R1–R2), facts test (R6) — https://llms-explorer.com/reference/attributes/#10-the-three-kinds-side-by-side
- [definition] The attribute rubric — Every attribute an llms file is judged on, with bars and severities. — https://llms-explorer.com/reference/attributes/#the-attribute-rubric

## Changelog: spec v1 to v2, and the hub pipeline
<https://llms-explorer.com/reference/changelog/>

- [parameter] Required elements: v1=H1 + blockquote + sections implied; v2=**H1 only** required; blockquote, prose and sections optional; Effect on an existing file=none required; the lint still scores a missing blockquote as Medium (I2) — quality, not validity — https://llms-explorer.com/reference/changelog/#the-spec-v1-v2-2026-08-10
- [parameter] Placement: v1=`/llms.txt` at the site root; v2=root **or any subpath**; a file covers the URLs under its path; **most-specific wins**; `/.well-known/` explicitly rejected; Effect on an existing file=enables families and split roots (`<section>/llms.txt`) — https://llms-explorer.com/reference/changelog/#the-spec-v1-v2-2026-08-10
- [parameter] Discovery: v1=none; v2=`Link: <…>; rel="describedby"` on files; `rel="alternate" type="text/markdown"` on HTML pages; as `<link>` or an HTTP header; Effect on an existing file=add the two headers ([usage](/reference/usage/) §1) — https://llms-explorer.com/reference/changelog/#the-spec-v1-v2-2026-08-10
- [parameter] Markdown twins: v1=`page.html.md`; v2=`page.html.md` **or** `page.md`; directories append `index.html.md` or `index.md`; Effect on an existing file=either form passes the twin probe (N6) — https://llms-explorer.com/reference/changelog/#the-spec-v1-v2-2026-08-10
- [parameter] `## Optional`: v1=mechanical: skippable when context is short; consumed by `llms_txt2ctx`; v2=a **convention** for secondary information; `llms_txt2ctx` and context-expansion removed from the proposal; Effect on an existing file=keep it last; build nothing that depends on it — https://llms-explorer.com/reference/changelog/#the-spec-v1-v2-2026-08-10
- [parameter] BOM: v1=—; v2=an optional byte-order mark is tolerated; Effect on an existing file=the lint strips it as hygiene (P14) — https://llms-explorer.com/reference/changelog/#the-spec-v1-v2-2026-08-10
- [parameter] Consumption model: v1=expand the whole file into context; v2="view or search the index, then follow links"; the index stays small; detail lives behind links; Effect on an existing file=the size ladder (small / full) becomes the producer's job — https://llms-explorer.com/reference/changelog/#the-spec-v1-v2-2026-08-10
- [parameter] Authoring guidance: v1=—; v2=concise language, informative link descriptions, no unexplained jargon, "test your file by asking an agent questions … giving it only your llms.txt"; Effect on an existing file=the agent test (R5, P12) is the spec's own test made numeric — https://llms-explorer.com/reference/changelog/#the-spec-v1-v2-2026-08-10
- [parameter] Acquire: V1 (to 2026-08-29)=trafilatura BFS crawl → banner mirror; V2 (from 2026-08-30)=the ladder: `llms-full.txt` → `llms.txt` + `.md` twins → `Accept: text/markdown` → docs API → structured crawl; the banner mirror stays the internal format — https://llms-explorer.com/reference/changelog/#the-hub-pipeline-v1-v2-2026-08-30
- [parameter] Clean: V1 (to 2026-08-29)=none (raw HTML → text); V2 (from 2026-08-30)=`docset_refine clean`: boilerplate lines, MDX → markdown, page classes (reference / guide / changelog / marketing / index) — https://llms-explorer.com/reference/changelog/#the-hub-pipeline-v1-v2-2026-08-30
- [parameter] Extract: V1 (to 2026-08-29)=`distill_offline.py bulk` — zero-LLM, output never consumed; V2 (from 2026-08-30)=`extract` (snippets, table rows → parameter, definitions, changelog → change; anchors to real headings) + `units` (local LLM, evidence rule) + `polish` — https://llms-explorer.com/reference/changelog/#the-hub-pipeline-v1-v2-2026-08-30
- [parameter] Export: V1 (to 2026-08-29)=none; V2 (from 2026-08-30)=`export_llms`: index (split over 10 KB) / full (Mintlify grammar) / small (≤ 200,000 chars) / facts / manifest; `topical`; `vocabulary` — https://llms-explorer.com/reference/changelog/#the-hub-pipeline-v1-v2-2026-08-30
- [parameter] Index: V1 (to 2026-08-29)=one raw vector layer; V2 (from 2026-08-30)=raw **and** facts vector layers, plus an FTS5 keyword layer per layer — https://llms-explorer.com/reference/changelog/#the-hub-pipeline-v1-v2-2026-08-30
- [parameter] Serve: V1 (to 2026-08-29)=`web-text-mirror --serve` (HTML); V2 (from 2026-08-30)=`llms_serve.py`: `/llms.txt`, `/d/<stem>/…`, `/m/<key>/…`, `/t/<slug>/…`, with the markdown headers — https://llms-explorer.com/reference/changelog/#the-hub-pipeline-v1-v2-2026-08-30
- [parameter] Gate: V1 (to 2026-08-29)=none; V2 (from 2026-08-30)=`llms_lint.py` (the deterministic passes) inside `docset_rollout cleanup`; `/ldo` for the model and live passes — https://llms-explorer.com/reference/changelog/#the-hub-pipeline-v1-v2-2026-08-30
- [parameter] Artifacts: V1 (to 2026-08-29)=`<stem>.pages/`, `_master.md`, `._distill_index.json`; V2 (from 2026-08-30)=`<stem>.reference/{pages.json, structured.jsonl, units.jsonl, all_units.jsonl}`, `<stem>.llms/` — https://llms-explorer.com/reference/changelog/#the-hub-pipeline-v1-v2-2026-08-30
- [parameter] 2026-08-30: `docset_refine` gains `clean / extract / units / polish / render / export`; the reference dir layout above — https://llms-explorer.com/reference/changelog/#dated-changes-to-the-hub-schema
- [parameter] 2026-08-30: `export_llms` writes the four-file ladder plus `manifest.json`; index split at 10,000 bytes; `PART_PAGES = 60` — https://llms-explorer.com/reference/changelog/#dated-changes-to-the-hub-schema
- [parameter] 2026-08-30: `llms_lint.py` ships the deterministic passes P0–P3, P5–P7, P9 and P14 and the `--json` CI output; `UNIT_RE` fixes the facts line grammar — https://llms-explorer.com/reference/changelog/#dated-changes-to-the-hub-schema
- [parameter] 2026-08-30: `llms_serve.py` sends `Content-Type: text/markdown`, `X-Markdown-Tokens`, `Link: rel="describedby"` — https://llms-explorer.com/reference/changelog/#dated-changes-to-the-hub-schema
- [parameter] 2026-08-30: `docset_refine topical` and `vocabulary`; tree nodes carry `slug` / `aliases`; `--register` writes `llmsFile` on a node — https://llms-explorer.com/reference/changelog/#dated-changes-to-the-hub-schema
- [parameter] 2026-08-31: `export_llms` honours `manifest.json["overrides"]` (`title`, `summary`, `section_order`, `note`) so hand inputs survive regeneration — https://llms-explorer.com/reference/changelog/#dated-changes-to-the-hub-schema
- [parameter] 2026-08-31: `llms_lint.py --kind vocabulary` lints the vocabulary line grammar — https://llms-explorer.com/reference/changelog/#dated-changes-to-the-hub-schema
- [definition] Changelog: spec v1 to v2, and the hub pipeline — What changed in the llms.txt proposal on 2026-08-10, rule by rule, with the effect on an existing file — and the dated changes to the hub schema behind this site. — https://llms-explorer.com/reference/changelog/#changelog-spec-v1-to-v2-and-the-hub-pipeline

## The concept tree: nodes, frontier, and how to read a node page
<https://llms-explorer.com/reference/concept-tree/>

- [parameter] `concept`: the node's name, and the string its parent and children link it by — https://llms-explorer.com/reference/concept-tree/#the-fields-on-a-node-page
- [parameter] `slug`: its URL segment; stable, and the key the API in step 3 will use — https://llms-explorer.com/reference/concept-tree/#the-fields-on-a-node-page
- [parameter] `parent`: the concept it hangs from — linked, unless the node is a root — https://llms-explorer.com/reference/concept-tree/#the-fields-on-a-node-page
- [parameter] `children`: the concepts it names, each either researched (linked) or frontier (greyed) — https://llms-explorer.com/reference/concept-tree/#the-fields-on-a-node-page
- [parameter] `aliases`: other names the same concept goes by; the filter matches these too — https://llms-explorer.com/reference/concept-tree/#the-fields-on-a-node-page
- [parameter] `researchedAt`: the date the research run that created the node finished — https://llms-explorer.com/reference/concept-tree/#the-fields-on-a-node-page
- [parameter] `sourcesCount`: how many sources that run read — https://llms-explorer.com/reference/concept-tree/#the-fields-on-a-node-page
- [parameter] `conceptsCount`: how many concepts that run identified under this one — https://llms-explorer.com/reference/concept-tree/#the-fields-on-a-node-page
- [parameter] `skillId`: the skill the research produced, when it produced one — https://llms-explorer.com/reference/concept-tree/#the-fields-on-a-node-page
- [parameter] `state`: `researched` for every node with a page; `frontier` only for a named child without one — https://llms-explorer.com/reference/concept-tree/#the-fields-on-a-node-page
- [definition] The concept tree: nodes, frontier, and how to read a node page — What the tree is, why frontier is derived rather than stored, what every field on a node means, and how the browser at /tree/ filters it. — https://llms-explorer.com/reference/concept-tree/#the-concept-tree-nodes-frontier-and-how-to-read-a-node-page
- [definition] What a node is — A node is one researched concept. The tree is stored as a **flat list of nodes linked by name** — each node names its parent and its children as strings, not as pointers — so a rename is a one-place edit and a reader can hold the whole file in mind. — https://llms-explorer.com/reference/concept-tree/#what-a-node-is
- [definition] Frontier is derived, never stored — A **frontier** concept on this site is a name that appears in some node's `childConcepts` and has no node of its own. It is computed on every build from the two sides of that comparison, never stored as a status: a stored status can disagree with the tree, and a derived one cannot. — https://llms-explorer.com/reference/concept-tree/#frontier-is-derived-never-stored
- [definition] How the filter works — The filter box on `/tree/` matches a substring against each concept **and its aliases**, and a branch survives if it or any descendant matches — so filtering hides non-matching branches without ever hiding the path to a hit. — https://llms-explorer.com/reference/concept-tree/#how-the-filter-works
- [definition] What is not here yet — Queueing a frontier concept for research, forking the tree, and attaching your own files to a node are per-user actions, and this site has no accounts yet. — https://llms-explorer.com/reference/concept-tree/#what-is-not-here-yet

## The directory and its grades
<https://llms-explorer.com/reference/directory/>

- [parameter] `A`: 0 High, 0 Medium — https://llms-explorer.com/reference/directory/#how-a-grade-is-derived
- [parameter] `B`: 0 High, 1–2 Medium — https://llms-explorer.com/reference/directory/#how-a-grade-is-derived
- [parameter] `C`: 0 High, 3 or more Medium — https://llms-explorer.com/reference/directory/#how-a-grade-is-derived
- [parameter] `D`: exactly 1 High — https://llms-explorer.com/reference/directory/#how-a-grade-is-derived
- [parameter] `F`: 2 or more High — https://llms-explorer.com/reference/directory/#how-a-grade-is-derived
- [parameter] `I`: Identity and shape — https://llms-explorer.com/reference/directory/#the-rubric-groups-on-a-score-card
- [parameter] `N`: Navigation — https://llms-explorer.com/reference/directory/#the-rubric-groups-on-a-score-card
- [parameter] `D`: Descriptions — https://llms-explorer.com/reference/directory/#the-rubric-groups-on-a-score-card
- [parameter] `C`: Content fidelity — https://llms-explorer.com/reference/directory/#the-rubric-groups-on-a-score-card
- [parameter] `P`: Provenance and trust — https://llms-explorer.com/reference/directory/#the-rubric-groups-on-a-score-card
- [parameter] `S`: Size and budget — https://llms-explorer.com/reference/directory/#the-rubric-groups-on-a-score-card
- [parameter] `R`: Retrieval readiness — https://llms-explorer.com/reference/directory/#the-rubric-groups-on-a-score-card
- [parameter] `F`: Family / nesting — https://llms-explorer.com/reference/directory/#the-rubric-groups-on-a-score-card
- [parameter] `H`: Hygiene and serving — https://llms-explorer.com/reference/directory/#the-rubric-groups-on-a-score-card
- [definition] The directory and its grades — What the directory measures, how the A–F grade is derived, and why the mirrored text is never republished. — https://llms-explorer.com/reference/directory/#the-directory-and-its-grades
- [definition] What the directory measures — One thing only: the output of `llms_lint` run over a copy of that site's `llms-full.txt`, with `kind="full"`. — https://llms-explorer.com/reference/directory/#what-the-directory-measures
- [definition] Which files are left out — Three exclusions, in the order they bite. — https://llms-explorer.com/reference/directory/#which-files-are-left-out
- [definition] How a grade is derived — The grade is arithmetic over the High and Medium counts, and nothing else. No weighting, no opinion, no manual override: — https://llms-explorer.com/reference/directory/#how-a-grade-is-derived
- [definition] The rubric groups on a score card — Each site page splits its High and Medium findings across the nine rubric groups, keyed by the first letter of the attribute id: — https://llms-explorer.com/reference/directory/#the-rubric-groups-on-a-score-card
- [definition] What the directory does not do — It does not republish anybody's text. The hub mirrors each file so it can be scored, and that copy stays in the hub: every directory page links the source's own file at the source's own URL. — https://llms-explorer.com/reference/directory/#what-the-directory-does-not-do
- [definition] How a site is added — The directory is generated, never hand-edited. `site/tools/gen_directory.py` reads the hub's catalog of known files and writes `src/data/directory.json`; the pages render that file. — https://llms-explorer.com/reference/directory/#how-a-site-is-added
- [definition] How a site is corrected or removed — Every entry names a real organisation and prints a public letter grade against it, so there is a way off the list and a way to fix a wrong one. — https://llms-explorer.com/reference/directory/#how-a-site-is-corrected-or-removed

## Ethos: what an llms file owes its reader
<https://llms-explorer.com/reference/ethos/>

- [parameter] An index — links and extractive descriptions: yes; it is a map of someone's public pages — https://llms-explorer.com/reference/ethos/#5-rights-are-explicit
- [parameter] A facts file — short anchored claims, each traceable: yes; quotation with attribution, bounded in length — https://llms-explorer.com/reference/ethos/#5-rights-are-explicit
- [parameter] Your own words — hand pages, essays, this site: yes — https://llms-explorer.com/reference/ethos/#5-rights-are-explicit
- [parameter] Third-party full text — a mirrored `llms-full.txt` of a site you do not own: served only to its owner, or under the internal marker; never on a public route — https://llms-explorer.com/reference/ethos/#5-rights-are-explicit
- [definition] Ethos: what an llms file owes its reader — Files are promises; generate, do not hand-edit; never instruct the reader; evidence is external; rights are explicit. — https://llms-explorer.com/reference/ethos/#ethos-what-an-llms-file-owes-its-reader
- [definition] 1. Files are promises — Every link resolves. Every fact is anchored to a heading that exists. — https://llms-explorer.com/reference/ethos/#1-files-are-promises
- [definition] 2. Generate, don't hand-edit — An llms file is an output. Its inputs are a mirror, a page list, extracted units, a concept tree, and a small overrides file (`title`, `summary`, `section_order`, `note`). — https://llms-explorer.com/reference/ethos/#2-generate-dont-hand-edit
- [definition] 3. Never instruct the reader — A docs file has no business telling a model what to say. The spec repository's issue #152 found 42.3% of a sample of wild files attempting exactly that ([evidence](/reference/evidence/)). — https://llms-explorer.com/reference/ethos/#3-never-instruct-the-reader
- [definition] 4. Evidence is external — A finding about an llms file cites something outside the file: the link check, the mirror span behind a unit, the probe result, the HTTP response. A finding with no external evidence is Low at most. — https://llms-explorer.com/reference/ethos/#4-evidence-is-external
- [definition] 5. Rights are explicit — Three tiers, and the tooling knows which is which: — https://llms-explorer.com/reference/ethos/#5-rights-are-explicit
- [definition] The test — If a stranger's agent, handed only the index, can answer eight of ten reasonable questions in two hops — below six is a High — and can check any facts line it relies on in one fetch, the file kept its promises. Nothing else on this site is a stronger claim than that. — https://llms-explorer.com/reference/ethos/#the-test

## Ecosystem evidence
<https://llms-explorer.com/reference/evidence/>

- [parameter] Feb→May 2025: Source=Chris Green; Sample=Majestic Million; Finding=15 → 105 valid files (~0.01%); ~100k crawl errors caveat[^4] — https://llms-explorer.com/reference/evidence/#2-adoption-measurements-dated
- [parameter] Jun 2025: Source=Originality.ai; Sample=3M+ sites; Finding=4,088 llms.txt[^5] — https://llms-explorer.com/reference/evidence/#2-adoption-measurements-dated
- [parameter] Jun 2025: Source=Rankability; Sample=Tranco top 1,000; Finding=0.3%[^6] — https://llms-explorer.com/reference/evidence/#2-adoption-measurements-dated
- [parameter] Jul 2025: Source=HTTP Archive (Burridge); Sample=top 10k; Finding=1.04% valid[^7] — https://llms-explorer.com/reference/evidence/#2-adoption-measurements-dated
- [parameter] Nov 2025: Source=SE Ranking; Sample=~300k domains; Finding=10.13% overall (9.88% low-traffic / 10.54% mid / 8.27% 100k+ visits)[^3] — https://llms-explorer.com/reference/evidence/#2-adoption-measurements-dated
- [parameter] Mar 2026: Source=Originality.ai via ppc.land; Sample=Fortune 500; Finding=7.4% (37/500)[^8] — https://llms-explorer.com/reference/evidence/#2-adoption-measurements-dated
- [parameter] May 2026: Source=Originality.ai; Sample=3M+ sites; Finding=36,120 llms.txt (8.8× YoY); llms-full.txt 23 → 2,463 (107×); ai.txt 397[^5] — https://llms-explorer.com/reference/evidence/#2-adoption-measurements-dated
- [parameter] May 2026: Source=Ahrefs; Sample=137,210 Ahrefs-Web-Analytics domains; Finding=28% publish a valid file (self-selected sample)[^1] — https://llms-explorer.com/reference/evidence/#2-adoption-measurements-dated
- [parameter] Jun 2026: Source=HTTP Archive (Burridge); Sample=top 1k / 10k / 100k / 1M; Finding=6.28% / 5.61% / 5.17% / 5.07% (~5.4× in 12 months)[^7] — https://llms-explorer.com/reference/evidence/#2-adoption-measurements-dated
- [parameter] Jun 2026: Source=Rankability; Sample=Tranco top 1,000; Finding=8.7% (87 files; 15 with llms-full.txt)[^6] — https://llms-explorer.com/reference/evidence/#2-adoption-measurements-dated
- [parameter] Ahrefs (2026-06-15): Window / sample=May 2026 logs, 137,210 domains; Finding=**97% of valid files got zero requests**; of requests, 96% bots, 77% of those non-AI (SEO auditors 21.7%); named AI bots 19.5%; AI training crawlers 5.3% (GPTBot 4.51%, ClaudeBot 0.8%); AI retrieval 1.1% (OAI-SearchBot 0.74%); **0 AI requests to non-existent files** (nobody probes speculatively); the `Claude-Code` UA… — https://llms-explorer.com/reference/evidence/#4-who-reads-server-log-studies
- [parameter] OtterlyAI (2026-02-05): Window / sample=90 days, one site; Finding=84 of 62,100 AI-bot requests hit /llms.txt (0.1%)[^15] — https://llms-explorer.com/reference/evidence/#4-who-reads-server-log-studies
- [parameter] Wislr (Feb–Mar 2026): Window / sample=48 days, one site; Finding=12,099 bot requests; robots.txt fetched hundreds of times (OAI-SearchBot 180, ClaudeBot 175); sitemap.xml too; **llms.txt 0**[^16] — https://llms-explorer.com/reference/evidence/#4-who-reads-server-log-studies
- [parameter] EZY Research (Apr–Jul 2026): Window / sample=83 sites, 12 weeks; Finding=robots vs llms: GPTBot 3,990/7, ClaudeBot 3,120/9, PerplexityBot 775/0, Googlebot 5,125/67, **Meta-ExternalAgent 172/193** (the only bot fetching it more than robots.txt)[^17] — https://llms-explorer.com/reference/evidence/#4-who-reads-server-log-studies
- [parameter] Hacker News thread (Feb 2026): Window / sample=anecdotal logs; Finding=only OVH/GCP-hosted tools (WebPageTest, BuiltWith), no ChatGPT/Claude UAs[^18] — https://llms-explorer.com/reference/evidence/#4-who-reads-server-log-studies
- [parameter] Cloudflare `Accept: text/markdown` (Mar–Apr 2026): Window / sample=44 days, one Worker; Finding=1,421 requests: headless Chrome 639, "Claude" (Anthropic infra) 500, axios 211; no GPTBot/PerplexityBot/ClaudeBot[^19] — https://llms-explorer.com/reference/evidence/#4-who-reads-server-log-studies
- [parameter] directory.llmstxt.cloud: Size="4,000 websites listed" (49M llms.txt tokens / 325M llms-full tokens); Notes=named in spec v2 — https://llms-explorer.com/reference/evidence/#6-directories-and-registries
- [parameter] llmstxthub.com: Size=~2,650 entries, 15–16 categories (David Dias); Notes=named in spec v2 — https://llms-explorer.com/reference/evidence/#6-directories-and-registries
- [parameter] llmstxt.site: Size=~1,000+ (≈170 in May 2025); columns product / website / llms.txt / llms-full.txt / **token counts**; `/submit`; Notes=named in spec v2 — https://llms-explorer.com/reference/evidence/#6-directories-and-registries
- [parameter] SecretiveShell/Awesome-llms-txt: Size=784 link lines (counted 2026-08-30); Notes=GitHub — https://llms-explorer.com/reference/evidence/#6-directories-and-registries
- [parameter] llms-text.com: Size="780+ verified implementations"; Notes=vendor's own directory — https://llms-explorer.com/reference/evidence/#6-directories-and-registries
- [parameter] llms-text.com/blog/sites-using-llms-txt: Author / date=Michael Vereb, 2025-07-25; Claims="780+ verified"; names Anthropic, Cloudflare, Supabase, Vercel, ElevenLabs, Firecrawl, Mintlify, Cursor, Aptos, GitBook, Wix; "no e-commerce adoption"; Grade=adopters check out on live probe; count uncorroborated — low for numbers, fine for examples[^31] — https://llms-explorer.com/reference/evidence/#7-vendor-sources-graded
- [parameter] llms-text.com/blog/what-is-llms-txt: Author / date=same; Claims="foundational pillar of GEO"; ChatGPT/Perplexity/Cursor/Windsurf/Claude Code consume it; "up to 114% more tokens" (incoherent arithmetic), "10–15% accuracy" — unattributed; Grade=GEO and ChatGPT/Perplexity-consumption claims contradicted by every log study — low[^32] — https://llms-explorer.com/reference/evidence/#7-vendor-sources-graded
- [parameter] llms-text.com/blog/llms-txt, /how-to-create-llms-txt: Author / date=same; Claims=MIME/200/UTF-8 rules; `Link: …; rel="describedby"`; "under 10 KB"; framework snippets; funnels to its generator/validator; Grade=useful mechanics (the `describedby` relation is now in spec v2), vendor numbers — medium[^33][^34] — https://llms-explorer.com/reference/evidence/#7-vendor-sources-graded
- [parameter] gitdoc.ai/blog/llms-txt-ai-readable-documentation: Author / date=Yadian Llada / GitDoc, 2026-05-22; Claims=GitBook: 41% of docs page requests from AI agents (unverified); permission / inventory / navigation distinction; curate 10–20 pages (quickstart, auth, per-resource reference, errors, changelog); regenerate in the build; llms-full for priority pages; Grade=sound guidance, unverified headline… — https://llms-explorer.com/reference/evidence/#7-vendor-sources-graded
- [definition] 1. The one-line verdict — Adoption is real and growing (≈5–10% of the general web by mid-2026, 28% among SEO-savvy sites, 8.8× year on year); *unsolicited* consumption is near zero (97% of files never get an AI request); the demonstrated use is agents that are pointed at the file — the `Claude-Code` UA out-fetched every AI… — https://llms-explorer.com/reference/evidence/#1-the-one-line-verdict
- [definition] 6. Directories and registries — Self-submitted, overlapping, unverified — lower bounds, not measurements:[^27][^28][^29][^30] — https://llms-explorer.com/reference/evidence/#6-directories-and-registries

## Formatting: the grammars side by side
<https://llms-explorer.com/reference/formatting/>

- [snippet] 1. The index — `llms.txt`: > One paragraph saying what this is and who it is for. — `# Product` — https://llms-explorer.com/reference/formatting/#1-the-index-llmstxt
- [snippet] 4. The facts line — `llms-facts.txt`: https://example.com/docs/install.md — `## Install` — https://llms-explorer.com/reference/formatting/#4-the-facts-line-llms-factstxt
- [snippet] 5. The vocabulary line — `llms-vocabulary.txt`: - **anchor** [llms.anchor] (noun): the `#fragment` on a facts-line URL that names the heading a claim came from — https: — `- **anchor** [llms.anchor] (noun): the `#fragment` on a facts-line URL that names the heading a clai` — https://llms-explorer.com/reference/formatting/#5-the-vocabulary-line-llms-vocabularytxt
- [parameter] Mintlify: Page block=`# Title` / `Source: <url>` / blank / body; blank lines between pages; Who=Mintlify sites, Claude Code docs, **the hub** (`GRAMMAR_NOTE`) — https://llms-explorer.com/reference/formatting/#2-the-full-file-llms-fulltxt
- [parameter] Anthropic YAML: Page block=site H1, `---`, per page `## Heading` + YAML (`title:` / `url:` / `description:`) + raw MDX; Who=platform.claude.com — https://llms-explorer.com/reference/formatting/#2-the-full-file-llms-fulltxt
- [parameter] Cloudflare frontmatter: Page block=YAML frontmatter, a "Documentation Index" blockquote, `# Title`, `[View as Markdown](…/index.md)`, body; Who=developers.cloudflare.com — https://llms-explorer.com/reference/formatting/#2-the-full-file-llms-fulltxt
- [definition] Formatting: the grammars side by side — The index, the three full-file grammars, the facts line, the small file, the vocabulary line, the split root and the manifest — on one page. — https://llms-explorer.com/reference/formatting/#formatting-the-grammars-side-by-side
- [definition] 1. The index — `llms.txt` — The only file the spec defines. Structure, in order: an optional BOM; **one H1** naming the site or product (the only required element); a blockquote summary of one to three sentences; free-form prose (no headings) about how to read the file; then H2 sections, each a list of links. — https://llms-explorer.com/reference/formatting/#1-the-index-llmstxt
- [definition] 2. The full file — `llms-full.txt` — Not in the spec; three grammars are in the wild. The hub emits the first and names it in a header comment so a parser never has to guess. — https://llms-explorer.com/reference/formatting/#2-the-full-file-llms-fulltxt
- [definition] 3. The budgeted file — `llms-small.txt` — Same grammar as the full file, different selection: reference-class pages first, then guides, until `SMALL_MAX_CHARS = 200_000` characters (about 50k tokens at `CHARS_PER_TOKEN = 4`) — the ceiling at which indexed docs become unstable in consumers such as Cursor. — https://llms-explorer.com/reference/formatting/#3-the-budgeted-file-llms-smalltxt
- [definition] 4. The facts line — `llms-facts.txt` — A hub extension: the checkable claims, one per line, each anchored to the heading it came from. — https://llms-explorer.com/reference/formatting/#4-the-facts-line-llms-factstxt
- [definition] 5. The vocabulary line — `llms-vocabulary.txt` — The lexical layer, spec-v2-shaped so any llms reader can open it: an H1 `<Family> — vocabulary`, a blockquote with the term count, then `## Terms`, `## Homonyms` and `## Named, not yet defined`. One line per term per sense: — https://llms-explorer.com/reference/formatting/#5-the-vocabulary-line-llms-vocabularytxt
- [definition] 6. Split roots and families — When an index would exceed 10 KB the sections become subpath indexes: the root keeps the H1, blockquote and a `## Sections` list of `<slug>/llms.txt` links, each line carrying page and token counts; a section with no further path structure is cut into `part-N` files of `PART_PAGES = 60` pages. — https://llms-explorer.com/reference/formatting/#6-split-roots-and-families
- [definition] 7. The manifest — `manifest.json` — Beside the files, never linked from them: `files{name: {bytes, tokens}}`, `chars_per_token`, `pages`, `units`, `sections`, `dropped_empty_pages`, `acquired` (how the mirror was obtained), and `overrides` — the hand inputs (`title`, `summary`, `section_order`, `note`) that survive regeneration. — https://llms-explorer.com/reference/formatting/#7-the-manifest-manifestjson
- [definition] Reading order — Index first, always. Fall through to `llms-small.txt` when you need whole pages and have a budget, to `llms-full.txt` when you have none, to `llms-facts.txt` when you need a claim with a place to check it. — https://llms-explorer.com/reference/formatting/#reading-order

## Glossary
<https://llms-explorer.com/reference/glossary/>

- [definition] Glossary — The terms of the field, one line each, in the sense this site uses them — with the contrasts that matter. — https://llms-explorer.com/reference/glossary/#glossary

## The passes
<https://llms-explorer.com/reference/passes/>

- [parameter] B0: Passes=P0; Kind=deterministic; Runs as=inline, first — every other pass keys off the detected kind — https://llms-explorer.com/reference/passes/#bundle-map-and-dispatch-rules
- [parameter] B1: Passes=P1 P2 P3 P5 P14; Kind=deterministic; Runs as=one `llms_lint.py` invocation, JSON findings — https://llms-explorer.com/reference/passes/#bundle-map-and-dispatch-rules
- [parameter] B2: Passes=P4 P9; Kind=model; Runs as=one subagent reading the index (+ sample pages) — https://llms-explorer.com/reference/passes/#bundle-map-and-dispatch-rules
- [parameter] B3: Passes=P6 P7; Kind=deterministic; Runs as=`llms_lint.py --full` / `--facts` (same invocation as B1 when the kind is full/facts) — https://llms-explorer.com/reference/passes/#bundle-map-and-dispatch-rules
- [parameter] B4: Passes=P8; Kind=model, sampled; Runs as=one subagent, 20 units re-read against source spans — https://llms-explorer.com/reference/passes/#bundle-map-and-dispatch-rules
- [parameter] B5: Passes=P10; Kind=deterministic + model; Runs as=only when kind = family or `--family` — https://llms-explorer.com/reference/passes/#bundle-map-and-dispatch-rules
- [parameter] B6: Passes=P11; Kind=live; Runs as=`docset_indexer.py keyword` + `query --layer facts` probes — https://llms-explorer.com/reference/passes/#bundle-map-and-dispatch-rules
- [parameter] B7: Passes=P12; Kind=live agent; Runs as=one fresh-context subagent given ONLY the file; opt-in `--agent-test`, default on for new files — https://llms-explorer.com/reference/passes/#bundle-map-and-dispatch-rules
- [parameter] B8: Passes=P13; Kind=live HTTP; Runs as=only when a URL is given or `--serve-check` — https://llms-explorer.com/reference/passes/#bundle-map-and-dispatch-rules
- [parameter] B9: Passes=P15; Kind=deterministic; Runs as=only when the export directory has a source mirror — https://llms-explorer.com/reference/passes/#bundle-map-and-dispatch-rules
- [definition] The passes — What the optimizer runs, in order, and how each pass is judged and fixed. — https://llms-explorer.com/reference/passes/#the-passes
- [definition] Severity resolution across passes — Same span flagged by several passes: take the highest severity; tie → lower pass number wins (P0 > P1 > …); tie → the more conservative fix (report over rewrite). — https://llms-explorer.com/reference/passes/#severity-resolution-across-passes
- [definition] N/A rules — A pass reports `N/A (<reason>)` — never silently skips — when: the kind excludes it (P6 on an index), the layer is absent (P11 without an indexed docset), the input is missing (P13 without a URL, P15 without a mirror), or the flag is off (P12 without `--agent-test` on a refresh run). — https://llms-explorer.com/reference/passes/#na-rules

## Reasoning: why the rules are what they are
<https://llms-explorer.com/reference/reasoning/>

- [definition] Reasoning: why the rules are what they are — Extractive descriptions, the size ladder, anchors, facts as the trusted layer, and the two-hop bar — each rule traced to the evidence that produced it. — https://llms-explorer.com/reference/reasoning/#reasoning-why-the-rules-are-what-they-are
- [definition] 1. Extractive descriptions beat generated ones — An index description exists so that a routing model can decide, without fetching, whether the page answers its question. That decision is made on tokens: the flag name, the error string, the endpoint path. — https://llms-explorer.com/reference/reasoning/#1-extractive-descriptions-beat-generated-ones
- [definition] 2. Size is a producer-side problem — Consumers do not truncate gracefully. — https://llms-explorer.com/reference/reasoning/#2-size-is-a-producer-side-problem
- [definition] 3. Anchors make facts checkable — A claim without a place to verify it is a rumour with a URL. The facts line carries `url#anchor`, and the anchor must resolve to a heading that exists on the page (attribute C6). — https://llms-explorer.com/reference/reasoning/#3-anchors-make-facts-checkable
- [definition] 4. The facts file is the trusted layer — Raw page text is untrusted input — the spec repository's own issue #152 found that 42.3% of a 100-file sample tried to steer the reader ([evidence](/reference/evidence/)). — https://llms-explorer.com/reference/reasoning/#4-the-facts-file-is-the-trusted-layer
- [definition] 5. Two hops, from the index alone — The spec's own test is the bar: "test your file by asking an agent questions about your content, giving it only your llms.txt as a starting point." The rubric makes it numeric — attribute R5, pass P12: ten questions, an agent that starts from the index and may follow links, at least eight answered… — https://llms-explorer.com/reference/reasoning/#5-two-hops-from-the-index-alone
- [definition] 6. Publish for agents, not for search — The evidence page holds the numbers: adoption at roughly 5–10% of the general web by mid-2026 and rising 8.8× year on year, yet 97% of valid files received zero AI requests in a month of logs, Google says it does not read the file, and a 300k-domain model found no relationship between having one… — https://llms-explorer.com/reference/reasoning/#6-publish-for-agents-not-for-search
- [definition] 7. Why the rubric is deterministic first — Every attribute names its measure: deterministic, model judgment, or live. The lint implements the deterministic passes P0–P3, P5–P7, P9 and P14 with no model call and gates CI on them; the model, live and family passes (P4, P8, P10–P13, P15) run under `/ldo` when someone asks. — https://llms-explorer.com/reference/reasoning/#7-why-the-rubric-is-deterministic-first

## Recreating and aggregating
<https://llms-explorer.com/reference/recreation/>

- [snippet] 6. Scale to a family: nested indexes, hub-and-spoke: > One index per product below; each product's own llms.txt is the authoritative map of that product. — `# Acme Platform docs` — https://llms-explorer.com/reference/recreation/#6-scale-to-a-family-nested-indexes-hub-and-spoke
- [definition] Recreating and aggregating — The acquisition ladder, lenient parsing, families, rights. — https://llms-explorer.com/reference/recreation/#recreating-and-aggregating
- [definition] 2. Acquire clean markdown — the ladder — Try in this order; each step is cheaper and cleaner than the next: 1. — https://llms-explorer.com/reference/recreation/#2-acquire-clean-markdown-the-ladder
- [definition] 3. Build the index for ONE product — Structure (spec v2): H1 = product name; blockquote = one-paragraph summary; optional prose "how to interpret the files"; H2 sections, each a list of `- [name](url): description`.[^4] — https://llms-explorer.com/reference/recreation/#3-build-the-index-for-one-product
- [definition] 6. Scale to a family: nested indexes, hub-and-spoke — Spec v2 gives the mechanism: "The file can be placed at the site root, or at any path within it, covering the pages under that path … where more than one file applies, agents should use the most specific one."[^4] The live exemplar is **Cloudflare**: `developers.cloudflare.com/llms.txt` holds ~105… — https://llms-explorer.com/reference/recreation/#6-scale-to-a-family-nested-indexes-hub-and-spoke

## llms.txt: the spec and its grammars
<https://llms-explorer.com/reference/spec/>

- [snippet] 2. The spec, v2: > Optional description goes here — `# Title` — https://llms-explorer.com/reference/spec/#2-the-spec-v2
- [parameter] Spec-conformant: Example=code.claude.com/docs/llms.txt, FastHTML; Notes=H1, blockquote, H2 sections, `.md` links — https://llms-explorer.com/reference/spec/#31-llmstxt-three-real-shapes
- [parameter] API-first: Example=docs.github.com/llms.txt; Notes=first H2 "How to use" lists JSON/markdown APIs (Page List, Article Body → markdown, Search) and the MCP server before any content links[^6] — https://llms-explorer.com/reference/spec/#31-llmstxt-three-real-shapes
- [parameter] Non-conformant prose: Example=docs.anthropic.com/llms.txt; Notes=H1, then prose and `## Root URL` / language lists, no blockquote[^6] — https://llms-explorer.com/reference/spec/#31-llmstxt-three-real-shapes
- [parameter] Mintlify: Page block=`# Title` / `Source: <url>` / blank / description / body; pages separated by blank lines only; Verified sample=code.claude.com/docs/llms-full.txt (191 pages, 8.5 MB); mintlify.com/docs/llms-full.txt[^7] — https://llms-explorer.com/reference/spec/#32-llms-fulltxt-not-in-the-spec-and-three-grammars
- [parameter] Anthropic platform: Page block=site H1, `---`, then per-page `## Heading` + YAML block (`title:` / `url:` / `description:`) + raw MDX; Verified sample=platform.claude.com/docs/llms-full.txt[^7] — https://llms-explorer.com/reference/spec/#32-llms-fulltxt-not-in-the-spec-and-three-grammars
- [parameter] Cloudflare: Page block=YAML frontmatter (`description:` / `title:` / `image:`), a "Documentation Index" blockquote pointing at the covering `/<product>/llms.txt`, `# Title`, a `[View as Markdown](…/index.md)` line, body; Verified sample=developers.cloudflare.com/llms-full.txt (57 MB)[^6] — https://llms-explorer.com/reference/spec/#32-llms-fulltxt-not-in-the-spec-and-three-grammars
- [parameter] Firecrawl generators: Page block=pages delimited by `<\|firecrawl-page-N-lllmstxt\|>`; Verified sample=create-llmstxt-py[^10] — https://llms-explorer.com/reference/spec/#32-llms-fulltxt-not-in-the-spec-and-three-grammars
- [parameter] Reference `llms_txt2ctx`: Behaviour=regex-parse, fetch every link, emit XML `<project title summary><docs><doc …>`; `--optional True` includes the Optional section; removed from the proposal in v2; Evidence=[^24][^2] — https://llms-explorer.com/reference/spec/#5-how-consumers-actually-use-it
- [parameter] LangChain `mcpdoc` (MCP): Behaviour=`list_doc_sources` + `fetch_docs`; the *agent* decides which links to follow; allowlists only the llms.txt's domain; Evidence=[^25] — https://llms-explorer.com/reference/spec/#5-how-consumers-actually-use-it
- [parameter] Claude Code: Behaviour=Anthropic publishes its docs index and points the agent at it; Ahrefs' logs show the `Claude-Code` UA out-fetching every AI retrieval bot bar two (statespace-indexer, GPTBot); no documented *automatic* lookup — it is fetched when directed; Evidence=[^14][^26] — https://llms-explorer.com/reference/spec/#5-how-consumers-actually-use-it
- [parameter] Cursor `@Docs`: Behaviour=crawls URLs; "cannot recognise llms.txt" request acknowledged (Jun 2025), no documented support; >50–60k tokens unstable; its own llms.txt once redirected to an HTML app shell; Evidence=[^17][^27] — https://llms-explorer.com/reference/spec/#5-how-consumers-actually-use-it
- [parameter] Windsurf, Copilot: Behaviour=`@docs` is a curated list; Copilot feature request unanswered as of Jul 2026; Evidence=[^28][^29] — https://llms-explorer.com/reference/spec/#5-how-consumers-actually-use-it
- [parameter] ChatGPT, Perplexity, Google: Behaviour=no statements of use; logs ≈ 0 requests; Google: "You don't need to create new machine readable files"; Evidence=[^12][^13][^14] — https://llms-explorer.com/reference/spec/#5-how-consumers-actually-use-it
- [parameter] `robots.txt`: Job=access control; now also carries Cloudflare **Content Signals** (`Content-Signal: search=yes, ai-input=…, ai-train=no`, 2025-09-24); Do AI bots fetch it?=yes, thousands of times per site; Content Signals: Google says "no effects whatsoever"[^36][^37] — https://llms-explorer.com/reference/spec/#7-related-files
- [parameter] `sitemap.xml`: Job=exhaustive inventory; no `.md` versions, no external links; Do AI bots fetch it?=yes (ClaudeBot, GPTBot, Bingbot)[^38] — https://llms-explorer.com/reference/spec/#7-related-files
- [parameter] `llms.txt`: Job=curated navigation for agents pointed at it; Do AI bots fetch it?=~0 speculative fetches; agents when directed[^14] — https://llms-explorer.com/reference/spec/#7-related-files
- [parameter] `ai.txt`: Job=opt-out preferences (IETF draft); Do AI bots fetch it?=397 instances found May 2026[^39] — https://llms-explorer.com/reference/spec/#7-related-files
- [parameter] `agents.md` / `/.well-known/ucp`: Job=Shopify's agent-commerce additions shipped with llms.txt to every store (May 2026); Do AI bots fetch it?=n/a[^40] — https://llms-explorer.com/reference/spec/#7-related-files
- [definition] llms.txt: the spec and its grammars — Spec v2, llms-full grammars, discovery, consumers. — https://llms-explorer.com/reference/spec/#llmstxt-the-spec-and-its-grammars
- [definition] 1. What it is, in one paragraph — A markdown file — `/llms.txt` at a site root or **at any subpath** — that gives a language model a curated, priority-ordered map of a site's LLM-friendly content: an H1, a blockquote summary, optional prose, then H2 sections of `- [name](url): description` links.[^1] The links should point at… — https://llms-explorer.com/reference/spec/#1-what-it-is-in-one-paragraph
- [definition] 2. The spec, v2 — Structure, in order (verbatim from llmstxt.org):[^1] — https://llms-explorer.com/reference/spec/#2-the-spec-v2
- [definition] 3.2 llms-full.txt — not in the spec, and three grammars — `llms-full.txt` (the whole docset inlined into one markdown file) appears nowhere in the v1/v2 spec text or the repo README.[^7] Mintlify says it "was developed by Mintlify in collaboration with customer Anthropic";[^8] Lab451 dates its popularisation to early 2025.[^9] There is **no single… — https://llms-explorer.com/reference/spec/#32-llms-fulltxt-not-in-the-spec-and-three-grammars
- [definition] 3.3 `.md` twins — Mintlify, Fern, GitBook and ReadMe all serve a `.md` twin per page and link them from llms.txt "so AI tools can fetch the Markdown version of each page directly".[^4][^11] Mintlify's twins begin with a blockquote — `> ## Documentation Index` / `Fetch the complete documentation index at… — https://llms-explorer.com/reference/spec/#33-md-twins

## Generation tooling
<https://llms-explorer.com/reference/tooling/>

- [parameter] Docs on Mintlify / GitBook / ReadMe / Fern: Use=nothing — it is automatic; Emits=llms.txt (+ full on Mintlify/GitBook) + `.md` twins — https://llms-explorer.com/reference/tooling/#1-pick-by-situation
- [parameter] Docusaurus, MkDocs, VitePress, Starlight, Sphinx, Nuxt: Use=the framework plugin (table §3); Emits=llms.txt + llms-full.txt (+ `.md`, `llms-small.txt` on Starlight) — https://llms-explorer.com/reference/tooling/#1-pick-by-situation
- [parameter] A live site you do not own: Use=crawl-based generator (§4) — `create-llmstxt-py`, `dotenvx/llmstxt`, or your own sitemap→markdown pipeline; Emits=llms.txt (+ full) with **extracted or AI-written** descriptions — https://llms-explorer.com/reference/tooling/#1-pick-by-situation
- [parameter] WordPress: Use=Yoast ≥25.3 / Rank Math / AIOSEO (§5); Emits=llms.txt only (AIOSEO Pro adds full + markdown posts) — https://llms-explorer.com/reference/tooling/#1-pick-by-situation
- [parameter] Webflow / Framer: Use=host a hand-written file; Emits=whatever you upload — https://llms-explorer.com/reference/tooling/#1-pick-by-situation
- [parameter] Any Cloudflare-proxied HTML site: Use="Markdown for Agents" toggle (§6); Emits=on-the-fly markdown on `Accept: text/markdown`, no llms.txt — https://llms-explorer.com/reference/tooling/#1-pick-by-situation
- [parameter] **Mintlify**: Emits=llms.txt, llms-full.txt, `.md` per page, `/.well-known/` copies, `/_llms/` split indexes; Descriptions from=frontmatter `description` (truncated at 300 chars), nav order from `docs.json`; optional `markdown.instructions` agent text; Notes=index capped at 100,000 chars → recursive `/_llms/<group>.md`; default language/version only; hidden/noindex pages excluded; hand-written… — https://llms-explorer.com/reference/tooling/#2-docs-platforms
- [parameter] **Fern**: Emits=llms.txt (root **and per-subdirectory**), `.md` per page; **no llms-full.txt**; Descriptions from=frontmatter `description`, fallback `subtitle`; adds OpenAPI/AsyncAPI links; Notes=dropped llms-full because it "exceeded most model context windows, added heavy serving overhead, saw little use"[^2] — https://llms-explorer.com/reference/tooling/#2-docs-platforms
- [parameter] **GitBook**: Emits=llms.txt (Jan 2025), llms-full.txt + `.md` per page (Jun 2025), `/sitemap.md`, `Accept: text/markdown`; Descriptions from=auto from page structure; Notes=zero-config; no curation controls documented; full export "will be more expensive"[^3][^4] — https://llms-explorer.com/reference/tooling/#2-docs-platforms
- [parameter] **ReadMe**: Emits=llms.txt (default on, all plans), `.md` per page; **no llms-full**; Descriptions from=project title + guide/API hierarchy; Notes=a custom file from the repo root disables auto-updates; hidden pages excluded[^5] — https://llms-explorer.com/reference/tooling/#2-docs-platforms
- [parameter] **GitDoc** (vendor claim): Emits=llms.txt + llms-full.txt "for the pages you mark as priority", regenerated in the build; Descriptions from=sidebar/nav; Notes=vendor blog, 2026-05-22[^6] — https://llms-explorer.com/reference/tooling/#2-docs-platforms
- [parameter] `docusaurus-plugin-llms` (rachfop): Emits=llms.txt, llms-full.txt, optional per-page `.md`, versioned + `customLLMFiles`; Input=source tree at `postBuild`; Descriptions / ordering=frontmatter → first heading → site fallback; `includeOrder` globs; Maturity & limits=144★, MIT; not run in `docusaurus start`; image rewrite only for bundled assets[^7] — https://llms-explorer.com/reference/tooling/#3-static-site-generator-plugins
- [parameter] `@signalwire/docusaurus-plugin-llms-txt`: Emits=llms.txt, `.md`, optional full; Input=**built HTML** (rehype/remark); Descriptions / ordering=manual `sections[].description`, `autoSectionDepth`; Maturity & limits=v1.2.2, ~10 months stale; ENOENT / "processed 0 documents" bug[^8][^9] — https://llms-explorer.com/reference/tooling/#3-static-site-generator-plugins
- [parameter] Docusaurus core: Emits=none; Input=—; Descriptions / ordering=—; Maturity & limits=issue #10899 open since Feb 2025[^10] — https://llms-explorer.com/reference/tooling/#3-static-site-generator-plugins
- [parameter] `mkdocs-llmstxt` (pawamoy): Emits=llms.txt, `.md`, optional `full_output`; Input=built HTML → BeautifulSoup → Markdownify; Descriptions / ordering=`sections:` dict with per-file descriptions; Maturity & limits=130★, v0.5.x, **maintenance mode, seeking maintainer**; needs `site_url`; mkdocstrings `show_source` mangles tables/code in the full file[^11][^12] — https://llms-explorer.com/reference/tooling/#3-static-site-generator-plugins
- [parameter] `vitepress-plugin-llms` (okineadev): Emits=llms.txt, llms-full.txt, `.md`; Input=VitePress source; Descriptions / ordering=frontmatter `description`; `<llm-only>` / `<llm-exclude>` tags; Maturity & limits=394★; used by Vite, Vue, Vitest, Rolldown; relative URLs break under redirects/domain moves[^13] — https://llms-explorer.com/reference/tooling/#3-static-site-generator-plugins
- [parameter] `starlight-llms-txt` (delucis): Emits=llms.txt, llms-full.txt, **llms-small.txt**; Input=Astro Starlight; Descriptions / ordering=`projectName`, `description`, `details`, `optionalLinks`, `customSets`, `promote`/`demote`; `minify` strips asides; Maturity & limits=110★, docs updated Aug 2026; needs `site`[^14] — https://llms-explorer.com/reference/tooling/#3-static-site-generator-plugins
- [parameter] `sphinx-llms-txt` (jdillard): Emits=llms.txt (markdown), llms-full.txt (**reStructuredText**); Input=Sphinx build; Descriptions / ordering=toctree titles; `llms_txt_summary`, `llms_txt_exclude`, `llms_txt_full_max_size`; Maturity & limits=v0.7.1; full file is RST; points to NVIDIA `sphinx-llm`[^15] — https://llms-explorer.com/reference/tooling/#3-static-site-generator-plugins
- [parameter] `nuxt-llms` / Nuxt Content: Emits=llms.txt (~5K tokens), opt-in llms-full.txt (~1M+ tokens); Input=Nuxt Content, runtime hooks; Descriptions / ordering=`sections` in `nuxt.config`; Maturity & limits=first-party; full file explicitly for 200K+-context tools[^16] — https://llms-explorer.com/reference/tooling/#3-static-site-generator-plugins
- [parameter] Next.js / Nextra: Emits=hand-rolled `app/llms.txt/route.ts` (force-static or dynamic); `next-llms-txt` adds per-page `.md` endpoints; Input=components; Descriptions / ordering="reads and parses readable text"; Maturity & limits=discussion #80692 unresolved; no Nextra built-in found (tentative)[^17][^18] — https://llms-explorer.com/reference/tooling/#3-static-site-generator-plugins
- [parameter] `llms-txt-action` (demodrive-ai): Emits=llms.txt, llms-full.txt, `.md`; Input=built HTML dir + sitemap.xml; Descriptions / ordering=local/offline or cloud LLM summaries via LiteLLM (default GPT-4o); Maturity & limits=16★; needs `--dirty` with `mkdocs gh-deploy`[^19] — https://llms-explorer.com/reference/tooling/#3-static-site-generator-plugins
- [parameter] Firecrawl `/llmstxt` API + llmstxt.firecrawl.dev: What it does=URL → async job → llms.txt (+ full); `maxUrls` 1–100 (default 10), 1 credit/URL, public pages only, 5,000-URL alpha cap; Limits=**deprecated in favour of the main endpoints** (page carries no date; still up); users pointed to the Python repo[^20][^21] — https://llms-explorer.com/reference/tooling/#4-crawl-based-generators-sites-you-do-not-own
- [parameter] `create-llmstxt-py` (Firecrawl, 320★): What it does=`/map` → scrape each page to markdown (batches of 10; failures skipped, no retry) → GPT-4o-mini writes a 3–4-word title + 9–10-word description → flat llms.txt; llms-full.txt concatenates under `<\|firecrawl-page-N-lllmstxt\|>`; Limits=default 20 URLs; memory issues on large sites; **sections are not inferred**; descriptions are AI-written and… — https://llms-explorer.com/reference/tooling/#4-crawl-based-generators-sites-you-do-not-own
- [parameter] `dotenvx/llmstxt` (147★, BSD-3): What it does=sitemap.xml → `- [Title](url): description` bullets; `--include-path` / `--exclude-path` globs; `--replace-title` regex; Limits=llms.txt only; titles extracted from HTML; description derivation undocumented[^23] — https://llms-explorer.com/reference/tooling/#4-crawl-based-generators-sites-you-do-not-own
- [parameter] Jina Reader `r.jina.ai/<url>`: What it does=headless Chrome or curl engine → Readability → Turndown; headers `x-respond-with`, `x-target-selector`, `x-retain-links`, `x-max-tokens`, `x-markdown-chunking`; Limits=per-page cleaner, no site/llms.txt mode; anonymous traffic rate-limited[^24] — https://llms-explorer.com/reference/tooling/#4-crawl-based-generators-sites-you-do-not-own
- [parameter] Screaming Frog v24.3: What it does=per-page `.md` via a Readability.js + Turndown custom-JS snippet; llms.txt via n8n/CSV converters; Limits=no native llms.txt export; thin pages return nothing; JS rendering slow[^25][^26] — https://llms-explorer.com/reference/tooling/#4-crawl-based-generators-sites-you-do-not-own
- [parameter] `plainsignal/llmstxt` Chrome extension: What it does=llms.txt + one `.md` per page + zip from sitemap or rendered DOM; meta description as blockquote; Limits=10★, HTTPS only[^27] — https://llms-explorer.com/reference/tooling/#4-crawl-based-generators-sites-you-do-not-own
- [parameter] SEO-tool generators (SEOmator etc.): What it does=robots.txt → sitemap discovery, index-sitemap expansion, LLM-written title+description per URL; Limits=vendor-claimed mechanics only[^28] — https://llms-explorer.com/reference/tooling/#4-crawl-based-generators-sites-you-do-not-own
- [parameter] llms-text.com generator/validator: What it does=crawls a domain and exports llms.txt + llms-full.txt ("deep-crawls up to 50 subpages"); validator checks syntax, links, UTF-8, headers; Limits=vendor; its guidance: 10–20 evergreen URLs, 4–7 H2s, 10–20-word descriptions, index under 10 KB, `Content-Type: text/plain — https://llms-explorer.com/reference/tooling/#4-crawl-based-generators-sites-you-do-not-own
- [parameter] Yoast SEO ≥25.3 (2025-06-10): Emits=llms.txt only, regenerated weekly; Descriptions=custom excerpt only — **no description otherwise**; 5 latest posts/pages/CPT (≤12 months, cornerstone first) + top-5 taxonomies; Limits=5-item cap; markdown chars escaped; a static file wins over the dynamic one[^31][^32] — https://llms-explorer.com/reference/tooling/#5-cms-and-site-builders
- [parameter] Rank Math: Emits=llms.txt only; Descriptions="intro text"; post types/taxonomies, limit default 100; custom lines; Limits=no full[^33] — https://llms-explorer.com/reference/tooling/#5-cms-and-site-builders
- [parameter] AIOSEO: Emits=llms.txt (free); llms-full.txt + markdown post conversion (Pro); Descriptions=site title/tagline; per-post-type limits, exclusions; Limits=paywall[^34] — https://llms-explorer.com/reference/tooling/#5-cms-and-site-builders
- [parameter] `website-llms-txt`, `llms-full-txt-generator`: Emits=llms.txt (+ full); Descriptions=titles + SEO-plugin descriptions; honour noindex; Limits=one shipped a broken-access-control CVE fix[^35] — https://llms-explorer.com/reference/tooling/#5-cms-and-site-builders
- [parameter] Joost de Valk "Markdown Alternate": Emits=`<link rel="alternate" type="text/markdown">` + `.md` URLs per post; Descriptions=—; Limits=negotiation, not an index[^36] — https://llms-explorer.com/reference/tooling/#5-cms-and-site-builders
- [parameter] Webflow / Framer: Emits=host an uploaded file (Framer: Pro/Enterprise "Hosting → Files"); a Framer marketplace plugin scans the CMS; Descriptions=manual; Limits=no generation[^37][^38] — https://llms-explorer.com/reference/tooling/#5-cms-and-site-builders
- [parameter] Shopify (Apr–May 2026, silent): Emits=auto `/llms.txt`, `/agents.md`, `/sitemap_agentic_discovery.xml`, `/.well-known/ucp` on every store; Descriptions=boilerplate: H1 store name, `/collections/all`, contact, UCP + MCP endpoints; Limits=`templates/llms.txt.liquid` **replaces, does not merge**; no changelog; 78.1% of top-10k Shopify hosts vs WordPress 8.7%[^39][^40][^41] — https://llms-explorer.com/reference/tooling/#5-cms-and-site-builders
- [definition] Generation tooling — Generators compared; why extractive descriptions win. — https://llms-explorer.com/reference/tooling/#generation-tooling
- [definition] 6. Edge content negotiation — Cloudflare "Markdown for Agents" (2026-02-12; Pro/Business/Enterprise; zone toggle under AI Crawl Control): on `Accept: text/markdown` the edge converts HTML → markdown (body + meta-derived YAML frontmatter + JSON-LD, nav/header/footer/scripts dropped) and returns `Content-Type: text/markdown… — https://llms-explorer.com/reference/tooling/#6-edge-content-negotiation

## Usage: serving, discovering and reading llms files
<https://llms-explorer.com/reference/usage/>

- [parameter] `Content-Type`: Value=`text/markdown; charset=utf-8`; Why=attribute H2; `text/plain` is tolerated, HTML is a High — https://llms-explorer.com/reference/usage/#1-serving
- [parameter] `X-Markdown-Tokens`: Value=`bytes // 4` — the same estimator `manifest.json` uses; Why=H4: cost known before fetch — https://llms-explorer.com/reference/usage/#1-serving
- [parameter] `Link`: Value=`</llms.txt>; rel="describedby"` — the index that covers this file; Why=H3, spec v2 discovery — https://llms-explorer.com/reference/usage/#1-serving
- [parameter] `hub_docset_index(key)`: the docset's `llms.txt` (or `llms-small.txt`, `llms-facts.txt`, `manifest.json`, `<section>/llms.txt`), with served URLs — https://llms-explorer.com/reference/usage/#5-claude-code-and-mcp
- [parameter] `hub_query_docset(key, q, mode=semantic\|keyword\|hybrid, layer=auto\|facts\|raw)`: ranked units or chunks; `layer=auto` prefers the facts layer — https://llms-explorer.com/reference/usage/#5-claude-code-and-mcp
- [parameter] `hub_llms_full_read(key, page=…)` or `(offset, limit)`: one page or a slice of a mirrored `llms-full.txt` — https://llms-explorer.com/reference/usage/#5-claude-code-and-mcp
- [parameter] `hub_llms_full_list(query, category, status, min_pages)`: which sites publish a full file, with sizes — https://llms-explorer.com/reference/usage/#5-claude-code-and-mcp
- [definition] Usage: serving, discovering and reading llms files — The headers to send, the .md twins to publish, how a reader discovers the family, how an agent reads an index, and how Claude Code and the hub MCP tools consume one. — https://llms-explorer.com/reference/usage/#usage-serving-discovering-and-reading-llms-files
- [definition] 1. Serving — Every markdown file in the family is served with: — https://llms-explorer.com/reference/usage/#1-serving
- [definition] 2. Markdown twins — Every page in the content sections — reference, essays, examples, blog — has a clean-markdown twin at the same route with `.md` appended: `/reference/usage/` → `/reference/usage.md`. — https://llms-explorer.com/reference/usage/#2-markdown-twins
- [definition] 4. Reading an index — The v2 consumption model: *view or search the index, then follow the relevant links; the detail lives behind the links and is fetched only when needed.* As a procedure: 1. Read the H1 and blockquote — is this the product you meant? — https://llms-explorer.com/reference/usage/#4-reading-an-index
- [definition] 5. Claude Code and MCP — Claude Code fetches an llms file when directed — Anthropic publishes its own docs index and points the agent at it, and the `Claude-Code` user agent shows up in server logs ahead of every AI retrieval bot but two. The pattern is a URL in a prompt or a `CLAUDE.md`, not automatic lookup. — https://llms-explorer.com/reference/usage/#5-claude-code-and-mcp

## Conceptual vs proprietary llms files
<https://llms-explorer.com/essays/cllms-vs-proprietary/>

- [parameter] file: source axis=`llms.txt`, `llms-full.txt`, `llms-facts.txt`; concept axis=`llms-concepts.txt`, `/t/<slug>/llms-facts.txt`, concept packs — https://llms-explorer.com/essays/cllms-vs-proprietary/#two-axes
- [parameter] the authority: source axis=the publisher; concept axis=the concept — https://llms-explorer.com/essays/cllms-vs-proprietary/#two-axes
- [parameter] the unit of truth: source axis=a page; concept axis=a claim about the concept — https://llms-explorer.com/essays/cllms-vs-proprietary/#two-axes
- [parameter] what "correct" means: source axis=the link resolves and says what the description said; concept axis=the claim survives comparison with every other source's claim — https://llms-explorer.com/essays/cllms-vs-proprietary/#two-axes
- [parameter] what a disagreement is: source axis=impossible: one publisher, one page; concept axis=routine: two sources, two claims, one concept — https://llms-explorer.com/essays/cllms-vs-proprietary/#two-axes
- [parameter] anyone (no account): may=read every file, every conflict record, the ladder; may not=submit — https://llms-explorer.com/essays/cllms-vs-proprietary/#governance
- [parameter] contributor (account): may=submit a unit with a source URL and anchor; it runs through `resolve`; may not=write a unit without a source; skip the ladder — https://llms-explorer.com/essays/cllms-vs-proprietary/#governance
- [parameter] maintainer: may=settle ties in the moderation queue with a note; reject a submission; may not=overwrite a ladder verdict without a note in the record — https://llms-explorer.com/essays/cllms-vs-proprietary/#governance
- [parameter] fork owner: may=keep a private tree with its own `precedence.json`; propose merges back as diffs of units; may not=push to the public tree directly — https://llms-explorer.com/essays/cllms-vs-proprietary/#governance
- [parameter] the lint: may=block any file with a High finding from being served; may not=be bypassed — https://llms-explorer.com/essays/cllms-vs-proprietary/#governance
- [definition] Conceptual vs proprietary llms files — Why a file no vendor owns can be trusted: the two axes, the precedence ladder that lets the most correct idea overwrite, and the rules that keep the losers visible. — https://llms-explorer.com/essays/cllms-vs-proprietary/#conceptual-vs-proprietary-llms-files
- [definition] Two axes — Every llms file sits on one of two axes. — https://llms-explorer.com/essays/cllms-vs-proprietary/#two-axes
- [definition] The most correct idea overwrites — On the concept axis, units compete. — https://llms-explorer.com/essays/cllms-vs-proprietary/#the-most-correct-idea-overwrites
- [definition] The precedence ladder — Higher rungs win; a tie on a rung falls to the next. The rungs, in order: 1. — https://llms-explorer.com/essays/cllms-vs-proprietary/#the-precedence-ladder
- [definition] Disagreements stay visible — A conflict that the ladder settles produces a winner in `llms-facts.txt` and a loser in a `## Disagreements` section of the pack's `llms-full.txt`. A conflict it cannot settle puts both there. — https://llms-explorer.com/essays/cllms-vs-proprietary/#disagreements-stay-visible
- [definition] Governance — Who can overwrite what, on the public tree: — https://llms-explorer.com/essays/cllms-vs-proprietary/#governance
- [definition] Rights — What a conceptual file may contain is narrower than what a proprietary one may: - **Links** — always. A link to a publisher's page is what the publisher wants. — https://llms-explorer.com/essays/cllms-vs-proprietary/#rights
- [definition] Honesty note — Three things the reader should know before trusting any of the above. — https://llms-explorer.com/essays/cllms-vs-proprietary/#honesty-note

## Semantic indexing: two legs and a fusion
<https://llms-explorer.com/essays/semantic-indexing/>

- [snippet] Run it yourself: docset_indexer.py keyword <docset> "CLAUDE_CODE_SYNC_SKILLS" — `# the cheap leg — FTS5/BM25, no model call` — https://llms-explorer.com/essays/semantic-indexing/#run-it-yourself
- [definition] Semantic indexing: two legs and a fusion — A docset can be asked a question by token or by meaning, and the two ways fail in opposite directions. What each leg costs, where each breaks, why reciprocal-rank fusion is the default, and how to read the recorded run at /demo/. — https://llms-explorer.com/essays/semantic-indexing/#semantic-indexing-two-legs-and-a-fusion
- [definition] The two legs — A docset in the hub is indexed twice over the same units. — https://llms-explorer.com/essays/semantic-indexing/#the-two-legs
- [definition] Fusing them — The hub's default is neither leg alone. `hub_query_docset(mode="hybrid")` runs both and fuses them with **reciprocal-rank fusion**: each hit scores `1 / (60 + rank)` in each list it appears in, and the scores add, keyed by `(url, seq)`. — https://llms-explorer.com/essays/semantic-indexing/#fusing-them
- [definition] What the recording shows — Read [`/demo/`](/demo/) with three questions in mind. — https://llms-explorer.com/essays/semantic-indexing/#what-the-recording-shows
- [definition] Run it yourself — Everything on the demo page comes from one command against an indexed docset, so the same comparison can be run over yours. The keyword index is built on first use; the facts layer is preferred automatically when a docset has one. — https://llms-explorer.com/essays/semantic-indexing/#run-it-yourself

## V2 vs V1
<https://llms-explorer.com/essays/v2-vs-v1/>

- [parameter] Required elements: v1=H1 + blockquote + sections implied; v2 (2026-08-10)=**H1 only** is required; blockquote, prose and sections are optional; Effect on an existing file=none is invalidated; the lint still scores a missing blockquote as Medium (I2) — a quality finding, not a validity one — https://llms-explorer.com/essays/v2-vs-v1/#the-spec-v1-v2
- [parameter] Placement: v1=`/llms.txt` at the site root; v2 (2026-08-10)=root **or any subpath** (`/docs/llms.txt`); a file covers the URLs under its path; where several apply, **the most specific wins**; Effect on an existing file=enables families and split roots (`<section>/llms.txt`) — https://llms-explorer.com/essays/v2-vs-v1/#the-spec-v1-v2
- [parameter] Discovery: v1=none; v2 (2026-08-10)=`Link: <…>; rel="describedby"` on the files; `rel="alternate" type="text/markdown"` on HTML pages; Effect on an existing file=add two headers (see the [serving reference](/reference/usage/#1-serving)) — https://llms-explorer.com/essays/v2-vs-v1/#the-spec-v1-v2
- [parameter] Markdown twins: v1=`page.html.md`; v2 (2026-08-10)=`page.html.md` **or** `page.md`; directories append `index.html.md` or `index.md`; Effect on an existing file=either form passes the twin probe — https://llms-explorer.com/essays/v2-vs-v1/#the-spec-v1-v2
- [parameter] `## Optional`: v1=mechanical: "can be skipped if a shorter context is needed", consumed by `llms_txt2ctx`; v2 (2026-08-10)=a **convention** for secondary information; `llms_txt2ctx` and its context-expansion mechanics are no longer part of the proposal; Effect on an existing file=keep it last; build nothing that depends on it — https://llms-explorer.com/essays/v2-vs-v1/#the-spec-v1-v2
- [parameter] BOM: v1=—; v2 (2026-08-10)=an optional BOM is tolerated; Effect on an existing file=the lint strips it as hygiene (P14) — https://llms-explorer.com/essays/v2-vs-v1/#the-spec-v1-v2
- [parameter] Consumption expectation: v1=expand the file into context; v2 (2026-08-10)="view or search the index, then follow the relevant links"; the index stays small; detail lives behind links; Effect on an existing file=the size ladder (small / full) becomes the producer's job — https://llms-explorer.com/essays/v2-vs-v1/#the-spec-v1-v2
- [parameter] `/.well-known/`: v1=—; v2 (2026-08-10)=explicitly rejected: well-known URIs exist only at the origin root, which defeats subpath scoping; Effect on an existing file=serve at the root or the subpath, not under `.well-known` — https://llms-explorer.com/essays/v2-vs-v1/#the-spec-v1-v2
- [parameter] Acquire: V1 (to 2026-08-29)=trafilatura BFS crawl → banner mirror; V2 (from 2026-08-30)=a ladder in `llms_acquire.py`: the site's `llms-full.txt` → its `llms.txt` + `.md` twins → `Accept: text/markdown` → a docs API → a structured crawl; the banner mirror stays the internal format — https://llms-explorer.com/essays/v2-vs-v1/#the-pipeline-v1-v2
- [parameter] Clean: V1 (to 2026-08-29)=none (raw HTML → text); V2 (from 2026-08-30)=`docset_refine clean`: boilerplate lines, MDX → markdown, page classes (reference / guide / changelog / marketing / index) — https://llms-explorer.com/essays/v2-vs-v1/#the-pipeline-v1-v2
- [parameter] Extract: V1 (to 2026-08-29)=`distill_offline.py bulk` — zero-LLM, output never consumed; V2 (from 2026-08-30)=`extract` (code snippets, table rows → `parameter`, definitions, changelog `change` units; anchors to real source headings) + `units` (local LLM under the evidence rule) + `polish` (Claude) — https://llms-explorer.com/essays/v2-vs-v1/#the-pipeline-v1-v2
- [parameter] Export: V1 (to 2026-08-29)=none; V2 (from 2026-08-30)=`export_llms`: index (split above 10 KB) / full (Mintlify grammar) / small (≤ ~50k tokens) / facts / `manifest.json`; `topical`; `vocabulary` — https://llms-explorer.com/essays/v2-vs-v1/#the-pipeline-v1-v2
- [parameter] Index: V1 (to 2026-08-29)=one raw vector layer (`nomic-embed-text` in `hub.db` for files; `mxbai-embed-large` for docsets); V2 (from 2026-08-30)=raw **and** facts vector layers, plus an FTS5 keyword layer beside each (`docset_indexer keyword-index`) — https://llms-explorer.com/essays/v2-vs-v1/#the-pipeline-v1-v2
- [parameter] Serve: V1 (to 2026-08-29)=`web-text-mirror --serve` (HTML); V2 (from 2026-08-30)=`llms_serve.py`: `/llms.txt`, `/d/<stem>/…` (with sections), `/m/<key>/…`, `/t/<slug>/…`, markdown headers on every response — https://llms-explorer.com/essays/v2-vs-v1/#the-pipeline-v1-v2
- [parameter] Gate: V1 (to 2026-08-29)=none; V2 (from 2026-08-30)=`llms_lint.py` (the deterministic passes P0–P3, P5–P7, P9, P14) inside `docset_rollout cleanup`; `/ldo` for the model, live and family passes — https://llms-explorer.com/essays/v2-vs-v1/#the-pipeline-v1-v2
- [parameter] Artifacts: V1 (to 2026-08-29)=`<stem>.pages/`, `_master.md`, `._distill_index.json`; V2 (from 2026-08-30)=`<stem>.reference/{pages.json, structured.jsonl, units.jsonl, all_units.jsonl}` and `<stem>.llms/` — https://llms-explorer.com/essays/v2-vs-v1/#the-pipeline-v1-v2
- [parameter] v1 file at root: Claude Code (`WebFetch` / hub MCP)=works; Cursor=works; generic MCP client=works; `llms_acquire`=works; lint=works (I2 Medium if no blockquote); Lighthouse agentic audit=works — https://llms-explorer.com/essays/v2-vs-v1/#compatibility-matrix
- [parameter] v2 file at root: Claude Code (`WebFetch` / hub MCP)=works; Cursor=works; generic MCP client=works; `llms_acquire`=works; lint=works; Lighthouse agentic audit=works — https://llms-explorer.com/essays/v2-vs-v1/#compatibility-matrix
- [parameter] v2 file at a subpath only: Claude Code (`WebFetch` / hub MCP)=works if given the URL; Cursor=degraded — no root discovery; generic MCP client=works if given the URL; `llms_acquire`=works — the ladder probes the given path; lint=works; Lighthouse agentic audit=degraded — expects the root — https://llms-explorer.com/essays/v2-vs-v1/#compatibility-matrix
- [parameter] split root (`## Sections`): Claude Code (`WebFetch` / hub MCP)=works — follows section links; Cursor=works — one extra hop; generic MCP client=works; `llms_acquire`=works — recurses by path then `part-N`; lint=works — `check DIR` walks sections; Lighthouse agentic audit=works — https://llms-explorer.com/essays/v2-vs-v1/#compatibility-matrix
- [parameter] family file (links only indexes): Claude Code (`WebFetch` / hub MCP)=works; Cursor=works; generic MCP client=works; `llms_acquire`=works; lint=works (F1 requires index targets); Lighthouse agentic audit=not evaluated — https://llms-explorer.com/essays/v2-vs-v1/#compatibility-matrix
- [parameter] `llms-full.txt`, Mintlify grammar: Claude Code (`WebFetch` / hub MCP)=works via `hub_llms_full_read(page=…)`; Cursor=degraded above ~50k tokens (the consumer ceiling); generic MCP client=works; `llms_acquire`=works — `split_llms_full` round-trips; lint=works; Lighthouse agentic audit=not evaluated — https://llms-explorer.com/essays/v2-vs-v1/#compatibility-matrix
- [parameter] `llms-full.txt`, YAML-block or Cloudflare frontmatter grammar: Claude Code (`WebFetch` / hub MCP)=works; Cursor=degraded above ~50k tokens; generic MCP client=works; `llms_acquire`=works — grammar detected from the header comment; lint=works; Lighthouse agentic audit=not evaluated — https://llms-explorer.com/essays/v2-vs-v1/#compatibility-matrix
- [parameter] `llms-full.txt` served **as** `llms.txt`: Claude Code (`WebFetch` / hub MCP)=degraded — the index is unreadable at that size; Cursor=breaks; generic MCP client=degraded; `llms_acquire`=works — detected and split; lint=**High** (I6); Lighthouse agentic audit=breaks — https://llms-explorer.com/essays/v2-vs-v1/#compatibility-matrix
- [parameter] no `.md` twins: Claude Code (`WebFetch` / hub MCP)=works — fetches HTML; Cursor=works; generic MCP client=works; `llms_acquire`=degraded — falls to the `Accept` probe or the crawl; lint=High (N6, with `--check-links`); Lighthouse agentic audit=degraded — https://llms-explorer.com/essays/v2-vs-v1/#compatibility-matrix
- [parameter] no `Link` headers: Claude Code (`WebFetch` / hub MCP)=works; Cursor=works; generic MCP client=works; `llms_acquire`=works; lint=Low (H3); Lighthouse agentic audit=degraded — https://llms-explorer.com/essays/v2-vs-v1/#compatibility-matrix
- [definition] V2 vs V1 — Two versioned things share a name: the llms.txt spec (v1 → v2, 2026-08-10) and the hub pipeline (V1 site dumps → V2 acquire, refine, dual index, gate). Both diffs, a migration guide, and what breaks. — https://llms-explorer.com/essays/v2-vs-v1/#v2-vs-v1
- [definition] The spec: v1 → v2 — The spec's structure did not change: an optional BOM, an H1, a blockquote, free markdown without headings, then H2 sections holding `- [name](url): notes` lines. What changed is which parts are required, where the file may live, and how a consumer is expected to use it. — https://llms-explorer.com/essays/v2-vs-v1/#the-spec-v1-v2
- [definition] The pipeline: V1 → V2 — The hub's V1 pipeline produced site dumps. It crawled with trafilatura, wrote one banner mirror per site, distilled that mirror with a zero-LLM bulk pass, and indexed the raw text in one vector layer. — https://llms-explorer.com/essays/v2-vs-v1/#the-pipeline-v1-v2
- [definition] Migration — For a **publisher** with a v1 file: 1. Run the migrate check (`llmsx migrate <url|file>`; today, `llms_lint.py check <file>` with `--check-links`). — https://llms-explorer.com/essays/v2-vs-v1/#migration
- [definition] Compatibility matrix — Rows are producer choices; columns are consumers. The cells are dated evidence (verified 2026-08-30) and need re-checking every 90 days, because consumer behaviour is the part of this table nobody controls. — https://llms-explorer.com/essays/v2-vs-v1/#compatibility-matrix
- [definition] What breaks — Honest list of what does not survive the two transitions. — https://llms-explorer.com/essays/v2-vs-v1/#what-breaks

## The vocabulary file
<https://llms-explorer.com/essays/vocabulary/>

- [snippet] The line grammar: > <n> terms of <family>; canonical name, definition, how it differs (not:), what people say instead (aka:). Each line an — `# <Family> — vocabulary` — https://llms-explorer.com/essays/vocabulary/#the-line-grammar
- [snippet] The line grammar: - **llms-small.txt** [llms.small] (noun): the budgeted variant of a full file — reference-class pages first, within abou — `- **llms-small.txt** [llms.small] (noun): the budgeted variant of a full file — reference-class page` — https://llms-explorer.com/essays/vocabulary/#the-line-grammar
- [snippet] The line grammar: - **llms-small.txt** — llms-small.txt is a small variant of a tokenized text file used to enforce size budgets on the pr — `- **llms-small.txt** — llms-small.txt is a small variant of a tokenized text file used to enforce si` — https://llms-explorer.com/essays/vocabulary/#the-line-grammar
- [snippet] Build one: PYTHONPATH=scripts .venv/bin/python -m docset_refine vocabulary \ — `PYTHONPATH=scripts .venv/bin/python -m docset_refine vocabulary \` — https://llms-explorer.com/essays/vocabulary/#build-one
- [parameter] `**term**`: required=yes; comes from=tree node or canonical token; rule=one line per term per sense — https://llms-explorer.com/essays/vocabulary/#the-line-grammar
- [parameter] `[sense-id]`: required=in a multi-family file; comes from=`<family-slug>.<term-slug>`; rule=disambiguates the pair (term × family) — https://llms-explorer.com/essays/vocabulary/#the-line-grammar
- [parameter] `(pos)`: required=no; comes from=part of speech; rule=noun unless stated — https://llms-explorer.com/essays/vocabulary/#the-line-grammar
- [parameter] `definition`: required=for a `## Terms` line; comes from=a kept unit; rule=must be extractive; its anchor is the line's source — https://llms-explorer.com/essays/vocabulary/#the-line-grammar
- [parameter] `— url#anchor`: required=with a definition; comes from=the unit's source; rule=resolves to a heading on the page (P7) — https://llms-explorer.com/essays/vocabulary/#the-line-grammar
- [parameter] `aka:`: required=no; comes from=surface forms in the pool; rule=never imported; the FTS5 layer expands through these — https://llms-explorer.com/essays/vocabulary/#the-line-grammar
- [parameter] `not:` … `— how`: required=no; comes from=contrast cues; rule=the neighbour and one clause on the difference — https://llms-explorer.com/essays/vocabulary/#the-line-grammar
- [parameter] `ant:`: required=no; comes from=explicit antonyms; rule=proposed extension — https://llms-explorer.com/essays/vocabulary/#the-line-grammar
- [parameter] `broader:` / `narrower:` / `related:`: required=no; comes from=the abstractor's relation taxonomy; rule=proposed extension — https://llms-explorer.com/essays/vocabulary/#the-line-grammar
- [parameter] `measure:`: required=no; comes from=the unit a quantity is stated in; rule=proposed extension — https://llms-explorer.com/essays/vocabulary/#the-line-grammar
- [parameter] `field:`: required=no; comes from=the family slug; rule=redundant with the sense id; kept for grep — https://llms-explorer.com/essays/vocabulary/#the-line-grammar
- [parameter] `verified-as-of:`: required=no; comes from=an actual re-fetch; rule=a date bump without a fetch is not evidence — https://llms-explorer.com/essays/vocabulary/#the-line-grammar
- [parameter] **assignment** — the topical builder's keyword pass: what it takes=`aka:` lists, merged into the concept-tree node's `aliases` by `--register` (add-only); what changes=a fact that says "session cookie" is filed under the node named "cookie" instead of falling to `## Shared` — https://llms-explorer.com/essays/vocabulary/#where-it-feeds
- [parameter] **keyword** — the FTS5 layer: what it takes=`aka:` surfaces of a matched term, OR-ed into the query (**designed**: an `expand` flag on `hub_query_docset`, which today takes only `docset, question, top, layer, mode`); what changes=an exact-token search for `X-Markdown-Tokens` would also find lines that wrote "the tokens header" — https://llms-explorer.com/essays/vocabulary/#where-it-feeds
- [parameter] **descriptions** — the index exporter: what it takes=the canonical definition; what changes=the one-liner after a link in `llms.txt` is the definition the pool agreed on, not a generated paraphrase — https://llms-explorer.com/essays/vocabulary/#where-it-feeds
- [definition] The vocabulary file — llms-vocabulary.txt is the lexical layer of a family: the words a field uses, what each means here, how it differs from its neighbours, and what people say instead. The line grammar, the sense model, where it feeds, and how to build one. — https://llms-explorer.com/essays/vocabulary/#the-vocabulary-file
- [definition] What a vocabulary file is — `llms-vocabulary.txt` is one line per term of a family, each line carrying: the canonical name, a definition taken from a kept unit, the neighbours it is easy to confuse it with (`not:`) and how it differs, the words people say instead (`aka:`), and the URL of the unit the definition came from. — https://llms-explorer.com/essays/vocabulary/#what-a-vocabulary-file-is
- [definition] The line grammar — The full grammar, with every optional field shown: — https://llms-explorer.com/essays/vocabulary/#the-line-grammar
- [definition] Senses across fields — A sense id is `<family-slug>.<term-slug>`. A term is disambiguated by the pair (term × family): *cookie* in the `web` family is `web.cookie`, in a folklore family `folklore.cookie-monster`, in a recipe family `food.cookie`. — https://llms-explorer.com/essays/vocabulary/#senses-across-fields
- [definition] Where it feeds — The vocabulary was built because three consumers were weak without it: — https://llms-explorer.com/essays/vocabulary/#where-it-feeds
- [definition] Build one — The walkthrough below builds the llms.txt family's own vocabulary — the terms are *index, full, small, facts, twin, describedby, family, split root, unit, anchor* and their neighbours. It is the same procedure for any field. — https://llms-explorer.com/essays/vocabulary/#build-one

## Which layer answers which question
<https://llms-explorer.com/examples/decision-table/>

- [parameter] Orientation before any retrieval: what does this site cover, where do I start: layer=`llms.txt` (≤ 10 KB) then ≤ 2 hops to a `.md` twin; cost class=~3k tokens, 3 requests, 0 embeddings; recipe=recipe-01 — https://llms-explorer.com/examples/decision-table/#the-table
- [parameter] Orientation on a site whose index split into sections (`## Sections` present): layer=split root: root index → `<slug>/llms.txt` → page; cost class=~3–5k tokens, 3–4 requests; recipe=recipe-02 — https://llms-explorer.com/examples/decision-table/#the-table
- [parameter] An exact token: an env var, a flag, a header name, an error string: layer=keyword layer (`mode="keyword"`, FTS5/BM25) over `llms-facts.txt`; cost class=microseconds, 0 model tokens, 0 embeddings; recipe=recipe-03 — https://llms-explorer.com/examples/decision-table/#the-table
- [parameter] A paraphrased question, or mixed / unsure whether the words match the source: layer=hybrid (`mode="hybrid"`, RRF over keyword + vector), or vector alone (`layer="facts"`); cost class=1 embedding, 0 generation tokens; recipe=recipe-04 — https://llms-explorer.com/examples/decision-table/#the-table
- [parameter] An agent that must find the right page from an MCP client without a search index: layer=index-first via `hub_docset_index` → `sections` → section index → page; cost class=~2k tokens read per hop, 0 embeddings; recipe=recipe-05 — https://llms-explorer.com/examples/decision-table/#the-table
- [parameter] A scripted check or query from a shell or a CI step: layer=the `llmsx` CLI (today: the hub scripts it wraps); cost class=seconds; 0 model tokens for lint / keyword; recipe=recipe-06 — https://llms-explorer.com/examples/decision-table/#the-table
- [parameter] Citation-grade answers inside your own RAG store: layer=`llms-facts.txt` units, one document each, `url#anchor` as metadata; cost class=1 embedding per unit at ingest; ~845k tokens for a 191-page site; recipe=recipe-07 — https://llms-explorer.com/examples/decision-table/#the-table
- [parameter] Keeping a published file honest on every push: layer=the lint as a GitHub Action gate (exit 1 on High); cost class=~10 s per file; network only with `--check-links`; recipe=recipe-08 — https://llms-explorer.com/examples/decision-table/#the-table
- [parameter] Serving the files so agents and the lint can find them: layer=headers: `text/markdown`, `X-Markdown-Tokens`, `Link: rel="describedby"`, `rel="alternate"` on HTML; cost class=one config block; verify with `curl -I`; recipe=recipe-09 — https://llms-explorer.com/examples/decision-table/#the-table
- [parameter] Whole-corpus reasoning, offline and private, within a token budget: layer=a local hub: Ollama + indexer + keyword layer + `llms_serve.py`; `llms-small.txt` for budgeted reads; cost class=one machine; ~50k tokens per small read, 0 API spend; recipe=recipe-10 — https://llms-explorer.com/examples/decision-table/#the-table
- [parameter] One concept across many sources, disagreements visible: layer=a topical file (`/t/<slug>/`) built from a fact pool; cost class=minutes to build; `--no-embed` for 0 embeddings; recipe=recipe-11 — https://llms-explorer.com/examples/decision-table/#the-table
- [parameter] Disambiguation: which sense of a word this family means, and its aliases: layer=`llms-vocabulary.txt` senses and `aka:` expansion before FTS5; cost class=free: string match, 0 model tokens; recipe=recipe-12 — https://llms-explorer.com/examples/decision-table/#the-table
- [definition] Which layer answers which question — The decision table for the cookbook: match the shape of your question to the cheapest llms layer that answers it, then open the recipe. — https://llms-explorer.com/examples/decision-table/#which-layer-answers-which-question
- [definition] When the table is the wrong tool — If the question is "is this file any good", none of these rows apply — that is the [lint](/reference/passes/), not a retrieval. — https://llms-explorer.com/examples/decision-table/#when-the-table-is-the-wrong-tool

## Recipe 01 — Two hops with requests
<https://llms-explorer.com/examples/recipe-01/>

- [snippet] Steps: import re, requests — `import re, requests` — https://llms-explorer.com/examples/recipe-01/#steps
- [snippet] Expected output: https://llms-explorer.pages.dev/examples/recipe-09/ https://llms-explorer.pages.dev/examples/recipe-09.md — `https://llms-explorer.pages.dev/examples/recipe-09/ https://llms-explorer.pages.dev/examples/recipe-` — https://llms-explorer.com/examples/recipe-01/#expected-output
- [definition] Recipe 01 — Two hops with requests — Read a site's llms.txt, pick a page by its description, fetch the .md twin, answer. The baseline every other recipe is measured against. — https://llms-explorer.com/examples/recipe-01/#recipe-01-two-hops-with-requests

## Recipe 02 — Split root: follow a section index
<https://llms-explorer.com/examples/recipe-02/>

- [snippet] Steps: import re, requests — `import re, requests` — https://llms-explorer.com/examples/recipe-02/#steps
- [snippet] Expected output: ('Agent Sdk', 'agent-sdk/llms.txt') — `('Agent Sdk', 'agent-sdk/llms.txt')` — https://llms-explorer.com/examples/recipe-02/#expected-output
- [definition] Recipe 02 — Split root: follow a section index — When the root llms.txt has a ## Sections block, let the counts on each section line decide which section index to fetch before touching a page. — https://llms-explorer.com/examples/recipe-02/#recipe-02-split-root-follow-a-section-index

## Recipe 03 — Keyword layer from Claude Code
<https://llms-explorer.com/examples/recipe-03/>

- [snippet] Steps: hub_query_docset(docset="codeclaudecom__codeclaudecom", question="CLAUDE_CODE_SYNC_SKILLS", mode="keyword", top=3) — `hub_query_docset(docset="codeclaudecom__codeclaudecom", question="CLAUDE_CODE_SYNC_SKILLS", mode="ke` — https://llms-explorer.com/examples/recipe-03/#steps
- [snippet] Steps: { — `{` — https://llms-explorer.com/examples/recipe-03/#steps
- [snippet] Steps: hub_llms_full_read(key="code.claude.com__docs", page="https://code.claude.com/docs/en/env-vars") — `hub_llms_full_read(key="code.claude.com__docs", page="https://code.claude.com/docs/en/env-vars")` — https://llms-explorer.com/examples/recipe-03/#steps
- [snippet] Steps: { — `{` — https://llms-explorer.com/examples/recipe-03/#steps
- [snippet] Steps: .venv/bin/python scripts/docset_indexer.py keyword codeclaudecom__codeclaudecom "CLAUDE_CODE_SYNC_SKILLS" --layer facts — `.venv/bin/python scripts/docset_indexer.py keyword codeclaudecom__codeclaudecom "CLAUDE_CODE_SYNC_SK` — https://llms-explorer.com/examples/recipe-03/#steps
- [definition] Recipe 03 — Keyword layer from Claude Code — Find an exact token — an env var, a flag, an error string — with hub_query_docset(mode=\"keyword\"), then open the page it came from. Zero model tokens. — https://llms-explorer.com/examples/recipe-03/#recipe-03-keyword-layer-from-claude-code

## Recipe 04 — Hybrid: keyword and vector fused
<https://llms-explorer.com/examples/recipe-04/>

- [snippet] Steps: hub_query_docset( — `hub_query_docset(` — https://llms-explorer.com/examples/recipe-04/#steps
- [snippet] Steps: { — `{` — https://llms-explorer.com/examples/recipe-04/#steps
- [snippet] Steps: .venv/bin/python scripts/docset_indexer.py query   codeclaudecom__codeclaudecom "which environment variable downloads my — `.venv/bin/python scripts/docset_indexer.py query   codeclaudecom__codeclaudecom "which environment v` — https://llms-explorer.com/examples/recipe-04/#steps
- [definition] Recipe 04 — Hybrid: keyword and vector fused — For a paraphrased or uncertain question, mode=\"hybrid\" runs the keyword and vector legs and fuses them with reciprocal-rank fusion; legs == 2 tells you both agreed on a hit. — https://llms-explorer.com/examples/recipe-04/#recipe-04-hybrid-keyword-and-vector-fused

## Recipe 05 — Index-first agent over MCP
<https://llms-explorer.com/examples/recipe-05/>

- [snippet] Steps: hub_docset_index(docset="codeclaudecom__codeclaudecom") — `hub_docset_index(docset="codeclaudecom__codeclaudecom")` — https://llms-explorer.com/examples/recipe-05/#steps
- [snippet] Steps: { — `{` — https://llms-explorer.com/examples/recipe-05/#steps
- [snippet] Steps: hub_docset_index(docset="codeclaudecom__codeclaudecom", file="agent-sdk/llms.txt") — `hub_docset_index(docset="codeclaudecom__codeclaudecom", file="agent-sdk/llms.txt")` — https://llms-explorer.com/examples/recipe-05/#steps
- [snippet] Steps: { — `{` — https://llms-explorer.com/examples/recipe-05/#steps
- [definition] Recipe 05 — Index-first agent over MCP — hub_docset_index → read sections → the section's llms.txt → the page. The pattern a concept-tree node page uses to find a source without any search index. — https://llms-explorer.com/examples/recipe-05/#recipe-05-index-first-agent-over-mcp

## Recipe 06 — The llmsx CLI
<https://llms-explorer.com/examples/recipe-06/>

- [snippet] Steps: llmsx lint ./docs/llms.txt --json — `llmsx lint ./docs/llms.txt --json` — https://llms-explorer.com/examples/recipe-06/#steps
- [snippet] Steps: llmsx query code.claude.com "CLAUDE_CODE_SYNC_SKILLS" --mode keyword — `llmsx query code.claude.com "CLAUDE_CODE_SYNC_SKILLS" --mode keyword` — https://llms-explorer.com/examples/recipe-06/#steps
- [snippet] Steps: llmsx export mirrors/code.claude.com.md — `llmsx export mirrors/code.claude.com.md` — https://llms-explorer.com/examples/recipe-06/#steps
- [snippet] Steps: llmsx tree show "llms.txt" — `llmsx tree show "llms.txt"` — https://llms-explorer.com/examples/recipe-06/#steps
- [snippet] Expected output: $ llmsx lint site/dist/llms.txt site/dist/llms-facts.txt --json | jq -c '.[] | {file, high: .counts.high}' — `$ llmsx lint site/dist/llms.txt site/dist/llms-facts.txt --json | jq -c '.[] | {file, high: .counts.` — https://llms-explorer.com/examples/recipe-06/#expected-output
- [definition] Recipe 06 — The llmsx CLI — Lint, query, export and inspect the tree from a shell: the llmsx commands and the hub scripts each one wraps today. — https://llms-explorer.com/examples/recipe-06/#recipe-06-the-llmsx-cli

## Recipe 07 — Facts into a RAG store
<https://llms-explorer.com/examples/recipe-07/>

- [snippet] Steps: import hashlib, re — `import hashlib, re` — https://llms-explorer.com/examples/recipe-07/#steps
- [snippet] Steps: import chromadb, requests — `import chromadb, requests` — https://llms-explorer.com/examples/recipe-07/#steps
- [snippet] Expected output: {'type': 'parameter', 'url': 'https://code.claude.com/docs/en/admin-setup', 'anchor': 'set-up-claude-code-for-your-organ — `{'type': 'parameter', 'url': 'https://code.claude.com/docs/en/admin-setup', 'anchor': 'set-up-claude` — https://llms-explorer.com/examples/recipe-07/#expected-output
- [definition] Recipe 07 — Facts into a RAG store — Parse llms-facts.txt with UNIT_RE, one document per unit with its url#anchor as metadata, embed with mxbai-embed-large — and never mix it with a 768-dimension model. — https://llms-explorer.com/examples/recipe-07/#recipe-07-facts-into-a-rag-store

## Recipe 08 — GitHub Action lint gate
<https://llms-explorer.com/examples/recipe-08/>

- [snippet] Steps: name: llms lint — `# .github/workflows/llms-lint.yml` — https://llms-explorer.com/examples/recipe-08/#steps
- [snippet] Expected output: [ — `[` — https://llms-explorer.com/examples/recipe-08/#expected-output
- [snippet] Expected output: {"pass": "P0", "attr": "I6", "severity": "high", "line": 0, — `{"pass": "P0", "attr": "I6", "severity": "high", "line": 0,` — https://llms-explorer.com/examples/recipe-08/#expected-output
- [definition] Recipe 08 — GitHub Action lint gate — Fail a pull request on any High finding in your llms files, and annotate the offending lines from the lint's JSON. — https://llms-explorer.com/examples/recipe-08/#recipe-08-github-action-lint-gate

## Recipe 09 — Serving with the right headers
<https://llms-explorer.com/examples/recipe-09/>

- [snippet] Steps: types { text/markdown md; } — `# inside the server {} block` — https://llms-explorer.com/examples/recipe-09/#steps
- [snippet] Steps: /*.md — `/*.md` — https://llms-explorer.com/examples/recipe-09/#steps
- [snippet] Steps: curl -sI https://docs.example.com/reference/attributes.md | grep -iE '^(content-type|x-markdown-tokens|link):' — `curl -sI https://docs.example.com/reference/attributes.md | grep -iE '^(content-type|x-markdown-toke` — https://llms-explorer.com/examples/recipe-09/#steps
- [snippet] Expected output: content-type: text/markdown; charset=utf-8 — `content-type: text/markdown; charset=utf-8` — https://llms-explorer.com/examples/recipe-09/#expected-output
- [snippet] Expected output: link: </reference/attributes.md>; rel="alternate"; type="text/markdown" — `link: </reference/attributes.md>; rel="alternate"; type="text/markdown"` — https://llms-explorer.com/examples/recipe-09/#expected-output
- [definition] Recipe 09 — Serving with the right headers — nginx and Cloudflare _headers blocks that serve .md twins as text/markdown with X-Markdown-Tokens and the two Link relations, verified with curl -I. — https://llms-explorer.com/examples/recipe-09/#recipe-09-serving-with-the-right-headers

## Recipe 10 — A local hub in miniature
<https://llms-explorer.com/examples/recipe-10/>

- [snippet] Steps: ollama pull mxbai-embed-large — `ollama pull mxbai-embed-large` — https://llms-explorer.com/examples/recipe-10/#steps
- [snippet] Steps: python3 scripts/llms_acquire.py https://code.claude.com/docs mirrors/code.claude.com.md — `python3 scripts/llms_acquire.py https://code.claude.com/docs mirrors/code.claude.com.md` — https://llms-explorer.com/examples/recipe-10/#steps
- [snippet] Steps: PYTHONPATH=scripts .venv/bin/python -m docset_refine all --no-units mirrors/code.claude.com.md — `PYTHONPATH=scripts .venv/bin/python -m docset_refine all --no-units mirrors/code.claude.com.md` — https://llms-explorer.com/examples/recipe-10/#steps
- [snippet] Steps: .venv/bin/python scripts/docset_indexer.py index mirrors/code.claude.com.md --name code.claude.com — `.venv/bin/python scripts/docset_indexer.py index mirrors/code.claude.com.md --name code.claude.com` — https://llms-explorer.com/examples/recipe-10/#steps
- [snippet] Steps: .venv/bin/python scripts/llms_serve.py --host 127.0.0.1 --port 8788 — `.venv/bin/python scripts/llms_serve.py --host 127.0.0.1 --port 8788` — https://llms-explorer.com/examples/recipe-10/#steps
- [snippet] Steps: .venv/bin/python scripts/docset_indexer.py keyword code.claude.com "CLAUDE_CODE_SYNC_SKILLS" --layer facts --mode phrase — `.venv/bin/python scripts/docset_indexer.py keyword code.claude.com "CLAUDE_CODE_SYNC_SKILLS" --layer` — https://llms-explorer.com/examples/recipe-10/#steps
- [snippet] Expected output: HTTP/1.0 200 OK — `HTTP/1.0 200 OK` — https://llms-explorer.com/examples/recipe-10/#expected-output
- [definition] Recipe 10 — A local hub in miniature — Ollama, the docset indexer, a keyword layer and llms_serve.py on one machine: the whole retrieval stack for one family, offline and private. — https://llms-explorer.com/examples/recipe-10/#recipe-10-a-local-hub-in-miniature

## Recipe 11 — Building a topical file
<https://llms-explorer.com/examples/recipe-11/>

- [snippet] Steps: PYTHONPATH=scripts .venv/bin/python -m docset_refine topical \ — `PYTHONPATH=scripts .venv/bin/python -m docset_refine topical \` — https://llms-explorer.com/examples/recipe-11/#steps
- [snippet] Steps: .venv/bin/python scripts/llms_lint.py check llms-topical/prompt-caching.llms/ --json — `.venv/bin/python scripts/llms_lint.py check llms-topical/prompt-caching.llms/ --json` — https://llms-explorer.com/examples/recipe-11/#steps
- [definition] Recipe 11 — Building a topical file — docset_refine topical turns a fact pool into a concept-axis llms.txt + llms-facts.txt with the subject's child concepts as sections; then /ldo --agent-test checks an agent can actually use it. — https://llms-explorer.com/examples/recipe-11/#recipe-11-building-a-topical-file

## Recipe 12 — Reading a vocabulary
<https://llms-explorer.com/examples/recipe-12/>

- [snippet] Steps: import re, requests — `import re, requests` — https://llms-explorer.com/examples/recipe-12/#steps
- [snippet] Expected output: ("/llms.txt" OR "/llmstxt" OR "llms.txt") discovery — `("/llms.txt" OR "/llmstxt" OR "llms.txt") discovery` — https://llms-explorer.com/examples/recipe-12/#expected-output
- [snippet] Expected output: - **cookie** [web.cookie] · [folklore.cookie-monster] · [food.cookie]: … — `- **cookie** [web.cookie] · [folklore.cookie-monster] · [food.cookie]: …` — https://llms-explorer.com/examples/recipe-12/#expected-output
- [definition] Recipe 12 — Reading a vocabulary — Expand a query through a family's aka: list before the FTS5 lookup, and pin the sense the family means. Free: string matching, no model. — https://llms-explorer.com/examples/recipe-12/#recipe-12-reading-a-vocabulary

## Abstracting one concept out of many docsets
<https://llms-explorer.com/blog/abstracting-one-concept/>

- [snippet] Commands: S=~/.claude/skills/llms-concept-abstractor/scripts/concept_abstract.py — `# cwd: ~/.global-ai-hub  (the skill's script; /lca wraps these steps for an agent)` — https://llms-explorer.com/blog/abstracting-one-concept/#commands
- [parameter] eval-1 with skill: Model tokens=336,335; Wall time=718 s; Grade=7/7 — https://llms-explorer.com/blog/abstracting-one-concept/#outputs
- [parameter] eval-1 baseline (ordinary tools): Model tokens=376,182; Wall time=571 s; Grade=3/7 — https://llms-explorer.com/blog/abstracting-one-concept/#outputs
- [parameter] eval-2 with skill (scope discovery): Model tokens=496,462; Wall time=2,281 s; Grade=6/6 — https://llms-explorer.com/blog/abstracting-one-concept/#outputs
- [definition] Abstracting one concept out of many docsets — /lca pulls 'indexing' out of nine database docsets and 'prompt caching' out of three API docs: lexicon expansion, a zero-token harvest, borderline classification, a facet-grouped pack — and what the evals measured. — https://llms-explorer.com/blog/abstracting-one-concept/#abstracting-one-concept-out-of-many-docsets

## Anchors that point nowhere
<https://llms-explorer.com/blog/anchors-that-point-nowhere/>

- [snippet] Commands: .venv/bin/python scripts/llms_lint.py check text-mirror/code.claude.com.llms/ \ — `# cwd: ~/.global-ai-hub` — https://llms-explorer.com/blog/anchors-that-point-nowhere/#commands
- [definition] Anchors that point nowhere — 1,124 of 11,965 units on the pilot were anchored to headings the site never renders — MDX <Step> and <Tab> titles that cleaning had turned into headings. The fix anchors every unit to the nearest real source heading. — https://llms-explorer.com/blog/anchors-that-point-nowhere/#anchors-that-point-nowhere

## Turning a customer's docs into an llms family
<https://llms-explorer.com/blog/customer-docs-to-llms-family/>

- [snippet] Commands: .venv/bin/python scripts/docset_rollout.py probe — `# cwd: ~/.global-ai-hub` — https://llms-explorer.com/blog/customer-docs-to-llms-family/#commands
- [parameter] developers.cloudflare.com: Pages=1,943; Acquired via=`llms-full.txt` (57 MB upstream, 2,000-page cap); Deterministic units=25,142 — https://llms-explorer.com/blog/customer-docs-to-llms-family/#inputs
- [parameter] developer.paypal.com: Pages=1,507; Acquired via=structured crawl (its `llms-full.txt` redirects to a 1.5 KB `llms.txt`); Deterministic units=38,710 — https://llms-explorer.com/blog/customer-docs-to-llms-family/#inputs
- [parameter] docs.claude.com (served from platform.claude.com): Pages=666; Acquired via=`llms.txt` + page `.md` twins; Deterministic units=13,432 — https://llms-explorer.com/blog/customer-docs-to-llms-family/#inputs
- [parameter] docs.langchain.com: Pages=529; Acquired via=`llms-full.txt`; Deterministic units=12,933 — https://llms-explorer.com/blog/customer-docs-to-llms-family/#inputs
- [parameter] developers.cloudflare.com: Root index (bytes)=9,241; Spoke indexes=243; Full (tokens)=4,162,267; Facts (tokens)=1,889,300 — https://llms-explorer.com/blog/customer-docs-to-llms-family/#outputs
- [parameter] developer.paypal.com: Root index (bytes)=4,104; Spoke indexes=193; Full (tokens)=2,921,259; Facts (tokens)=1,680,485 — https://llms-explorer.com/blog/customer-docs-to-llms-family/#outputs
- [parameter] docs.claude.com: Root index (bytes)=1,977; Spoke indexes=73; Full (tokens)=7,493,540; Facts (tokens)=768,209 — https://llms-explorer.com/blog/customer-docs-to-llms-family/#outputs
- [parameter] docs.langchain.com: Root index (bytes)=1,508; Spoke indexes=15; Full (tokens)=1,552,458; Facts (tokens)=738,488 — https://llms-explorer.com/blog/customer-docs-to-llms-family/#outputs
- [definition] Turning a customer's docs into an llms family — A product docset becomes index / full / small / facts, split hub-and-spoke at 10 KB — Cloudflare, PayPal, Claude and LangChain, with the real byte and token counts. — https://llms-explorer.com/blog/customer-docs-to-llms-family/#turning-a-customers-docs-into-an-llms-family

## Hub-and-spoke indexes
<https://llms-explorer.com/blog/hub-and-spoke-indexes/>

- [snippet] Commands: PYTHONPATH=scripts .venv/bin/python -m docset_refine export text-mirror/developers.cloudflare.com.md — `# cwd: ~/.global-ai-hub` — https://llms-explorer.com/blog/hub-and-spoke-indexes/#commands
- [parameter] developers.cloudflare.com: Pages=1,943; Root index (bytes)=9,241; Spokes=243 — https://llms-explorer.com/blog/hub-and-spoke-indexes/#inputs
- [parameter] developer.paypal.com: Pages=1,507; Root index (bytes)=4,104; Spokes=193 — https://llms-explorer.com/blog/hub-and-spoke-indexes/#inputs
- [parameter] docs.claude.com: Pages=666; Root index (bytes)=1,977; Spokes=73 — https://llms-explorer.com/blog/hub-and-spoke-indexes/#inputs
- [parameter] docs.langchain.com: Pages=529; Root index (bytes)=1,508; Spokes=15 — https://llms-explorer.com/blog/hub-and-spoke-indexes/#inputs
- [parameter] code.claude.com: Pages=191; Root index (bytes)=1,136; Spokes=6 — https://llms-explorer.com/blog/hub-and-spoke-indexes/#inputs
- [parameter] mongodb.com: Pages=82; Root index (bytes)=3,624; Spokes=30 — https://llms-explorer.com/blog/hub-and-spoke-indexes/#inputs
- [definition] Hub-and-spoke indexes — Why the 10 KB index rule is a split rule and not a truncation rule: the root keeps one line per section with page and token counts, every section becomes a spec-v2 index of its own, nothing is dropped, and /ldo refuses to improve the prose. — https://llms-explorer.com/blog/hub-and-spoke-indexes/#hub-and-spoke-indexes

## Keyword plus vector: the cheap path
<https://llms-explorer.com/blog/keyword-plus-vector/>

- [snippet] Commands: .venv/bin/python scripts/docset_indexer.py keyword-index codeclaudecom__codeclaudecom --layer facts — `# cwd: ~/.global-ai-hub` — https://llms-explorer.com/blog/keyword-plus-vector/#commands
- [definition] Keyword plus vector: the cheap path — An FTS5 (BM25) table beside the embeddings: exact tokens like CLAUDE_CODE_SYNC_SKILLS or --append-system-prompt cost no embedding call, and a reciprocal-rank hybrid fixes the queries the vector layer ranks below troubleshooting rows. — https://llms-explorer.com/blog/keyword-plus-vector/#keyword-plus-vector-the-cheap-path

## Six months of hand-made llms files
<https://llms-explorer.com/blog/six-months-of-hand-made-llms/>

- [snippet] Commands: .venv/bin/python scripts/pipeline_manager.py run          # mirror (trafilatura) → distill → index — `# cwd: ~/.global-ai-hub` — https://llms-explorer.com/blog/six-months-of-hand-made-llms/#commands
- [parameter] code blocks and tab panels dropped: `**macOS, Linux, WSL:**` followed by nothing; 122 fences in 37k lines; `curl -fsSL` twice on a site whose install page is built on it — https://llms-explorer.com/blog/six-months-of-hand-made-llms/#inputs
- [parameter] site chrome kept: 22 % of non-blank lines are duplicates (28,740 unique of 37,033); one FAQ paragraph appears 53 times — https://llms-explorer.com/blog/six-months-of-hand-made-llms/#inputs
- [parameter] link-only lines: 3,144 bare `[text](url)` lines, 8.5 % of the file — https://llms-explorer.com/blog/six-months-of-hand-made-llms/#inputs
- [parameter] one page is 11 % of the mirror: `/docs/en/changelog`, 535 KB, no date structure left — https://llms-explorer.com/blog/six-months-of-hand-made-llms/#inputs
- [parameter] the "distilled" output: 4.65 MB against a 4.74 MB mirror: 17,816 bullets, punctuation scrubbed, regex-bucketed, consumed by nothing — https://llms-explorer.com/blog/six-months-of-hand-made-llms/#inputs
- [parameter] V1 raw trafilatura: Mirror=4,744,720 B; Pages=228; Code fences=122; `curl -fsSL` lines=2; Score=**11 / 20** — https://llms-explorer.com/blog/six-months-of-hand-made-llms/#outputs
- [parameter] V2 after `llms-full.txt` acquisition: Mirror=8,547,884 B; Pages=191; Code fences=5,250; `curl -fsSL` lines=36; Score=— — https://llms-explorer.com/blog/six-months-of-hand-made-llms/#outputs
- [parameter] V2 facts layer (11,965 units: 5,034 parameters, 3,573 definitions, 2,624 snippets, 380 changes, 354 LLM): Mirror=—; Pages=191; Code fences=—; `curl -fsSL` lines=—; Score=**14 / 20** (partial LLM pass) — https://llms-explorer.com/blog/six-months-of-hand-made-llms/#outputs
- [definition] Six months of hand-made llms files — What the ecosystem's llms files actually look like when you download 608 of them, what our own V1 pipeline was producing, and why the answer to both was a facts layer instead of a better site dump. — https://llms-explorer.com/blog/six-months-of-hand-made-llms/#six-months-of-hand-made-llms-files

## The lint that gates the estate
<https://llms-explorer.com/blog/the-lint-that-gates-the-estate/>

- [snippet] Commands: .venv/bin/python scripts/llms_lint.py detect text-mirror/mongodb.com.llms/llms.txt          # kind + grammar — `# cwd: ~/.global-ai-hub` — https://llms-explorer.com/blog/the-lint-that-gates-the-estate/#commands
- [parameter] P0 detect: Kind=all; What fails High=none — reports kind and grammar so the right passes run — https://llms-explorer.com/blog/the-lint-that-gates-the-estate/#outputs
- [parameter] P1 structure: Kind=index, family; What fails High=no H1; more than one H1 — https://llms-explorer.com/blog/the-lint-that-gates-the-estate/#outputs
- [parameter] P2 links: Kind=index, family; What fails High=a relative target that does not exist (spoke split), a link with no target — https://llms-explorer.com/blog/the-lint-that-gates-the-estate/#outputs
- [parameter] P3 descriptions: Kind=index; What fails High=— (Medium: empty, duplicate, restated title) — https://llms-explorer.com/blog/the-lint-that-gates-the-estate/#outputs
- [parameter] P5 size ladder: Kind=all; What fails High=an index over 100,000 bytes — a full file wearing the wrong name — https://llms-explorer.com/blog/the-lint-that-gates-the-estate/#outputs
- [parameter] P6 full-file fidelity: Kind=full; What fails High=a grammar detected but zero page blocks parsed — https://llms-explorer.com/blog/the-lint-that-gates-the-estate/#outputs
- [parameter] P7 facts shape: Kind=facts; What fails High=a line with no source URL; a type outside the twelve; no unit lines at all — https://llms-explorer.com/blog/the-lint-that-gates-the-estate/#outputs
- [parameter] P9 provenance, rights and steering: Kind=all; What fails High=a real credential or PEM key body in copied text (attribute `P5`); third-party full text with no `<!-- internal -->` marker (attribute `P3`). A suspected instruction to the reading model is attribute `P4` and only a Medium — the model pass confirms it — https://llms-explorer.com/blog/the-lint-that-gates-the-estate/#outputs
- [parameter] P14 hygiene: Kind=all; What fails High=never High (excluded from Medium+ credit) — https://llms-explorer.com/blog/the-lint-that-gates-the-estate/#outputs
- [definition] The lint that gates the estate — llms_lint.py runs the deterministic passes of /ldo and exits 1 on any High; docset_rollout cleanup now runs it across 15 docsets and 652 files at 0 High — and what calibrating it on real docs taught about placeholder keys, PEM headers and quoted injection phrases. — https://llms-explorer.com/blog/the-lint-that-gates-the-estate/#the-lint-that-gates-the-estate

## A topical llms file from a pool of facts
<https://llms-explorer.com/blog/topical-llms-from-a-fact-pool/>

- [snippet] Commands: PYTHONPATH=scripts .venv/bin/python -m docset_refine topical \ — `# cwd: ~/.global-ai-hub` — https://llms-explorer.com/blog/topical-llms-from-a-fact-pool/#commands
- [parameter] `llms.txt`: Bytes=6,144; Tokens=1,523 — https://llms-explorer.com/blog/topical-llms-from-a-fact-pool/#outputs
- [parameter] `llms-facts.txt`: Bytes=74,210; Tokens=18,271 — https://llms-explorer.com/blog/topical-llms-from-a-fact-pool/#outputs
- [parameter] `llms-vocabulary.txt`: Bytes=10,313; Tokens=2,532 — https://llms-explorer.com/blog/topical-llms-from-a-fact-pool/#outputs
- [parameter] keyword match on section name / aliases: 30 — https://llms-explorer.com/blog/topical-llms-from-a-fact-pool/#outputs
- [parameter] file affinity (the spoke the fact came from): 122 — https://llms-explorer.com/blog/topical-llms-from-a-fact-pool/#outputs
- [parameter] embedding nearest-centroid: 9 — https://llms-explorer.com/blog/topical-llms-from-a-fact-pool/#outputs
- [parameter] `## Shared`: 7 — https://llms-explorer.com/blog/topical-llms-from-a-fact-pool/#outputs
- [definition] A topical llms file from a pool of facts — docset_refine topical builds sections from a concept-tree node's children and files every fact by keyword, then file affinity, then embedding centroid, then ## Shared — the llms.txt family pilot, with the assignment counts. — https://llms-explorer.com/blog/topical-llms-from-a-fact-pool/#a-topical-llms-file-from-a-pool-of-facts

## Your account
<https://llms-explorer.com/account/>

- [definition] Your account — Who you are signed in as, which plan you are on, and the sign-in methods and private tree forks attached to the account — all fetched in the browser. — https://llms-explorer.com/account/#your-account
- [definition] What the account holds — Three things the public site has no place for: the plan and its quotas, the API keys that authenticate the hosted MCP endpoint, and the private tree forks whose changes are proposed back rather than published. Deleting the account revokes every key with it. — https://llms-explorer.com/account/#what-the-account-holds

## Semantic indexing, recorded
<https://llms-explorer.com/demo/>

- [definition] Semantic indexing, recorded — One question set run three ways against one indexed docset — keyword (BM25), vector, and the fusion of both — hits and timings as recorded. — https://llms-explorer.com/demo/#semantic-indexing-recorded

## The directory of known llms files
<https://llms-explorer.com/directory/>

- [definition] The directory of known llms files — Every mirrored llms-full.txt that splits into pages, scored against the attribute rubric by llms_lint and graded A–F. — https://llms-explorer.com/directory/#the-directory-of-known-llms-files

## This site's llms family
<https://llms-explorer.com/family/>

- [definition] This site's llms family — The five files an agent reads, what each one is for, and the index rendered as clickable links rather than the raw text/markdown a browser cannot follow. — https://llms-explorer.com/family/#this-sites-llms-family
- [definition] What is on it — A table of the five members and what each is for, the index fetched and rendered with its links clickable, and a note on the `.md` twin every content page publishes beside itself. — https://llms-explorer.com/family/#what-is-on-it

## API keys
<https://llms-explorer.com/keys/>

- [definition] API keys — Create, list and revoke the scoped keys that authenticate the hosted MCP endpoint; the plaintext is shown once, at creation, and stored only as a hash. — https://llms-explorer.com/keys/#api-keys
- [definition] Shown once, stored hashed — What the API keeps is a non-secret lookup prefix and an Argon2id hash of the rest, so a key can be listed and revoked forever but never displayed twice. Losing one means issuing another and revoking the old, not recovering it. — https://llms-explorer.com/keys/#shown-once-stored-hashed

## Sign in
<https://llms-explorer.com/login/>

- [definition] Sign in — Sign in with a passkey, GitHub or Google; the API sets an HttpOnly session cookie that the account, keys and usage pages send back on every call. — https://llms-explorer.com/login/#sign-in
- [definition] Why an account exists — Only the metered surfaces need one — your own docsets, the hosted MCP endpoint, and private forks of the concept tree. Every published page, including the whole llms family, stays readable and unmetered without it. — https://llms-explorer.com/login/#why-an-account-exists

## concept-family-explorer
<https://llms-explorer.com/skills/concept-family-explorer/>

- [parameter] Parent / super-domain: What broader field is this a specialization of? — https://llms-explorer.com/skills/concept-family-explorer/#the-five-neighborhoods
- [parameter] Siblings: What sits at the same level under the same parent? — https://llms-explorer.com/skills/concept-family-explorer/#the-five-neighborhoods
- [parameter] Children / sub-concepts: What does this decompose into? — https://llms-explorer.com/skills/concept-family-explorer/#the-five-neighborhoods
- [parameter] Adjacent / cross-over: What neighboring domains overlap or interface here? — https://llms-explorer.com/skills/concept-family-explorer/#the-five-neighborhoods
- [parameter] Frontier / emerging: What is new, contested, or rising in this space? — https://llms-explorer.com/skills/concept-family-explorer/#the-five-neighborhoods
- [definition] concept-family-explorer — Gap-discovery layer above /dr — maps a subject's full conceptual family (parent, siblings, children, adjacent fields, frontier), scores what's missing, and researches every worthwhile gap to saturation. — https://llms-explorer.com/skills/concept-family-explorer/#concept-family-explorer
- [definition] The five neighborhoods — Every subject gets decomposed into five neighborhoods before anything is scored: — https://llms-explorer.com/skills/concept-family-explorer/#the-five-neighborhoods
- [definition] Where it sits — Breadth, not depth: it maps everything *around* a subject and stops at one useful pass per neighbor. Its narrow inverse — going *inside* one concept instead of around it — is [rabbithole](/skills/rabbithole/). — https://llms-explorer.com/skills/concept-family-explorer/#where-it-sits

## /dr — deep-research
<https://llms-explorer.com/skills/dr/>

- [definition] /dr — deep-research — Multi-source deep research using firecrawl and exa, synthesizing findings into cited reports with inline attribution, confidence ratings, and explicit knowledge gaps. — https://llms-explorer.com/skills/dr/#dr-deep-research

## full-suite
<https://llms-explorer.com/skills/full-suite/>

- [definition] full-suite — Exhaustively covers a subject end to end — maps the full concept family, researches every worthwhile gap to saturation, and compiles per-concept plus rollup llms-family files with keyword and semantic indexes. — https://llms-explorer.com/skills/full-suite/#full-suite

## notes-to-llms-txt
<https://llms-explorer.com/skills/notes-to-llms-txt/>

- [definition] notes-to-llms-txt — Turns disorganized notes — a scratch file, a run of meeting notes, a mixed-topic dump — into a well-formed llms.txt family, by segmenting, clustering by topic, and drafting a source-anchored entry per topic. — https://llms-explorer.com/skills/notes-to-llms-txt/#notes-to-llms-txt

## rabbithole
<https://llms-explorer.com/skills/rabbithole/>

- [definition] rabbithole — The narrow inverse of concept-family-explorer — takes one named concept and exhausts it completely, drilling down through mechanism, edge cases, and primary sources until a pass finds nothing new. — https://llms-explorer.com/skills/rabbithole/#rabbithole
- [definition] The six deepening questions — Every pass asks each of these against every claim still standing: 1. **Why is this true?** — the mechanism underneath a stated fact. — https://llms-explorer.com/skills/rabbithole/#the-six-deepening-questions
- [definition] Saturation, measured — Each pass's claims are diffed against the accumulated list from every prior pass. The **new-information rate** — new atomic claims this pass ÷ total claims after this pass — is recorded per pass, and the run stops after **two consecutive passes** each score below 5%. — https://llms-explorer.com/skills/rabbithole/#saturation-measured

## skill-tree-architect
<https://llms-explorer.com/skills/skill-tree-architect/>

- [definition] skill-tree-architect — Whole-tree architect for a skill library's hub-and-spoke taxonomy — audits description-cap headroom, hub balance, and cross-hub placement, then rebalances for a new family. — https://llms-explorer.com/skills/skill-tree-architect/#skill-tree-architect

## The concept tree
<https://llms-explorer.com/tree/>

- [definition] The concept tree — Every researched concept in the hub's tree, one page each, with its parent, its children and the frontier names below it. — https://llms-explorer.com/tree/#the-concept-tree

## Usage and credits
<https://llms-explorer.com/usage/>

- [definition] Usage and credits — The metered work on your account this period — jobs, tokens and embeddings, each row priced from the append-only ledger — and the credit balance left against your quota. — https://llms-explorer.com/usage/#usage-and-credits
- [definition] What metering counts — Model tokens on the refine and vocabulary passes, embedding calls on indexing, and the wall time of a job holding a worker. Querying an already-built index is not metered. — https://llms-explorer.com/usage/#what-metering-counts
