MDB Context Hub — Project Briefing
Published
Project article; see sources and editorial standards.
Version 1.0.39 · Node.js + TypeScript · Local MCP server · Internal Tool
I wrote this self-contained briefing for several audiences. Each section header names its primary audience: leadership sections use plain language, developer and reviewer sections use precise technical terms. Every command, path, version number, count and code reference comes from files in the repository, and the figures describe version 1.0.39 as of 17 June 2026.
1. Executive Summary (leadership)
The MDB Context Hub is a local-first Model Context Protocol (MCP) server and knowledge registry for MongoDB Technical Account Managers (TAMs), built by me, Mitch Hudson, a TAM. It is the backend behind the tam_* tool family: a single place that holds the skill guides, prompt templates, reference catalogs, and coding patterns a MongoDB TAM or Solutions Architect relies on. It exposes them, searchable and rankable, inside any MCP-capable AI coding client (Claude Code, Cursor, Gemini CLI).
The problem it solves is cold-start and tribal knowledge. Skill guides and prompts are otherwise scattered across repos and people’s heads, and every AI session starts blank. The hub centralizes them into registries (664 skills, 123 MCP tools across 23 domains, 282 saved prompts, a 431-node concept tree, 41 coding patterns, 15 roles, and 3 agents) and serves them locally from the operator’s own machine.
It is a local-first tool: it runs entirely on the operator’s machine (the HTTP transport binds to 127.0.0.1), stores no credentials, and holds no customer data. Customer-context ingestion is the sibling mdb-tam dashboard’s job and is deliberately out of scope here. The only outbound network calls are an optional Anthropic API call for prompt optimization, reachability checks for stale reference URLs, and a one-time download of the embedding model on first use.
I treat this as a production-quality internal tool, not a prototype: 251 automated tests across 8 suites, a single-workflow CI pipeline that gates build, skill-source validation, embedding-sync, and tool-inventory freshness, a documented STRIDE security model, and a documentation suite of about 33 Markdown files. Section 4 lists the coverage gaps.
2. Key Features (all)
- 664-skill library. Normalized skill manifests plus context docs, with explicit provenance and deduplication. Source split: 636 authored or researched locally and 28 sourced from
mdb-case-assistant. Categories: 325 developer, 255 custom, 84 mongodb. - 123 MCP tools across 23 domains. Skills, prompts, forge, URL/repo/shared/MCP libraries, meta (cross-catalog search and status), files (allowlisted file listing, metadata, and reads), file-analysis, telemetry, case-assistant, error-log, dependency-graph, prompt-variants, staleness, concept-tree, coding-patterns, operations-registry, tool-inventory, agents, roles, and tool-search (
docs/tool-inventory.json). - Two transports. stdio (
mcp-server/src/index.ts) and a loopback HTTP daemon on127.0.0.1:3939(mcp-server/src/http.ts,/mcpPOST and/healthzGET). - Skill recommendation and bundling.
tam_recommend_skillsranks skills against a natural-language query;tam_build_skill_bundleandtam_build_skill_bundle_fullassemble a ready-to-paste context bundle. - Semantic skill scoring. On-device embeddings via
Xenova/bge-small-en-v1.5(@huggingface/transformers);npm run embed:skillsbuilds them andnpm run embed:checkgates drift. No query or skill text leaves the machine; the model files are downloaded once from the Hugging Face Hub (section 6). - Prompt optimizer.
tam_optimize_promptinterprets and critiques a raw prompt, recommends skills and MCPs, optionally refines it through the Anthropic API, and auto-saves it toprompts/saved/when it looks reusable. - Prompt library. 282 entries in the saved-prompt registry (240 of kind
saved, 42 of kindworkflow) and 11 generated prompts (5 template and 6 bundle), with A/B prompt variants and comparison (tam_create_prompt_variant,tam_compare_prompt_variants). - Per-call telemetry. Every tool is wrapped by
instrumentedRegisterTool, recording status, duration, arguments with secret values masked, and suggested fixes tofile-analysis/telemetry.jsonl; surfaced throughtam_get_call_historyandtam_get_call_card. - Concept tree. A 431-node parent/child map (
concept-tree/tree.json) tying researched topics back to skills, with a generated visualization. - Coding-pattern catalog. 41 language- and category-tagged reusable patterns (
coding-patterns/registry.json). - Roles and agents. 15 roles that resolve to merged skill sets (
tam_role_resolve_skills) and 3 registered agents. - Reference libraries. 84 URL, 24 repo, 12 MCP-server, and 8 shared-library entries, each searchable and recommendable.
- Dependency, impact, and staleness analysis. A skill dependency graph with orphan and impact analysis, plus a staleness detector that HTTP-HEAD-checks reference-URL reachability.
- Operations registry. 40 named operations with a 5-standard external-call audit (
docs/operations-registry.json). It is a catalog and audit surface: no runners were wired at this version, sotam_ops_runreturnsskipped. - Sibling-repo integration. A generated telemetry feed for the
mdb-case-assistantdashboard (the registration patch and the dashboard-side polling were both still pending at this version; see section 6), and a strict-validation relay CLI into the case-assistant tool surface.
3. Problems Solved (leadership + team)
| Pain point | How the hub addresses it |
|---|---|
| Context switching: leaving the editor to find a runbook or skill guide | tam_recommend_skills and tam_build_skill_bundle surface ranked context in-editor from a natural-language query |
| Tribal knowledge: guidance lost on role change or departure | Skills, prompts, and catalogs are committed to a shared repo; every installer gets the same base |
| Prompt drift: each AI session starts blank, so output quality varies by operator | tam_optimize_prompt applies a consistent structure and saves reusable prompts to prompts/saved/ |
| Manual lookups: which internal repo, shared module, or MCP server applies | tam_recommend_repo_libraries / tam_recommend_shared_libraries / tam_recommend_mcps return scored recommendations |
| Onboarding friction: weeks to learn the toolset | The skill pack documents every tool with a description and a when-to-use, readable by humans and agents alike |
| Telemetry blindness: no record of which tool calls ran or failed | Every tool is wrapped by instrumentedRegisterTool; call history and call cards expose status and suggested fixes |
| Knowledge-coverage tracking: no map of what has been researched | The 431-node concept tree ties researched topics back to skills |
| Stale reference URLs: links rot silently | The staleness detector HTTP-HEAD-checks URL reachability (tam_staleness_scan) |
4. Scope of Work (leadership + reviewers)
I designed and built this project as an internal productivity tool. It is not a vendor product, prototype, or proof of concept. Its package.json declares the ISC license.
| Component | Path | Approx. lines |
|---|---|---|
| Service layer (all business logic) | mcp-server/src/service.ts |
~3,920 |
| Server / tool registration | mcp-server/src/server.ts |
~2,740 |
| Telemetry middleware | mcp-server/src/telemetry.ts |
~620 |
| Staleness detector | mcp-server/src/staleness.ts |
~540 |
| Operations registry | mcp-server/src/operations-registry.ts |
~530 |
| Dependency graph | mcp-server/src/dependency-graph.ts |
~390 |
| Prompt variants | mcp-server/src/prompt-variants.ts |
~260 |
| Remaining MCP-server modules | mcp-server/src/* |
balance of ~9,980 total |
| Sync pipeline | scripts/*.mjs |
~2,410 |
| Test suite | tests/*.ts |
~2,700 |
Line counts are raw file lines from wc -l, intended as scope indicators rather than SLOC.
Engineering quality markers:
- CI pipeline. A single GitHub Actions workflow (
.github/workflows/ci.yml, Node 20) gates, in order:npm run build(tsc) →node scripts/validate-skill-sources.mjs→npm run embed:check(embedding-sync) →npm test→npm audit --audit-level=moderate(non-blocking) → a tool-inventory freshness check that boots the server and fails on drift againstdocs/tool-inventory.json. - Tests. 251 test cases across 8 Vitest suites; the largest is
tests/mcp-server.test.ts(174 cases). Run withnpm test. The suites cover the service layer and the sync pipeline; the repository’s own docs (docs/TESTING.md,docs/external-calls.md) list the stdio and HTTP transports and the registry file watchers as untested. - Structured logging. Per-call records go to
file-analysis/telemetry.jsonlanderrors.jsonl; conventions are indocs/logging.mdanddocs/tool-call-telemetry.md. - Documentation suite. About 33 Markdown files in
docs/(excluding thedocs/superpowers/plans), including 6 operational runbooks underdocs/runbooks/.
5. Security Posture (reviewers + leadership)
Summary for reviewers: The hub is local-first. It stores and accepts no credentials. The one credential it touches, ANTHROPIC_API_KEY, is read from the environment for the optional prompt-refine step and sent only to the Anthropic API. The HTTP transport binds to loopback, and stdio is reachable only by the parent process that launches the server; the trust boundary is “loopback binding plus a trusted parent process.” All tool inputs are Zod-validated. The full STRIDE model is in docs/SECURITY.md.
Trust model
The hub trusts a single local user. There is no authentication, authorization, rate limiting, or audit logging by design: the security model is loopback binding plus a trusted parent process, nothing more. Per-call telemetry records the tool, status, and duration but not the caller, so it is not an audit log. On stdio, anything that can pipe to stdin can call any tool; over HTTP, any local process that can reach 127.0.0.1:3939 can call every tool, including the mutating save tools. This is acceptable only because the surface never leaves the machine.
Loopback-only binding
The HTTP transport binds to host = MDB_CONTEXT_HUB_HOST || '127.0.0.1' and port = MDB_CONTEXT_HUB_PORT || 3939 (mcp-server/src/http.ts). Requests that carry an Origin header are checked against the loopback origins and rejected with 403 otherwise. This mitigates DNS rebinding only and is not authentication. Binding the server off-loopback breaks the trust model entirely, because no auth layer exists behind it.
Secret handling
The repo stores no credentials. ANTHROPIC_API_KEY is read from the environment at call time for tam_optimize_prompt and is never persisted. The residual risk is that the free-form save tools (tam_save_prompt, tam_save_mcp_server, tam_save_url) write operator text to tracked files; .gitignore excludes .env and the telemetry/error logs but not the save catalogs, so review diffs before committing in case a secret was pasted.
Input validation
Every tool input is validated with Zod. Queries are capped at 500 characters, result limits at 100 items per call, and context reads at 20,000 characters; saved IDs are slugified to [a-z0-9-] (max 48 chars). The case-assistant relay CLI checks every argument before it runs: against a SAFE_ARG allowlist regex or, for --args= payloads, as valid JSON. It launches the CLI with execFile, so no shell is involved.
What this tool does not defend against
- A network-exposed deployment: binding off-loopback removes the only boundary there is.
- Multiple users or a malicious local process running as the same user.
- Secrets pasted into a save tool and then committed.
- Anything under the directories you list in
MDB_CONTEXT_HUB_FILE_ROOTS: the files tools are allowlisted by that variable, but any caller can then read those files (up to 10 MB each). - A compromised sibling
mdb-case-assistantrepo: the relay runs its code with the operator’s environment and credentials.
The full STRIDE threat model and trust boundary are in docs/SECURITY.md in the internal repository.
6. Architecture Overview (reviewers + team)
The hub is a layered, file-backed MCP server with a single in-memory registry cache. There is no database: the catalogs are git-tracked JSON and Markdown on disk.
Layered model
Upstream docs (mdb-case-assistant, mdb-tam) + local-sources/
│ scripts/sync-skill-pack.mjs (filesystem only, no network)
▼
Generated registries (skills/, prompts/, docs/ — "Do not edit by hand")
│ registry.ts (loadSkillPack → in-memory cache; HTTP hot-reload via fs.watch)
▼
service.ts (all MCP-facing operations: search, recommend, bundle, optimize, save)
│
server.ts (registers all 123 tools via instrumentedRegisterTool + Zod schemas)
│
├─ index.ts → stdio transport
└─ http.ts → StreamableHTTPServerTransport on 127.0.0.1:3939
Storage
Plain JSON and Markdown files in the git tree, with no MongoDB, Atlas, SQLite, or external database. Registry catalogs are git-tracked JSON arrays; per-call telemetry and errors are JSONL (gitignored).
Outbound network calls
tam_optimize_promptcalls the Anthropic API (modelclaude-haiku-4-5-20251001) for the optional LLM refine step.- The staleness detector sends HTTP HEAD reachability checks for reference URLs.
- The first time the embedder loads,
@huggingface/transformersdownloads the model files from the Hugging Face Hub and caches them (the library’s documented default, which the repo does not override).
Embedding inference itself runs on-device.
Sibling-repo integration
- mdb-case-assistant.
scripts/generate-case-assistant-registry.mjsemits one telemetry “external-call” card pertam_*tool. The dashboard is meant to polltam_get_call_historyand render them;docs/case-assistant-integration.mdrecords that the registration patch and the polling task were not yet applied in case-assistant at this version. A relay CLI (mcp-server/src/case-assistant-cli.ts) shells into the case-assistant CLI with strict argument validation. The MDB Case Assistant project pitch covers the case-assistant side. - mdb-tam (dashboard). The sync pipeline reads upstream docs from the dashboard repo by filesystem path. Customer-data ingestion and storage remain the dashboard’s responsibility.
Full diagrams and ADRs are in docs/ARCHITECTURE.md in the internal repository.
7. Installation & Quick Start (new users)
The steps below assume a checkout of the internal repository.
Prerequisites
- Node.js ≥ 20 (
package.jsonengines;.nvmrcpresent) - npm and Git
ANTHROPIC_API_KEYin the environment, only needed for thetam_optimize_promptrefine step
No Python, MongoDB, or database is required.
Install steps
- Run
npm installfrom the repository root. - Run
npm run build, thennpm test. - Start a transport:
npm run mcp:server:http(loopback daemon on127.0.0.1:3939) ornpm run mcp:server(stdio, for clients that spawn the server). - Point your client at it. The repository ships a
.mcp.jsonthat registersmdb_context_hubathttp://127.0.0.1:3939/mcp.
npm run sync:skills regenerates the registries from the upstream docs. It reads absolute source paths from scripts/skill-pack.config.mjs, so it runs only where those paths exist.