<!-- llms-explorer concept facts · https://llms-explorer.com/tree/atlas-diagnostics-expert/ · pack 2026-09-08 · ~5961 tokens -->

# Atlas Diagnostics Expert

> - SKIP (description-overflow seed, Glean 1000-char cap): WiredTiger storage-engine root-cause internals — cache-fill/eviction/checkpoint/MVCC mechanics behind a live perf symptom → mongodb-expert (ref

Parent: [MongoDB Expert Knowledge](https://llms-explorer.com/tree/mongodb-expert-knowledge/) · 19 facets · 92 facts · page: https://llms-explorer.com/tree/atlas-diagnostics-expert/

## Routing detail

- SKIP (description-overflow seed, Glean 1000-char cap): WiredTiger storage-engine root-cause internals - cache-fill/eviction/checkpoint/MVCC mechanics behind a live perf symptom → mongodb-expert (references/mongodb-wiredtiger-internals.md) — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#routing-detail)

## When to use this skill

- Atlas diagnostics and triage workflows — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#when-to-use-this-skill)
- internal single-pane triage tooling and adjacent internal support tools — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#when-to-use-this-skill)
- FTDC, log, metrics, alert, and explain-plan investigation choices — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#when-to-use-this-skill)
- KB-backed Atlas troubleshooting guidance — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#when-to-use-this-skill)
- Designing or reviewing new Atlas diagnostic tooling — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#when-to-use-this-skill)

## When NOT to use this skill

- Data-plane query/index/schema design not live perf troubleshooting - use mongodb-expert — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#when-not-to-use-this-skill)
- Atlas platform config/architecture (control plane, tiers, networking, security posture) - use mongodb-atlas-expert — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#when-not-to-use-this-skill)
- Backups, DR, migration, or security architecture - use mongodb-operations-expert — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#when-not-to-use-this-skill)
- KB article lookup - use misc-catch-all (references/mongodb-kb.md) — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#when-not-to-use-this-skill)

## Skill guidance

- Prefer documented Atlas diagnostic workflow before improvising. — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#skill-guidance)
- Call out what directly documented vs inferred when evidence thin. — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#skill-guidance)
- Use misc-catch-all (references/mongodb-kb.md) alongside when need article-level troubleshooting playbooks or customer-shareable links. — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#skill-guidance)

## Sub-skill routing table

- Consolidates 8 diagnostics/performance sub-skills as on-demand references - Read listed references/…md file before answering deep questions. — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#sub-skill-routing-table)

## Public Atlas docs

- Atlas metrics: <https://www.mongodb.com/docs/atlas/review-available-metrics/> — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#public-atlas-docs)
- Atlas alerts: <https://www.mongodb.com/docs/atlas/configure-alerts/> — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#public-atlas-docs)
- Performance Advisor: <https://www.mongodb.com/docs/atlas/performance-advisor/> — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#public-atlas-docs)

## Core principle

- Move from curated summary to raw evidence: — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#core-principle)
  - Start with a fast curated view (an internal single-pane triage tool, Atlas UI summaries, Performance Advisor, alerts, metrics) — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#core-principle)
  - Gather focused artifacts (logs, FTDC, explain plans, profiler samples) — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#core-principle)
  - Use specialized internal analyzers when first-pass evidence is insufficient — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#core-principle)
  - Package findings into repeatable escalation record using Atlas Diagnostic Checklist and Template — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#core-principle)

## What the internal docs establish directly

- An internal single-pane-of-glass tool is the first stop for Atlas project and cluster triage. — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#what-the-internal-docs-establish-directly)
- Atlas UI investigation still required for disk usage, IOPS, node state, query targeting, scan-and-order, oplog window, upgrade/election context. — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#what-the-internal-docs-establish-directly)
- Logs and FTDC are core raw artifacts behind deeper troubleshooting. — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#what-the-internal-docs-establish-directly)

## Atlas Diagnostic Checklist

- Purpose: Structured manual validation before escalation. — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#atlas-ui-atlas-diagnostic-checklist)
- Inputs: Project ID, node URI, cluster/node pages, logs/FTDC download links, observed symptoms and timestamps — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#atlas-ui-atlas-diagnostic-checklist)
- Outputs: Escalation-ready summary with cluster size, node status, storage, IOPS, CPU, oplog, query-targeting, restart attempts — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#atlas-ui-atlas-diagnostic-checklist)
  - Disk usage and write-blocking risk — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#atlas-ui-atlas-diagnostic-checklist)
  - Whether writes still accepted — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#atlas-ui-atlas-diagnostic-checklist)
  - Node down / recovering / upgrade state — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#atlas-ui-atlas-diagnostic-checklist)
  - Query targeting and scan-and-order behavior — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#atlas-ui-atlas-diagnostic-checklist)
  - Oplog window / fall-off risk — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#atlas-ui-atlas-diagnostic-checklist)
  - CPU pressure and OOM indicators — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#atlas-ui-atlas-diagnostic-checklist)
- Thresholds: Query targeting >100 red flag; >1000 urgent. Scan-and-order stay near 0; >25 warrants investigation. — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#atlas-ui-atlas-diagnostic-checklist)
- Cautions: Some node downtime during upgrades expected. Sanitize customer data before sharing log excerpts. — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#atlas-ui-atlas-diagnostic-checklist)
- Beyond the checklist and the Atlas UI/metrics/Performance Advisor surfaces, MongoDB support engineers also use a set of internal-only diagnostic tools (FTDC analyzers, log analyzers, Atlas Search explain-plan tooling, and an internal debugging-tool gateway) — not detailed here since they aren't publicly available. — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#atlas-diagnostic-checklist)

## Atlas metrics quick reference

  - High cache usage → working set or write pressure — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#atlas-metrics-quick-reference)
  - High disk latency / queue depth → storage bottleneck — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#atlas-metrics-quick-reference)
  - High connections → tier limits or pooling problem — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#atlas-metrics-quick-reference)
  - High execution time → query/index investigation — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#atlas-metrics-quick-reference)
- Atlas alerting RBAC-gated at org/project scope; severity levels: Critical, Error, Warning, Info. Alert state is diagnostic evidence, not just notification plumbing. — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#atlas-metrics-quick-reference)
- Performance Advisor works from slow-query evidence and suggests indexes based on query shape. Index recommendations still need read-vs-write tradeoff review before applying. — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#atlas-metrics-quick-reference)

## KB-guided troubleshooting posture

- Use KB for repeatable symptom-to-playbook mapping, especially when need customer-safe article or want to confirm known Atlas issue shape. Check visibility before sharing links externally. — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#kb-guided-troubleshooting-posture)
- Useful KB categories for Atlas diagnostics: — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#kb-guided-troubleshooting-posture)
  - Connection and TLS issues — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#kb-guided-troubleshooting-posture)
  - Oplog sizing / falling off the oplog — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#kb-guided-troubleshooting-posture)
  - Disk-usage interpretation — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#kb-guided-troubleshooting-posture)
  - Search/vector alert interpretation — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#kb-guided-troubleshooting-posture)
  - Network latency investigations — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#kb-guided-troubleshooting-posture)

## Standards for building new Atlas diagnostic tooling

- Prefer public Atlas Admin APIs first; use private/internal only when capability not exposed publicly. — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#standards-for-building-new-atlas-diagnostic-tooling)
- Decide consumer model up front: internal UI, CLI/programmatic tool, or agent-facing system. One API shape not fit every consumer. — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#standards-for-building-new-atlas-diagnostic-tooling)
- Use supported auth patterns: service accounts / OAuth, Digest for legacy Admin APIs, or approved internal auth flows. — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#standards-for-building-new-atlas-diagnostic-tooling)
- Make RBAC explicit - role annotations required, not implied. — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#standards-for-building-new-atlas-diagnostic-tooling)
- Add intentional rate limiting for fan-out or expensive diagnostic endpoints. — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#standards-for-building-new-atlas-diagnostic-tooling)
- Keep telemetry privacy-safe - avoid logging request/response bodies due to PII risk. — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#standards-for-building-new-atlas-diagnostic-tooling)
- Favor versioned and better-governed public APIs when long-term tool stability matters. — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#standards-for-building-new-atlas-diagnostic-tooling)
- Treat logs, FTDC, sample queries, and explains as potentially sensitive customer data; minimize storage and exposure. — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#standards-for-building-new-atlas-diagnostic-tooling)
- Preserve TS operational pattern: summary surface first, raw artifacts second, specialized analyzers third. — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#standards-for-building-new-atlas-diagnostic-tooling)

## Directly documented

- Atlas Diagnostic Checklist thresholds and escalation posture — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#directly-documented)
- Atlas Diagnostic Checklist thresholds and escalation posture — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#directly-documented)
- Atlas metrics / alerts / Performance Advisor high-level behavior — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#directly-documented)
- Internal API/auth/RBAC/privacy constraints (not detailed here) — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#directly-documented)
- Atlas metrics / alerts / Performance Advisor high-level behavior — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#directly-documented)
- Internal API/auth/RBAC/privacy constraints from MMS API Landscape — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#directly-documented)

## Lightly documented or partly inferred

  - Internal diagnostic-tool operating detail (not covered here — those tools aren't publicly available) — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#lightly-documented-or-partly-inferred)
  - Whether a given internal tool is currently recommended, maintained, or only historically available — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#lightly-documented-or-partly-inferred)
  - Whether given internal tool currently recommended, maintained, or only historically available — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#lightly-documented-or-partly-inferred)
- When extending this context, read tool's current README or operator guide before making prescriptive claims. — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#lightly-documented-or-partly-inferred)
- <!-- cross-hub-map --> — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#lightly-documented-or-partly-inferred)

## Cross-hub map — where every MongoDB topic lives

- All MongoDB knowledge split across four hubs (plus misc-catch-all for KB-article lookups via references/mongodb-kb.md). If task's deep material not in this hub's Sub-skill routing table, it is reference file under sibling hub - activate that hub or Read its references/<name>.md directly. — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#cross-hub-map-where-every-mongodb-topic-lives)
- High-overlap routing notes: — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#cross-hub-map-where-every-mongodb-topic-lives)
  - Performance symptom triage (high CPU, cache pressure, slow queries, latency spikes) starts at atlas-diagnostics-expert, but storage-engine root-cause internals (WiredTiger cache fill / dirty trigger / eviction threads / reconciliation / checkpoints) owned by mongodb-expert - cross-load mongodb-expert/references/mongodb-wiredtiger-internals.md (and mongodb-wiredtiger.md) for depth. — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#cross-hub-map-where-every-mongodb-topic-lives)
  - Migration symptoms vs migration execution: live-cluster diagnosis → atlas-diagnostics-expert; migration/mongosync runbook → mongodb-operations-expert. — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#cross-hub-map-where-every-mongodb-topic-lives)
  - Atlas Search/Vector query syntax & index design → mongodb-atlas-expert; slowness triage of running search → atlas-diagnostics-expert. — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#cross-hub-map-where-every-mongodb-topic-lives)
  - Host-OS memory tuning for self-managed mongod host (transparent hugepages disable - THP/defrag=never, vm.swappiness=1, swap sizing, kernel OOM killer and oom_score_adj, NUMA placement / interleave for WiredTiger cache, vm.max_map_count) lives under the devops-infra router's devops-linux-internals sub-hub → cross-load devops-linux-internals/references/linux-memory-numa.md. This skill owns MongoDB-side cache-pressure symptom triage; that reference owns Linux memory/NUMA mechanisms and sysctls beneath it. — [source](https://llms-explorer.com/sources/mdb-context-hub/atlas-diagnostics-expert/#cross-hub-map-where-every-mongodb-topic-lives)

## Where this helps

- Triaging a live Atlas cluster issue (outage, node-down, degraded performance) and needing to know which surface to reach for first — a fast curated summary, then logs/FTDC, then a specialized analyzer. — [source](https://llms-explorer.com/tree/atlas-diagnostics-expert/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Investigating a slow-query or query-targeting problem and deciding between the Atlas Performance Advisor, the profiler, or an Atlas Search/Vector Search explain-plan tool. — [source](https://llms-explorer.com/tree/atlas-diagnostics-expert/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Building an escalation-ready record for a support case using the Atlas Diagnostic Checklist's documented thresholds (query targeting >100 is a red flag, >1000 is urgent; scan-and-order should stay near 0). — [source](https://llms-explorer.com/tree/atlas-diagnostics-expert/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Deciding whether a symptom belongs in this diagnostic workflow or should route to a sibling hub — e.g. WiredTiger cache-fill/eviction root-cause internals go to mongodb-expert, not here. — [source](https://llms-explorer.com/tree/atlas-diagnostics-expert/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*

## How to apply this

- Follow the documented core principle: start with the fastest curated view (an internal triage tool, Atlas UI, Performance Advisor, alerts), then gather focused raw artifacts (logs, FTDC, explain plans), then reach for a specialized analyzer only if the first two passes are insufficient. — [source](https://llms-explorer.com/tree/atlas-diagnostics-expert/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Use the internal single-pane-of-glass tool as the first stop for Atlas project and cluster triage, but treat its quick diagnostics as a triage accelerator, not a replacement for validating against the Atlas UI, logs, and FTDC directly. — [source](https://llms-explorer.com/tree/atlas-diagnostics-expert/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Match the incident type to the documented workflow — cluster health/outage starts with the internal triage tool then the checklist; query/index issues start with Performance Advisor and the profiler; log-heavy incidents lean on internal log-analysis tooling. — [source](https://llms-explorer.com/tree/atlas-diagnostics-expert/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- When building new internal Atlas diagnostic tooling, follow this pack's own standards: prefer public Atlas Admin APIs first, make RBAC explicit, add rate limiting on expensive endpoints, and avoid logging request/response bodies due to PII risk. — [source](https://llms-explorer.com/tree/atlas-diagnostics-expert/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*

## Antipatterns

- Escalating a case straight to raw FTDC/log analysis without first checking the fast curated views (internal triage tool, Atlas UI, Performance Advisor) that would have answered the question in a fraction of the time. — [source](https://llms-explorer.com/tree/atlas-diagnostics-expert/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Sharing raw log excerpts with a customer or in an external ticket without sanitizing customer data first, which this pack explicitly flags as a caution. — [source](https://llms-explorer.com/tree/atlas-diagnostics-expert/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Treating a specialized analyzer's LLM-assisted output as ground truth without validating the hypothesis against the underlying explain data it was derived from. — [source](https://llms-explorer.com/tree/atlas-diagnostics-expert/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Building new internal diagnostic tooling that logs full request/response bodies for convenience, ignoring the PII-risk guidance this pack sets as a standard for new tooling. — [source](https://llms-explorer.com/tree/atlas-diagnostics-expert/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*

## Known issues

- Several internal tools in this pack's toolkit are only lightly documented or partly inferred, so treat guidance on them as directional. — [source](https://llms-explorer.com/tree/atlas-diagnostics-expert/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- It's not always clear from the corpus whether a given internal tool is currently recommended, actively maintained, or only historically available — the pack advises reading the tool's current README before making prescriptive claims. — [source](https://llms-explorer.com/tree/atlas-diagnostics-expert/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Access to most of these internal tools is RBAC-gated internally, so the workflow described here assumes access this pack doesn't itself grant. — [source](https://llms-explorer.com/tree/atlas-diagnostics-expert/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- This is a routing/triage hub, not the place for storage-engine root-cause internals — WiredTiger cache-fill, eviction, checkpoint, and MVCC mechanics behind a live performance symptom route to mongodb-expert instead. — [source](https://llms-explorer.com/tree/atlas-diagnostics-expert/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*

## Context files

- [Atlas Diagnostics Expert](https://llms-explorer.com/downloads/sources/mdb-context-hub/atlas-diagnostics-expert.md)
