<!-- llms-explorer concept facts · https://llms-explorer.com/tree/analytics-node-monitoring/ · pack 2026-10-02 · ~8964 tokens -->

# Analytics Node Monitoring

> Depth-first rabbithole dossier for Analytics Node Monitoring; source-anchored research pack.

Parent: [MongoDB Atlas Analytics Node](https://llms-explorer.com/tree/mongodb-atlas-analytics-node/) · 6 facets · 52 facts · page: https://llms-explorer.com/tree/analytics-node-monitoring/

## Definitions

- 44. On Atlas Infinite clusters, the oplog does not serve replication. Replication Lag there means the time for writes to become visible on the node. Replication Headroom and Opcounters – Repl apply to Atlas Core only. https://www.mongodb.com/docs/atlas/review-available-metrics.md 45. Atlas Infinite (public preview) supports analytics nodes and an analytics tier, but runs them in one region only. https://www.mongodb.com/docs/atlas/cluster-config/multi-cloud-distribution.md — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/mechanism.md#g-edition-boundary`

## How it works

- 42. With `secondaryPreferred` plus `nodeType:ANALYTICS`, BI reads first fell back to ordinary secondaries without anyone noticing. Later, when the analytics node stopped answering, the reads fell back to the primary. Primary read load rose, an election followed, and user-facing services degraded. *(Inherited parent host; not counted toward the source gate.)* — https://pmclain.com/mongodb/2022/09/28/fun-with-mongodb-atlas-analytics-nodes.html 43. **[inference]** An analytics-node incident can show up only as primary or electable-secondary opcounters. Watch the analytics node's own opcounters fo — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/edge-cases.md#h-misrouting-the-monitoring-symptom-shows-up-on-another-node`
- 34. Field report: a `secondaryPreferred` + `nodeType:ANALYTICS` connection string fell back to the primary when analytics nodes were unavailable. This raised primary read load and contributed to an election. The team fixed it with `readPreference=secondary`. — https://pmclain.com/mongodb/2022/09/28/fun-with-mongodb-atlas-analytics-nodes.html 35. Field report: a connection-string typo sent the "analytics" workload to operational secondaries for months before anyone noticed. — https://pmclain.com/mongodb/2022/09/28/fun-with-mongodb-atlas-analytics-nodes.html 36. **Implication:** monitoring shoul — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/practice.md#f-monitoring-the-isolation-itself`
- Eight specific gaps remain, listed in the file, such as the Admin API alert values and whether an M40+ analytics tier triggers premium monitoring. I also listed six sibling concepts for concept-family-explorer to pick up. The concept tree was not edited. — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/rabbithole-synthesis.md`
- 1. Atlas host alerts can target "all hosts or … specific type of host, such as primaries or config servers". The documented examples include no analytics host type. — https://www.mongodb.com/docs/atlas/reference/alert-conditions/ 2. In the alert-configuration matcher, `TYPE_NAME` accepts only `PRIMARY`, `SECONDARY`, `STANDALONE`, `CONFIG` and `MONGOS`. There is no `ANALYTICS` value. — https://raw.githubusercontent.com/mongodb/terraform-provider-mongodbatlas/master/docs/resources/alert_configuration.md 3. **[inference]** A Replication Lag alert scoped to `SECONDARY` covers both electable second — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/edge-cases.md#a-alert-targeting`
- **In scope:** how you observe the health and load of MongoDB Atlas analytics nodes. That covers metrics, replication lag, oplog window and headroom, alerts, the profiler and RTPP, analytics-tier auto-scaling signals, and when each capability appeared. — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/history.md#scope`

## Measurements and reference values

- 10. Analytics nodes scale up when average Normalized System CPU **or** System Memory Utilization exceeds 75% over the past hour. https://www.mongodb.com/docs/atlas/cluster-autoscaling-compute-core/ 11. Analytics nodes scale down when average CPU **and** memory stay below 50% over the past 24 hours. https://www.mongodb.com/docs/atlas/cluster-autoscaling-compute-core/ 12. "Atlas doesn't apply the Queued or Rejected Operations criterion to analytics nodes." Predictive auto-scaling also does not apply to analytics nodes. https://www.mongodb.com/docs/atlas/cluster-autoscaling-compute-core/ 13. You — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/history.md#c-auto-scaling-metrics-act-as-implicit-monitoring`
- - Passes: pass 0 (metrics, alerts, configuration pages) produced 28 claims. Pass 1 (the manual's lag mechanics, auto-scaling, RTPP, Query Profiler, Performance Advisor) added 13 new claims, a rate of 13/41 ≈ 32%. Pass 2 (matcher enum, Prometheus, Datadog, default-alert list) added 4 new claims and one disagreement, a rate of 4/45 ≈ 9%. - Verdict: **BUDGET_EXHAUSTED (single-run brief)**, not SATURATED-DEPTH. The rate is falling but has not had two consecutive passes under 5%. One more pass would likely still find: the exact Admin API process-metric names for analytics hosts, the new Metrics UI' — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/mechanism.md#saturation-and-quality-gate`
- 30. Analytics nodes scale up when average Normalized System CPU **or** System Memory Utilization stays above 75% for one hour. Operational M30+ nodes also scale on 90% for 10 minutes. — https://www.mongodb.com/docs/atlas/cluster-autoscaling-compute-core/ 31. Atlas does not apply the Queued or Rejected Operations (IWM load-shedding) criterion to analytics nodes. — https://www.mongodb.com/docs/atlas/cluster-autoscaling-compute-core/ 32. **[inference]** An analytics node can stay overloaded for up to about an hour before scaling starts. Burst BI queries need your own CPU and lag alerts. — source — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/edge-cases.md#f-auto-scaling-as-a-monitoring-signal`
- - New-information rate per pass: P0 = 1.00 (8 claims), P1 = 0.58 (+11), P2 = 0.47 (+17), P3 = 0.16 (+7). The rate is falling but never went below 5%. **Verdict: BUDGET_EXHAUSTED (soft stop), not SATURATED-DEPTH.** One or two more passes (Admin API alert schema, `maxStalenessSeconds` monitoring, Prometheus integration labels) would probably still add claims. - **Source gate: met for the mechanisms, not met for claims specific to analytics nodes.** The independent hosts are percona.com, docs.datadoghq.com, github.com (DataDog issue) and oneuptime.com. They corroborate flow control, lag measureme — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/edge-cases.md#saturation-and-source-gate`
- | Date | Event | Source | |---|---|---| | ~2019-06 (unverified) | MongoDB blog "Workload Isolation with Atlas and Charts" covers analytics nodes. The date comes from the Owler report-ID timestamp 1559878320932 ms, which is 2019-06-07; the page itself timed out. | https://www.owler.com/reports/mongodb/mongodb-blog-workload-isolation-with-atlas-and-cha/1559878320932 | | 2021-07 | Data Federation can target analytics nodes. | https://www.mongodb.com/docs/atlas/release-notes/changelog/ | | 2021-08-03 | 10-second premium metrics for projects with an M40+ cluster. This is the granularity that analyt — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/history.md#f-timeline`
- - **Pass rates** (new claims ÷ running total): pass 0 = 9/9; pass 1 = 8/17 (47%); pass 2 = 6/23 (26%); pass 3 = 1/24 (4%). - **Verdict: BUDGET_EXHAUSTED.** This was a soft stop, not saturation. Only one pass fell below 5%, and the tool restrictions blocked further primary fetches. - **Gaps that remain open:** - When analytics nodes were first introduced, and in which original announcement. The Owler and Plushcap mirrors returned a timeout and a 404. - Whether the Admin API alert matcher `TYPE_NAME` accepts an analytics value. The API spec page did not render. - Whether Performance Advisor runs — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/history.md#gaps-and-saturation`

## Problems, failure modes and limitations

- **Most useful additions over the parent:** - Lag is sampled only every 85 s, even with 10-second premium monitoring. - A lagging analytics node doesn't slow writes on the primary, because it doesn't vote. This is a strong inference; MongoDB never says it outright. - Auto-scaling ignores lag and acts only after an hour of high CPU or memory, so it can't stand in for lag alerts. — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/rabbithole-synthesis.md`
- 20. `w:"majority"` counts only data-bearing **voting** members (`votes > 0`). — https://www.mongodb.com/docs/manual/reference/write-concern/ 21. Flow control throttles the primary only to keep **majority-committed** lag under `flowControlTargetLagSeconds`. Lag can grow without flow control engaging, for example on an unresponsive secondary. — https://www.mongodb.com/docs/manual/tutorial/troubleshoot-replica-sets/ ; mechanism corroborated by https://www.percona.com/blog/mongodb-queries-slow-due-to-flow-control-but-no-replication-lag/ (Malkowski, 2023-10-12) 22. **[inference]** Analytics nodes h — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/edge-cases.md#d-no-backpressure-flow-control-ignores-analytics-lag`
- 1. Atlas collects metrics at 1-minute granularity by default. Replication Lag and Replication Headroom are the exception: Atlas collects them at an 85-second granularity. https://www.mongodb.com/docs/atlas/monitor-cluster-metrics.md 2. The Replication Lag and Replication Headroom charts keep the 85-second interval "regardless of your project's default granularity level". https://www.mongodb.com/docs/atlas/review-available-metrics.md 3. If any cluster in a project is M40 or larger, Atlas enables premium monitoring (10-second granularity) for every cluster in the project. https://www.mongodb.com — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/mechanism.md#a-metric-collection-mechanism`
- 11. Replication alert conditions are `Replication Lag is`, `Replication Oplog Window is`, `Replication Headroom is`, and `Oplog Data Per Hour is`. Headroom is the sync source's oplog window minus the secondary's replication lag. — https://www.mongodb.com/docs/atlas/reference/alert-conditions/ 12. The supported metric name is `OPLOG_REPLICATION_LAG_TIME`. `OPLOG_SLAVE_LAG_MASTER_TIME` is deprecated. — https://www.mongodb.com/docs/atlas/reference/alert-host-metrics/ 13. Replication Lag and Replication Headroom are sampled at 85-second granularity. Other metrics are sampled at 1 minute, or at 10 — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/practice.md#b-metrics-and-alert-conditions`
- **What this changes in the parent:** "separately monitored" is only partly true. Atlas tags analytics nodes in the Datadog export and has analytics-only auto-scaling alerts. But its own alert rules have no analytics node type, so alerts can't target analytics nodes. Nor can the Metrics page or Prometheus labels pick them out. — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/rabbithole-synthesis.md`
- 8. Atlas sends Datadog a `nodeType` tag with the values `ELECTABLE`, `READ_ONLY` or `ANALYTICS`. This allows per-analytics-node lag monitors that Atlas's own `TYPE_NAME` matcher cannot express. — https://www.mongodb.com/docs/atlas/tutorial/datadog-integration/ 9. The Datadog integration is M10+ only and push-based. It exports `mongodb.atlas.replset.replicationlag` and `mongodb.atlas.replset.oplograte`. — https://docs.datadoghq.com/integrations/mongodb-atlas/ 10. Datadog's `mongodb.atlas.replstatus.health` can be wrong in some cases. If every node is unresponsive, Datadog keeps reporting them a — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/edge-cases.md#b-third-party-export-datadog`
- 13. Atlas collects Replication Lag and Replication Headroom every 85 seconds, even though other metrics use 1-minute or 10-second (premium) granularity. — https://www.mongodb.com/docs/atlas/monitor-cluster-metrics/ 14. **[inference]** Lag spikes shorter than about 85 s may not show in the lag chart. Short-lived staleness on BI reads can go undetected. — source in 13. 15. Premium (10-second) monitoring turns on when the project has "at least one cluster that's M40 or larger". The docs do not say whether an M40+ analytics tier on an M30 base cluster qualifies. See Unresolved. — https://www.mongo — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/edge-cases.md#c-lag-measurement-and-granularity`
- 39. Secondaries log slow oplog application as `applied op: <oplog entry> took <num>ms` under `REPL`. The profiler does not capture these entries. — https://www.mongodb.com/docs/manual/tutorial/troubleshoot-replica-sets/ 40. Atlas raises a default-on alert when the Database Profiler is at level 2, or level 1 with `slowms <= 0`. That setting is a tempting but harmful way to catch every BI query. — https://www.mongodb.com/docs/atlas/reference/alert-conditions/ 41. Atlas keeps MongoDB log data at no more than 2000 lines per 2 minutes. **[inference]** On a busy analytics node, slow-query evidence c — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/edge-cases.md#g-slow-operations-on-the-analytics-node`
- 1. Atlas reports **Replication Lag** as "the approximate number of seconds the secondary is behind the primary in write application." It collects this metric every 85 seconds, whatever the project's granularity setting. https://www.mongodb.com/docs/atlas/review-available-metrics/ 2. **Replication Headroom** is the primary's oplog window minus the secondary's replication lag. A node needs a full resync if its lag exceeds the oplog window and its headroom reaches zero. https://www.mongodb.com/docs/atlas/review-available-metrics/ 3. The Atlas metrics reference describes these metrics only in term — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/history.md#a-replication-lag-is-the-main-health-signal-for-an-analytics-node`
- 15. The Metrics page shows servers through **Toggle Members**, which uses `P` and `S` icons. The docs do not describe an analytics-specific icon or filter. https://www.mongodb.com/docs/atlas/monitor-cluster-metrics/ 16. Standard granularity is 1 minute. Premium 10-second granularity turns on automatically for every cluster in a project that has an M40+ cluster. Lag and headroom stay at 85 seconds either way. https://www.mongodb.com/docs/atlas/monitor-cluster-metrics/ 17. Retention: 10-second data for 8 hours, 1-minute and 5-minute data for 48 hours, 1-hour data for 63 days, 1-day data forever. — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/history.md#d-per-node-observation-tools`
- 22. In one production incident, the client used `readPreference=secondaryPreferred&readPreferenceTags=nodeType:ANALYTICS`. When the analytics node stopped serving reads, the analytics load fell back onto the primary. That degraded the primary and caused a failover. Lag on the analytics node had gone unnoticed until then. https://pmclain.com/mongodb/2022/09/28/fun-with-mongodb-atlas-analytics-nodes.html 23. The same author says Atlas gives the operational secondary "higher priority resource access" than analytics nodes. On that account, analytics nodes see high lag under heavy write throughput. — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/history.md#e-practitioner-evidence-independent-and-disconfirming`
- The announcement blog for analytics tiers (2022-08-03) contains no monitoring or lag guidance. The lag warning exists only in the docs. https://www.mongodb.com/company/blog/product-release-announcements/introducing-ability-independently-scale-atlas-analytics-node-tiers — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/history.md#f-timeline`
- **In scope:** how Atlas observes analytics nodes, covering per-node metrics, replication-lag signals, alert targeting, slow-operation tooling, auto-scaling telemetry, and external metric export. It also covers the invariants and blind spots of each. — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/mechanism.md#scope`
- 8. Replication Lag is "the approximate number of seconds the secondary is behind the primary in write application". https://www.mongodb.com/docs/atlas/review-available-metrics.md 9. Replication Headroom is the primary's oplog window minus the secondary's replication lag. It predicts whether that member will fall off the oplog. https://www.mongodb.com/docs/atlas/review-available-metrics.md 10. A node needs a full resync if its lag exceeds the oplog window and its headroom reaches zero. https://www.mongodb.com/docs/atlas/review-available-metrics.md 11. Atlas warns that an analytics tier "signifi — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/mechanism.md#b-replication-health-signals-the-core-analytics-node-risk`
- 17. RTPP shows data at one-second granularity. Its Replication Lag chart and Lag Time column are "available only for secondary members". https://www.mongodb.com/docs/atlas/architecture/current/monitoring-alerts.md ; https://www.mongodb.com/docs/atlas/real-time-performance-panel.md 18. RTPP requires M10+, is enabled by default, and caps each `db.currentOp()` sample at 4 MB. Its Kill Op button runs `db.killOp()` on the selected operation. https://www.mongodb.com/docs/atlas/real-time-performance-panel.md 19. The Query Profiler can filter "to one or more specific hosts or to view primary or second — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/mechanism.md#c-real-time-and-slow-operation-tooling`
- After any topology change, re-check hostname-scoped analytics alerts against `rs.conf()` tags. 29. The replication-lag alert fires when the secondary's lag "meets the specified threshold". Atlas computes lag with the method in the manual's "Check the Replication Lag" section. https://www.mongodb.com/docs/atlas/reference/alert-conditions.md 30. The Replication Headroom alert compares the sync source's oplog window with the secondary's lag. The Oplog Window alert measures the **primary's** oplog. https://www.mongodb.com/docs/atlas/reference/alert-conditions.md 31. A new project's default alerts — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/mechanism.md#d-alerting-mechanism-and-its-limits`
- 41. Prometheus integration requires M10+. It discovers targets through `/prometheus/v1.0/groups/<group_id>/discovery` and scrapes each node as its own `instance`. Labels include `cl_name`, `rs_nm`, `process_type` and `replica_state_name`. The page documents no node-type or analytics label. https://www.mongodb.com/docs/atlas/tutorial/prometheus-integration.md 42. **[inference]** Neither `replica_state_name` nor `process_type` can tell an analytics secondary from an operational one, because analytics nodes report as SECONDARY. To isolate analytics series in Prometheus, join on the host name, wit — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/mechanism.md#f-external-export`
- In scope: how to observe MongoDB Atlas analytics nodes. This covers metrics, replication lag and oplog health, alerting, slow-operation tooling, auto-scaling signals, and third-party export. — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/practice.md#scope`
- 1. Atlas warns that if the analytics tier is set "significantly below" the base tier, replication lag may result, and the analytics node "might fall off the oplog altogether." — https://www.mongodb.com/docs/atlas/cluster-config/multi-cloud-distribution/ 2. A node that has fallen off the oplog reports `STARTUP2` / `RECOVERING` for a long time. It logs `We are too stale to use <node>:27017 as a sync source.` It needs an initial sync to recover. — https://www.mongodb.com/docs/atlas/reference/alert-resolutions/replication-oplog/ 3. On the Metrics page, a dark-yellow `R` icon marks a member that ha — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/practice.md#a-replication-lag-is-the-primary-health-signal`
- 19. Query Profiler shows operations across the whole cluster by default. A "Filter by Hosts" dropdown narrows the view to specific hosts or to primary/secondary groups. — https://www.mongodb.com/docs/atlas/tutorial/query-profiler/ 20. By default Atlas sets the slow-op threshold per `mongod` from that host's average execution time. You can opt out to a fixed 100 ms through the Admin API. Profiler settings reset after a node restart. — https://www.mongodb.com/docs/atlas/tutorial/query-profiler/ 21. **Implication:** an analytics node runs long aggregations. Its dynamic threshold therefore sits hi — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/practice.md#c-slow-operations-and-index-advice`
- 25. Analytics nodes scale up one tier when average Normalized System CPU or System Memory Utilization exceeds 75% for the past hour. The Queued/Rejected Operations criterion does not apply to them. — https://www.mongodb.com/docs/atlas/cluster-autoscaling-compute-core/ 26. Analytics nodes scale down when average CPU and memory over the past 24 h are below 50%. — https://www.mongodb.com/docs/atlas/cluster-autoscaling-compute-core/ 27. Predictive auto-scaling does not apply to analytics or search nodes. — https://www.mongodb.com/docs/atlas/cluster-autoscaling-compute-core/ 28. Analytics-tier auto — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/practice.md#d-auto-scaling-signals-analytics-tier`

## Comparisons and alternatives

- 35. Analytics nodes scale up when average Normalized System CPU or System Memory Utilization stays above 75% for one hour. The Queued-or-Rejected-Operations (IWM) criterion does not apply to analytics nodes. https://www.mongodb.com/docs/atlas/cluster-autoscaling-compute-core.md 36. Analytics nodes scale down only when average CPU and memory stay below 50% over the past 24 hours. Operational nodes use 10-minute and 4-hour windows. https://www.mongodb.com/docs/atlas/cluster-autoscaling-compute-core.md 37. Predictive auto-scaling does not apply to analytics nodes. https://www.mongodb.com/docs/atl — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/mechanism.md#e-auto-scaling-as-a-monitoring-consumer`
- - **Why analytics nodes lag (two explanations).** The pmclain.com post says "the secondary node has higher priority resource access than the analytics nodes", which causes analytics lag under heavy writes. MongoDB docs blame only tier under-sizing ("significantly below the base tier"). No MongoDB doc describes resource prioritization between node types. — https://pmclain.com/mongodb/2022/09/28/fun-with-mongodb-atlas-analytics-nodes.html vs https://www.mongodb.com/docs/atlas/cluster-config/multi-cloud-distribution/ - **Default lag-alert threshold.** OneUptime (2026-03-31) says the Atlas Replica — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/edge-cases.md#unresolved-disagreements-and-gaps`
- 1. **Why analytics nodes lag.** - pmclain.com says Atlas gives the operational secondary higher-priority resource access. - Official docs attribute lag to undersizing (an analytics tier far below the base tier) and to the generic causes in the manual. - No official source confirms a resource-priority mechanism. Both views stand. 2. **Strict isolation or availability on the read path.** - pmclain.com moved to strict `readPreference=secondary` to protect the primary. - The Atlas docs recommend `secondary` plus an empty-tag fallback to any secondary. That keeps analytics reads available, but duri — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/history.md#unresolved-disagreements`
- - **Default replication-lag alert.** OneUptime (2026-03-31) says "Atlas alerts fire when lag exceeds a configurable threshold (default: 60 seconds)" and recommends 30–60 s thresholds: https://oneuptime.com/blog/post/2026-03-31-mongodb-respond-to-replication-lag-alerts/view. Atlas's own default-alert list has no lag alert: https://www.mongodb.com/docs/atlas/configure-alerts.md. The Architecture Center recommends > 240 s and > 1 h: https://www.mongodb.com/docs/atlas/architecture/current/monitoring-alerts.md. The vendor docs are the primary source, and on them the 60 s "default" does not hold. Th — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/mechanism.md#unresolved-disagreements`
- **What I re-checked against the live MongoDB docs (2026-10-02):** - **Datadog:** the mechanism report said the Datadog docs don't identify analytics nodes; the edge-cases report said they do. The Atlas Datadog page confirms a `nodeType` tag with values `ELECTABLE`, `READ_ONLY` and `ANALYTICS`. But Atlas puts its tags only "on certain metrics", so the page doesn't confirm the tag is on the lag metric. - **Default lag alert:** a new project's default alerts include no replication-lag alert, only `Replication Oplog Window is below 1 hours`. That refutes OneUptime's "60 s default". Which threshold — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/rabbithole-synthesis.md`
- - **Resource priority vs. tier sizing.** pmclain.com says "the secondary node has higher priority resource access than the analytics nodes" (https://pmclain.com/mongodb/2022/09/28/fun-with-mongodb-atlas-analytics-nodes.html). Atlas docs attribute analytics lag to a smaller analytics tier (https://www.mongodb.com/docs/atlas/cluster-config/multi-cloud-distribution/). The docs give no scheduling-priority mechanism. Both are kept; the blog claim is unverified. - **Real-Time Performance Panel coverage.** A summary of https://www.mongodb.com/docs/atlas/real-time-performance-panel/ returned both "cli — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/practice.md#unresolved-disagreements-and-gaps`

## Facts and statements

- Out of scope: analytics-node configuration and sizing, read-preference routing, BI Connector and Atlas SQL, and general Atlas monitoring that has no analytics-node delta. Parent facts are not repeated. The parent says analytics nodes are "separately monitored"; this report adds the specifics. — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/practice.md#scope`
- 31. Atlas pushes metrics to Datadog on M10+. The replication metrics are `mongodb.atlas.replset.replicationlag` (seconds) and `mongodb.atlas.replset.oplograte`. — https://docs.datadoghq.com/integrations/mongodb_atlas/ 32. Atlas Prometheus targets carry the labels `cl_name`, `rs_nm`, `rs_state`, `cl_role`, `process_port`, and `group_id`, on M10+. No node-type label is documented. — https://www.mongodb.com/docs/atlas/tutorial/prometheus-integration/ 33. Atlas applies the immutable replica-set tag `nodeType:ANALYTICS`. It applies `workloadType:OPERATIONAL` to non-analytics nodes. These are server — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/practice.md#e-export-and-integrations`
- **Source gate:** met for the general replication mechanics, which Datadog, Percona and OneUptime confirm independently. Not met for anything specific to analytics nodes: every such claim comes from MongoDB alone. I left pmclain.com (inherited) and Owler (an aggregator) out of the count. — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/rabbithole-synthesis.md`
- **In scope:** how you observe, alert on, and auto-scale MongoDB Atlas analytics nodes. That covers replication lag, oplog headroom, alert targeting, metric granularity, auto-scaling signals, third-party metric export, and the blind spots in each. — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/edge-cases.md#scope`
- **Out of scope:** analytics-node sizing, read-preference design, BI Connector, Atlas SQL, Data Federation, Search Nodes, and general replica-set monitoring. Those are sibling or parent concepts. — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/edge-cases.md#scope`
- **Not repeated from the parent:** analytics nodes are M10+, `priority:0`/`votes:0`, read-only and tagged `nodeType:ANALYTICS`. Claims marked **[inference]** combine two or more cited facts. They are not stated directly by any one source. — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/edge-cases.md#scope`
- 25. If the analytics tier is "significantly below" the base tier, "replication lag might result. The analytics node might fall off the oplog altogether." — https://www.mongodb.com/docs/atlas/cluster-config/multi-cloud-distribution/ 26. Replication Headroom is "the difference between the sync source member's oplog window and the replication lag time on the secondary". It is the alert that directly predicts falling off. The Oplog Window alert measures only the primary. — https://www.mongodb.com/docs/atlas/reference/alert-conditions/ 27. Signs a node has fallen off: the log line `We are too stale — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/edge-cases.md#e-falling-off-the-oplog`
- **Out of scope:** what analytics nodes are, sizing, read-preference routing, BI Connector, SQL Interface and Data Federation. Those belong to the parent and sibling concepts. Facts inherited from the parent are not repeated here, except where this report corrects or limits them. — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/history.md#scope`
- 8. The relevant alert conditions are `Replication Lag is`, `Replication Oplog Window is`, `Replication Headroom is` and `Oplog Data Per Hour is`. Conditions can target all hosts or a host type "such as primaries or config servers." The docs list no analytics host type as a matcher. https://www.mongodb.com/docs/atlas/reference/alert-conditions/ 9. Four alert or event conditions are specific to analytics nodes, and all cover analytics-tier auto-scaling: - compute auto-scaling initiated for analytics tier - scale-down blocked by storage requirements - scale-up blocked by the maximum configured ti — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/history.md#b-alerts`
- Run: /rabbithole depth pass, 2026-10-02. Parent: MongoDB Atlas Analytics Node. — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/mechanism.md`
- **Out of scope:** analytics node sizing and configuration, read-preference routing design, BI Connector, Atlas SQL, Data Federation, and Search Nodes. Those are sibling frontier items. Parent facts are not repeated here. They include priority 0, votes 0, M10+, read-only, the `nodeType:ANALYTICS` tag, and the `secondaryPreferred` fallback incident. — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/mechanism.md#scope`
- 26. Host-alert matchers accept `TYPE_NAME`, `HOSTNAME`, `PORT`, `HOSTNAME_AND_PORT` and `REPLICA_SET_NAME`. `TYPE_NAME` accepts only `PRIMARY`, `SECONDARY`, `STANDALONE`, `CONFIG` and `MONGOS`. It has no analytics value. https://raw.githubusercontent.com/mongodb/terraform-provider-mongodbatlas/master/docs/resources/alert_configuration.md 27. Atlas does not guarantee that a host name keeps pointing to an analytics node after a topology change such as scaling node count or regions. https://www.mongodb.com/docs/atlas/cluster-config/multi-cloud-distribution.md 28. **[inference]** Claims 26 and 27 — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/mechanism.md#d-alerting-mechanism-and-its-limits`
- - **Distinct hosts used:** mongodb.com (docs and company blog), pmclain.com, registry.terraform.io, raw.githubusercontent.com and owler.com. - **Independent authors:** only two, MongoDB Inc. and pmclain.com. The Terraform provider and its changelog are MongoDB-authored, and Owler only aggregates. - **Result: partly met.** There are at least 3 hosts, but only 2 independent organizations. - **Disconfirming source:** pmclain.com, which contradicts the assumption that tagged analytics reads are fully isolated and attributes lag differently from the docs. — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/analytics-node-monitoring-b7c48f1ec6/reports/history.md#quality-gate`

## Related concepts

- Analytics — is a part of Analytics Node Monitoring
- Node — is a part of Analytics Node Monitoring
- Monitoring — is a part of Analytics Node Monitoring
