<!-- llms-explorer concept facts · https://llms-explorer.com/tree/ako-troubleshooting/ · pack 2026-10-02 · ~9990 tokens -->

# AKO Troubleshooting

> Depth-first rabbithole dossier for AKO Troubleshooting; source-anchored research pack.

Parent: [Atlas Kubernetes Operator](https://llms-explorer.com/tree/atlas-kubernetes-operator/) · 5 facets · 59 facts · page: https://llms-explorer.com/tree/ako-troubleshooting/

## How it works

- - **The runbook does not match the defect history.** The runbook (claim 22) leads with configuration causes: a mislabelled credentials secret, then logs. The changelog and issues (claims 29–34) show that many "stuck" or infinite-reconcile cases were AKO bugs: diffs on read-only or defaulted fields, and autoscaling comparisons. Checking the secret label does not find those. The third step, reading logs, plus upgrading AKO is the only route the record supports. The sources do not resolve this; both sides are kept. - **The CVE was published long after the fix.** CVE-2023-0436 was published on 202 — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/history.md#unresolved-disagreements-and-gaps`
- 28. Since 2.0, deleting a custom resource leaves the Atlas object in place by default; AKO only stops managing it. The global switch is `--object-deletion-protection` / `OBJECT_DELETION_PROTECTION` (default `true`). The per-resource override is the annotation `mongodb.com/atlas-resource-policy: "delete"` or `"keep"`. https://www.mongodb.com/docs/atlas/operator/current/custom-resources/ 29. AKO 2.1.0 disabled `--subobject-deletion-protection`, because with it enabled users could not modify existing resources. https://www.mongodb.com/docs/atlas/operator/current/ak8so-changelog/ 30. AKO 2.0.1 cou — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/edge-cases.md#e-deletion-finalizers-and-destructive-side-effects`
- 36. `mongodb.com/atlas-reconciliation-policy: "skip"` stops AKO from starting reconciliation for that resource until you remove the annotation. MongoDB documents it as a way to make manual changes without AKO undoing them. https://www.mongodb.com/docs/atlas/operator/current/custom-resources/ 37. The skip annotation pauses sync but does not change resource state. Reconciliation resumes when you remove the annotation, so any manual Atlas changes made meanwhile get reconciled back to the spec. https://github.com/mongodb/mongodb-atlas-kubernetes/blob/v1.9.3/docs/annotations.md 38. `mongodb.com/atl — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/edge-cases.md#f-skip-annotation-and-version-mismatch-annotation`
- Out of scope: AKO installation, Helm, GitOps, CRD lifecycle, IaC comparison, and the Enterprise/Community/MCK operators. Those are sibling frontier items. Claims inherited from the parent brief are not repeated here; for example, the parent already says AKO uses custom condition types, so that point is not restated as a new finding. — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/history.md#scope`
- 26. AKO exposes the standard controller-runtime metrics on `:8080/metrics` (the `--metrics-bind-address` flag changes the port). They include reconcile totals and errors per controller, queue length, reconcile latency, and process and Go runtime metrics. — https://www.mongodb.com/docs/atlas/operator/current/ak8so-metrics/ ; https://www.mongodb.com/docs/atlas/operator/current/production-notes/ 27. The documented stuck-resource alert is `100 * rate(controller_runtime_reconcile_errors_total{controller="AtlasProject"}[1m]) / rate(controller_runtime_reconcile_total{controller="AtlasProject"}[1m])`. — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/mechanism.md#f-metrics`

## Measurements and reference values

- I merged the four AKO Troubleshooting reports (mechanism, history, edge cases, practice) into one set of 101 claims. Each claim is a single fact with its source URL and the report claims it came from. **Verdict: BUDGET_EXHAUSTED (a soft stop), not SATURATED-DEPTH.** All four reports stopped for the same reason. Their last-pass new-information rates were 15%, 6%, about 16% and 26%. None had two passes in a row under 5%. — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/rabbithole-synthesis.md`
- **Verdict: BUDGET_EXHAUSTED (soft stop), not SATURATED-DEPTH.** The curve is falling but still above 5%. One or two more passes would likely pay off. Targets: AKO's own retry and backoff on 429, controller log message catalogue, `observedGeneration` semantics per CR, and GitHub issues from after 2024. - **Tool limits:** Firecrawl search and scrape were not available in this run. Research used WebSearch and WebFetch. The shared Helm-charts cache page held no troubleshooting content. — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/practice.md#quality-gate-and-saturation`
- Verdict: **BUDGET_EXHAUSTED.** This was a history-scoped brief, not depth saturation. The rate fell to 6% but never had two passes under 5%. One or two more passes look likely to pay off: GitHub release notes for v2.9–v2.17, and how AKO handles Atlas API 429 errors in source. — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/history.md#depth-passes-new-information-rate`
- 32. Infinite reconcile loops from spec/Atlas normalisation mismatches recur. Examples: autoscaling `min/maxInstanceSize` set while compute autoscaling is disabled (fixed in 2.8.1 and 2.8.2), `terminationProtectionEnabled` set in the UI, and `ebsVolumeType`. — https://www.mongodb.com/docs/atlas/operator/current/ak8so-changelog/ ; https://github.com/mongodb/mongodb-atlas-kubernetes/releases 33. Some resources got permanently stuck because of concurrency. Simultaneous FlexCluster and Group creation caused `DUPLICATE_CLUSTER_NAME` (fixed in v2.15.0, #3367). A 2.2.2 concurrency bug made AKO miss CR — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/mechanism.md#h-failure-classes-in-the-record`
- **Verdict: BUDGET_EXHAUSTED (soft stop), not SATURATED-DEPTH.** The rate fell 100 → 57 → 40 → 15%, but it never reached two passes under 5%. Bash and Firecrawl were denied in this session, so the passes used WebSearch/WebFetch only, and the github.com 404s cut off direct reading of source files. One or two more passes would probably still pay off. They would read the controller source on requeue and 429 handling, and the per-reason remedies in the skill's `references/troubleshooting.md`. — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/mechanism.md#pass-log-new-information-rate`

## Problems, failure modes and limitations

- - **Independent hosts:** 5 — mongodb.com, raw.githubusercontent.com and github.com (MongoDB-owned repo plus user-filed issues), book.kubebuilder.io, kubernetes.io. The gate of 3 independent sources is met by host. **Caveat:** nearly all AKO-specific claims trace to MongoDB as author, through its docs, repo and changelog. The independent sources (kubebuilder, kubernetes.io, user issues) corroborate mechanism and symptoms, not AKO's own guidance. No third-party field report on AKO troubleshooting was found. - **Disconfirming source sought:** kubernetes.io on `kubectl wait`, which partly contradi — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/practice.md#quality-gate-and-saturation`
- 1. AKO v0.8.0 was published on 2022-03-04. https://api.github.com/repos/mongodb/mongodb-atlas-kubernetes/releases/tags/v0.8.0 2. v0.8.0 added the `mongodb.com/atlas-reconciliation-policy=skip` annotation, which makes AKO skip reconciliation for a specific resource. https://www.mongodb.com/docs/atlas/operator/current/ak8so-changelog/ 3. From v0.8.0, AKO does not mark an `AtlasProject` Ready until the project IP access list exists in Atlas. https://api.github.com/repos/mongodb/mongodb-atlas-kubernetes/releases/tags/v0.8.0 4. While the skip annotation is present, AKO stops syncing that resource, — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/history.md#timeline-of-the-troubleshooting-surface`
- 45. If Atlas rate-limits a request, it returns HTTP 429 with `errorCode: RATE_LIMITED_TOKEN_BUCKET`. It may send `Retry-After`, `RateLimit-Limit`, and `RateLimit-Remaining`, but these headers "may not always be present." https://www.mongodb.com/docs/atlas/api/api-rate-limit/ 46. Limits apply per scope (GROUP, ORGANIZATION, USER, IP) and per endpoint set. The tightest sets are tiny: Organization Settings has capacity 10 and refills 5 per 60s. Clusters are generous: capacity 10,000, refill 5000 per 60s per project. https://www.mongodb.com/docs/atlas/api/api-rate-limit/ 47. MongoDB says rate limi — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/edge-cases.md#h-atlas-api-rate-limits-the-atlas-side-of-ako-errors`
- 1. If reconciliation hits an error, AKO writes the error into `status.conditions` on the custom resource. https://www.mongodb.com/docs/atlas/operator/current/production-notes/ 2. A failed sub-step shows up as its own condition type, with the raw Atlas API error in `message`. The documented example has `type: IPAccessListReady`, `reason: ProjectIPAccessListNotCreatedInAtlas`, and the message `POST …/accessList: 400 (request "INVALID_IP_ADDRESS_OR_CIDR_NOTATION")`. https://www.mongodb.com/docs/atlas/operator/current/custom-resources/ 3. One reconcile of one custom resource can make several Atlas — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/edge-cases.md#a-status-conditions-as-the-diagnostic-surface`
- 10. MongoDB's runbook lists two symptoms: the resource is not `Ready`, and the controller has a high error rate. It says this "can occur with every Atlas Kubernetes Operator resource type." https://www.mongodb.com/docs/atlas/operator/v2.14/ak8so-metrics/ 11. The documented error-rate query is `100 * rate(controller_runtime_reconcile_errors_total{controller="AtlasProject"}[1m]) / rate(controller_runtime_reconcile_total{controller="AtlasProject"}[1m])`. Metrics are served on `http://localhost:8080/metrics`. https://www.mongodb.com/docs/atlas/operator/v2.14/ak8so-metrics/ 12. The runbook has thre — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/edge-cases.md#b-official-stuck-in-reconciliation-runbook`
- 25. In AKO 1.1.0, after Atlas autoscaled a cluster from M10 to M40, the operator's PATCH failed with `400 (request 'ATTRIBUTE_READ_ONLY') The attribute createDate is read-only`. The operator had echoed read-only fields back to Atlas. The bug was fixed by PR #615. https://github.com/mongodb/mongodb-atlas-kubernetes/issues/606 26. In AKO 1.4.1, raising `minInstanceSize` above the current base size produced `BASE_INSTANCE_SIZE_MUST_MATCH` without `readOnlySpecs`. With `readOnlySpecs`, it produced a 500 `UNEXPECTED_ERROR` after the log warning "The instance size is below the minimum autoscaling co — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/edge-cases.md#d-autoscaling-drift-errors-atlas-rejects-the-operator-s-patch`
- 39. Dry run is **public preview**, and MongoDB says it "might change at any time." https://www.mongodb.com/docs/atlas/operator/current/ak8so-dry-run/ 40. Dry run emits events only for mutating verbs (`POST`, `PATCH`, `PUT`, `DELETE`). It cannot show read-path failures such as a credential or GET error that never reaches a write. https://www.mongodb.com/docs/atlas/operator/current/ak8so-dry-run/ 41. Dry run events can be filtered with `kubectl -n mongodb-atlas-system get events --field-selector reason=DryRun`. A run ends with `Done` and then `Finished`. https://www.mongodb.com/docs/atlas/operat — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/edge-cases.md#g-dry-run-as-a-diagnostic`
- 48. The upgrade to 2.0.1 removed `advancedDeploymentSpec` in favor of `deploymentSpec` but kept `apiVersion: atlas.mongodb.com/v1`. The reporter called it a silent breaking change with no automated migration. The issue was closed "not planned" with no maintainer reply. https://github.com/mongodb/mongodb-atlas-kubernetes/issues/1330 49. 2.15.0 fixed duplicate `status.projects` entries left over after upgrading. Stale status after an upgrade is a known artifact, not necessarily a live fault. https://github.com/mongodb/mongodb-atlas-kubernetes/releases 50. `AtlasProject` subresources (`spec.proje — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/edge-cases.md#i-upgrade-induced-failures`
- 17. AKO serves the standard controller-runtime metrics at `http://localhost:8080/metrics`. https://www.mongodb.com/docs/atlas/operator/current/ak8so-metrics/ 18. The metrics cover reconcile errors and successes per controller, queue length per controller, reconcile latency, process resources, and Go runtime stats. https://www.mongodb.com/docs/atlas/operator/current/ak8so-metrics/ 19. `controller_runtime_reconcile_total` counts reconciliations per controller. `controller_runtime_reconcile_errors_total` counts reconcile errors per controller. Both are counters. https://book.kubebuilder.io/refere — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/history.md#current-mechanism-what-an-operator-checks`
- **In scope:** how AKO tells you a reconcile has failed, which signals you can read (status conditions, reasons, messages, events, logs, metrics, dry-run output), how those signals are produced inside the operator, the failure classes they report, and where the signals stop being useful. — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/mechanism.md#scope`
- 1. Every AKO custom resource status embeds a shared `Common` struct. It holds `Conditions` ("the list of statuses showing the current state of the Atlas Custom Resource") and `ObservedGeneration`. — https://pkg.go.dev/github.com/mongodb/mongodb-atlas-kubernetes/v2/api 2. `ObservedGeneration` "indicates the generation of the resource specification of which the Atlas Operator is aware". If it is lower than `metadata.generation`, AKO has not yet processed the latest spec, so the conditions describe an older spec. — https://pkg.go.dev/github.com/mongodb/mongodb-atlas-kubernetes/v2/api 3. A `Condit — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/mechanism.md#a-status-conditions-the-primary-signal`
- 22. Two flags control logging: `--log-level` (`debug|info|warn|error|dpanic|panic|fatal`, default `info`) and `--log-encoder` (`json|console`, default `json`). — https://www.mongodb.com/docs/atlas/operator/current/production-notes/ 23. With `--log-level=debug`, AKO logs a field-level diff between the current state and the desired state. This shows which field drives a repeated PATCH. Debug mode also adds GET requests to Atlas, so the docs say to use it only outside production. — https://www.mongodb.com/docs/atlas/operator/current/ak8so-dry-run/ 24. Debug logs have leaked secrets in the past. C — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/mechanism.md#e-logs`
- 37. The Atlas Admin API rate-limits with a token bucket per endpoint set, scoped to GROUP (project), ORGANIZATION, USER or IP. A rejected request returns HTTP 429 with `errorCode: RATE_LIMITED_TOKEN_BUCKET` and a `detail` string that names the capacity and refill rate. — https://www.mongodb.com/docs/atlas/api/api-rate-limit/ 38. Atlas may return `RateLimit-Limit`, `RateLimit-Remaining` and `Retry-After` headers. It recommends honouring `Retry-After`, then exponential backoff with jitter. Rate limits are "not part of the API contract and may change without notice". — https://www.mongodb.com/doc — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/mechanism.md#i-rate-limiting-atlas-side-mechanism`
- | Pass | Focus | New claims | Total | Rate | |---|---|---|---|---| | 0 | official troubleshooting/metrics + quick start | 9 | 9 | 100% | | 1 | source: condition types, reasons, Result semantics | 12 | 21 | 57% | | 2 | production notes, dry run, changelog failure classes | 14 | 35 | 40% | | 3 | Atlas rate limits, kubectl wait (disconfirming), issues | 6 | 41 | 15% | — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/mechanism.md#pass-log-new-information-rate`
- - Reconciliation-skip annotation `mongodb.com/atlas-reconciliation-policy=skip` (v2.6.1+): separate frontier item. - Deletion-protection flags (`--object-deletion-protection`, the disabled `--subobject-deletion-protection`). - Resource version label / `ResourceVersionIsValid` and upgrade paths: belongs to AKO upgrade/install. - Atlas Admin API error-code catalogue: https://www.mongodb.com/docs/atlas/reference/api-errors/ — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/mechanism.md#handoffs-adjacent-concepts-not-chased`
- 1. Every AKO custom resource carries a `Ready` condition. On success it is `status: "True"`. On failure the condition is `False`, and resource-specific conditions hold the error detail. https://www.mongodb.com/docs/atlas/operator/current/custom-resources/ 2. Failed conditions embed the raw Atlas API call and error code in `message`, and give a typed `reason`. Example: `reason: ProjectIPAccessListNotCreatedInAtlas` on condition type `IPAccessListReady`, with message `POST …/accessList: 400 (request "INVALID_IP_ADDRESS_OR_CIDR_NOTATION")`. Most Atlas-side failures can therefore be diagnosed from — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/practice.md#primary-signal-status-conditions`
- 7. MongoDB's SRE runbook defines the problem as a resource that is not `Ready` and has a high reconcile error rate. It applies to every AKO resource type, not only `AtlasProject`. https://raw.githubusercontent.com/mongodb/mongodb-atlas-kubernetes/main/docs/sre-runbook/resource_stuck_in_reconciliation.md 8. The runbook's error-rate query is `100 * rate(controller_runtime_reconcile_errors_total{controller="AtlasProject"}[1m]) / rate(controller_runtime_reconcile_total{controller="AtlasProject"}[1m])`. https://raw.githubusercontent.com/mongodb/mongodb-atlas-kubernetes/main/docs/sre-runbook/resourc — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/practice.md#official-runbook-resource-stuck-in-reconciliation`
- 12. The Helm chart raises verbosity through `extraArgs`, for example `--log-level=debug`. https://raw.githubusercontent.com/mongodb/helm-charts/main/charts/atlas-operator/values.yaml 13. The manager binary also accepts `--log-encoder=json`, as shown in the dry-run Job spec. https://www.mongodb.com/docs/atlas/operator/current/ak8so-dry-run/ 14. Debug logging has a security cost. CVE-2023-0436 (fixed in 1.7.1) let DEBUG-mode AKO print secrets such as GCP service-account keys and API integration secrets. DEBUG is off by default. Treat debug logs from older versions as sensitive. https://www.mongo — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/practice.md#logs-and-log-level`
- 16. AKO records Kubernetes Events per resource. If its ServiceAccount lacks `create` on `events` in the target namespace, those events fail silently while the reconcile itself succeeds. Exact error: `events is forbidden: User "system:serviceaccount:atlas-operator:mongodb-atlas-operator" cannot create resource "events" … in the namespace "operator-sandbox"` (issue #348, v0.6.1, 2021-11-25). If `kubectl describe` shows no events, check RBAC before assuming there was no activity. https://github.com/mongodb/mongodb-atlas-kubernetes/issues/348 — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/practice.md#events`
- 17. Dry run (public preview since 2.8.0) runs the manager with `--dry-run`, as a Job or through `atlas kubernetes dry-run --targetNamespace=… --watch`. It emits `Would <verb> (<METHOD>) <Atlas URL>` events, and only for mutating verbs (`POST`, `PATCH`, `PUT`, `DELETE`). https://www.mongodb.com/docs/atlas/operator/current/ak8so-dry-run/ 18. To read dry-run output, run `kubectl -n mongodb-atlas-system get events --field-selector reason=DryRun`. The run ends with `Done`, then `Finished`. https://www.mongodb.com/docs/atlas/operator/current/ak8so-dry-run/ 19. Dry run with `--log-level=debug` logs f — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/practice.md#dry-run-as-a-diagnostic`
- 27. When an Atlas Admin API rate limit is hit, the API returns HTTP 429 with `errorCode: "RATE_LIMITED_TOKEN_BUCKET"` and may return `Retry-After`. Limits are token buckets scoped per GROUP, ORGANIZATION, USER or IP. https://www.mongodb.com/docs/atlas/api/api-rate-limit/ 28. MongoDB states that rate limits "are not part of the API contract and may change without prior notice". `GET /api/atlas/v2/rateLimits?groupId=…` reports the current limits. https://www.mongodb.com/docs/atlas/api/api-rate-limit/ 29. Implication (inference from 2, 10 and 27): one AKO instance that reconciles many resources i — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/practice.md#atlas-api-rate-limiting-seen-through-ako`
- 30. Since AKO 2.0, deleting a CR does not delete the Atlas object. AKO stops managing it instead. A "the cluster still exists in Atlas" report is therefore expected behavior, not a fault. https://www.mongodb.com/docs/atlas/operator/current/atlasdeployment-custom-resource/ 31. The chart switches for this behavior are `objectDeletionProtection: true` and `subobjectDeletionProtection: true`. The second stops AKO from overwriting sub-resources it did not create. That can explain a resource that is `Ready` while Atlas still holds out-of-band settings. https://raw.githubusercontent.com/mongodb/helm- — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/practice.md#deletion-and-missing-resources`
- 33. Diagnostic routes, cheapest first: - condition `message`: zero cost, includes the Atlas error code (2) - events: can be silently missing because of RBAC (16) - controller metrics: show fleet-wide rates but no root cause (8, 11) - debug logs: show the diff, but carry a secret-exposure risk on old versions (14, 15) - dry run: makes no change, but is preview, needs its own cluster for upgrades, and has had coverage gaps (17–21) Sources: https://www.mongodb.com/docs/atlas/operator/current/custom-resources/ and https://www.mongodb.com/docs/atlas/operator/current/ak8so-dry-run/ 34. Version matte — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/practice.md#operational-trade-offs-evaluation`

## Comparisons and alternatives

- - **D1 — `kubectl wait` vs polling `status.conditions`.** The parent skill says to "poll status.conditions instead of kubectl wait." That advice conflicts with kubectl's documented behavior: `--for=condition=Ready` matches any condition type, and AKO does set a `Ready` type (claims 4 and 8). The real risks are the 30s default timeout and the missing per-condition `observedGeneration` (claims 5 and 9). It is not a custom-type incompatibility. Neither side is authoritative on the stale-Ready race, because MongoDB documents no guidance on it. https://kubernetes.io/docs/reference/kubectl/generated — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/edge-cases.md#unresolved-disagreements-and-gaps`
- **Contradictions I kept side by side (12 in the file). The main ones:** - **`kubectl wait`:** the parent says to poll `status.conditions` instead of using `kubectl wait`. Three reports independently found that `kubectl wait` does work on AKO's generic `Ready` condition. The real risks are its 30-second default timeout and a stale `Ready=True` left over from the previous spec. - **Deployment readiness:** the docs show a deployment with `Ready=True` while `ClusterReady=False`, which would make waiting on `Ready` alone unsafe. Separately, the docs call the condition `ClusterReady`, but the source — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/rabbithole-synthesis.md`
- 15. In versions before 2.8.2, AKO reconciled `AtlasDeployment` forever in one case: `autoscaling.compute.enabled: false` while `minInstanceSize`/`maxInstanceSize` were still set. A region-config comparison never matched. https://www.mongodb.com/docs/atlas/operator/current/ak8so-changelog/ 16. In versions before 2.8.1, a similar infinite loop occurred when compute auto-scaling was enabled but `minInstanceSize` was not cleared. https://www.mongodb.com/docs/atlas/operator/current/ak8so-changelog/ 17. In versions before 2.14.0, setting `terminationProtectionEnabled: true` **in the Atlas UI** could — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/edge-cases.md#c-infinite-or-stuck-reconcile-loops-version-bound-bugs`
- 29. In issue #606, opened on 2022-07-15 against v1.1.0, `DeploymentReady=False` showed reason `DeploymentNotUpdatedInAtlas`. The Atlas error was `400 (request "ATTRIBUTE_READ_ONLY") The attribute createDate is read-only`. The cause was AKO sending a read-only field in its PATCH. PR #615 fixed it. https://github.com/mongodb/mongodb-atlas-kubernetes/issues/606 30. v2.2.1 fixed a reconcile loop that kept recreating serverless private endpoints when they failed to sync with Atlas. https://www.mongodb.com/docs/atlas/operator/current/ak8so-changelog/ 31. v2.4.1 fixed a bug where AKO sometimes skippe — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/history.md#history-of-stuck-infinite-reconcile-defects-failure-modes-outside-the-runbook`
- **Out of scope:** installing AKO, Helm, independent CRDs, GitOps, how the reconciliation-skip annotation works, AKO vs MCK vs Terraform. Each one is a separate frontier item. This report names them only as handoffs. — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/mechanism.md#scope`
- - **kubectl wait vs polling.** The parent says "poll status.conditions instead of kubectl wait". The kubectl reference (claim 40) shows `--for=condition=Ready` and `--for=jsonpath` work on any condition type, and AKO always sets a generic `Ready` (claim 5). AKO docs poll (claim 41) but never say `kubectl wait` fails. Most likely reconciliation: `kubectl wait` works on `Ready`. It can still return early on a stale `Ready=True` left over from a previous generation (claim 2), so pair it with an `observedGeneration` check. Neither side is settled by a primary source, so both positions stand. - **D — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/mechanism.md#unresolved-disagreements-and-gaps`
- **Out of scope:** these are covered by separate frontier items: - AKO install or upgrade (Helm), independent CRDs, and GitOps - the reconciliation-skip annotation as a feature in its own right - the parent's AKO vs Terraform comparison — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/practice.md#scope`
- 22. Leaving fields to Atlas defaults can cause a reconcile loop that keeps the resource from reaching `READY`. MongoDB's mitigation is to state autoscaling explicitly (`compute.enabled`, `scaleDownEnabled`, `minInstanceSize`, `maxInstanceSize`). Otherwise AKO keeps reapplying a static `instanceSize` against Atlas autoscaling. https://www.mongodb.com/docs/atlas/operator/current/custom-resources/ 23. Known AtlasDeployment loop bugs. Version 2.8.1 fixed a loop when the `minInstanceSize` flag is enabled but the field is not cleared. Version 2.8.2 fixed a loop caused by wrong region-config comparis — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/practice.md#reconcile-loops-causes-and-known-bugs`
- - **`kubectl wait` vs polling `status.conditions`.** The parent skill says AKO's custom condition types mean you should poll `status.conditions` instead of using `kubectl wait`. Two sources cut against this: - The Kubernetes docs describe `kubectl wait --for=condition=<type>` and `--for=jsonpath='{.status.conditions[?(@.type=="Ready")].status}'=True` as reading `.status.conditions` generically. https://kubernetes.io/docs/reference/kubectl/generated/kubectl_wait/ - AKO does expose a standard `type: Ready`. https://www.mongodb.com/docs/atlas/operator/current/custom-resources/ — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/practice.md#unresolved-disagreements`

## Facts and statements

- - **Hosts used:** mongodb.com (official docs); github.com, together with api.github.com and raw.githubusercontent.com (official AKO repo, releases and issues); cveawg.mitre.org (CVE record); book.kubebuilder.io (controller-runtime metric definitions). - **Independence:** the gate is met on host count. Independence of content is weak. The docs and repo both come from MongoDB, and the CVE record restates MongoDB's advisory. Only kubebuilder is a truly independent source, and it supports only the metric definitions. - **Disconfirming source:** I looked for one and found it. Issue #606 and the inf — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/history.md#quality-gate`
- Run: `/rabbithole` depth pass, 2026-10-01. Concept: **AKO Troubleshooting** (parent: Atlas Kubernetes Operator). — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/mechanism.md`
- - The newer `Cluster` and `Group` custom resources (`/docs/atlas/operator/current/cluster-custom-resource/`, `/group-custom-resource/`) appear to be a new CRD family. They need their own troubleshooting coverage. - AKO metrics and observability as a standalone concept, and AKO upgrade or migration paths. Both are siblings, not deepened here. — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/edge-cases.md#handoffs-out-of-scope-for-concept-family-explorer`
- In scope: the troubleshooting surface that the MongoDB Atlas Kubernetes Operator (AKO) gives an operator. That covers status conditions, logs and log levels, the reconcile-skip annotation, dry run, controller metrics, the SRE runbook, and the history of "stuck" or "infinite reconcile" defects. — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/history.md#scope`
- **Inherited, not repeated:** the parent says AKO runs a controller-runtime reconcile loop, uses custom condition types, and has a `references/troubleshooting.md` covering stuck clusters, secret format, rate limiting, dry run and upgrades. The claims below add only child-specific detail, corrections and limits. — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/mechanism.md#scope`
- **Source gate:** met with three independent origins: MongoDB, the Kubernetes project, and the people who filed the GitHub issues. I didn't count inherited sources: the shared Helm-charts cache page has no troubleshooting content, and I treated the Helm chart `values.yaml` as the same origin. The independence is weak on content, though: all AKO-specific *guidance* comes from MongoDB alone. — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/rabbithole-synthesis.md`
- **Biggest open gaps:** whether AKO honours `Retry-After` or backs off on HTTP 429 (no report found a source), what requeue policy the newer controllers use, and AKO's finalizer key or a procedure for a resource stuck Terminating. The next pass should read AKO's controller source and the v2.9–v2.17 release notes. — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/rabbithole-synthesis.md`
- **Scope.** This report covers how to diagnose and recover from Atlas Kubernetes Operator (AKO) failures only: status conditions, stuck or looping reconciliation, credential secrets, deletion and finalizer behavior, the skip annotation, dry run, and API rate limits as AKO meets them. It does not cover AKO installation, CRD authoring, GitOps, or comparisons with other IaC tools; those are separate frontier items. Inherited parent facts are not repeated here. — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/edge-cases.md`
- **Quality gate.** The report uses 4 source hosts: mongodb.com, github.com, pkg.go.dev, and kubernetes.io. Only kubernetes.io is independent of MongoDB. pkg.go.dev renders MongoDB's own Go source, and github.com hosts MongoDB's repo and issues. The gate of 3 independent sources is therefore **only partly met**: the general Kubernetes claims have an independent source, but the AKO-specific claims rest on vendor material, including user-filed issues. I found no third-party write-up that contradicts AKO's docs. Two disconfirming findings come from comparing vendor sources with each other (D1, D2 b — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/edge-cases.md`
- Shared cache: the `mongodb.github.io/helm-charts` page lists charts only and has no troubleshooting content. This report does not use it as evidence. — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/history.md#scope`
- Handoffs to concept-family-explorer (not researched here): AKO upgrade and CRD-version testing; the Atlas CLI `atlas kubernetes` command group; Prometheus/Grafana dashboards for AKO (`docs/grafana/`, `docs/metrics-via-grafana.md`); deletion protection and the `atlas-resource-policy: keep` annotation. — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/history.md#depth-passes-new-information-rate`
- 19. AKO looks for credentials in this order. First `spec.connectionSecretRef.name` (by default in the AtlasProject's namespace, or in `spec.connectionSecretRef.namespace`). If that is not set, it uses the global secret `<operator-deployment-name>-api-key`, whose name the `--global-api-secret-name` flag can override. — https://www.mongodb.com/docs/atlas/operator/current/production-notes/ 20. The referenced secret must carry the label `atlas.mongodb.com/type=credentials`. The official runbook checks this label as step 2, after the status message. — https://www.mongodb.com/docs/atlas/operator/cur — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/mechanism.md#d-credentials-the-most-common-precondition-failure`
- 29. Dry run (public preview since v2.8.0) runs the `/manager` binary with `--dry-run`, usually as a Kubernetes Job. You can also start it with `atlas kubernetes dry-run --targetNamespace=... --watch`. — https://www.mongodb.com/docs/atlas/operator/current/ak8so-dry-run/ 30. Dry run emits Kubernetes events with `reason=DryRun`. It emits a `Would <verb> (<HTTP-METHOD>) <Atlas URL>` event only for mutating methods (POST, PATCH, PUT, DELETE), followed by `Done` and `Finished`. To read them: `kubectl -n <ns> get events --field-selector reason=DryRun`. — https://www.mongodb.com/docs/atlas/operator/cu — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/mechanism.md#g-dry-run-preview-before-you-change-anything`
- 40. `kubectl wait --for=condition=<name>[=<value>]` and `--for=jsonpath='{.status.conditions[?(@.type=="Ready")].status}'=True` are generic kubectl features. The condition named in `--for=condition` can be any condition type in `status.conditions`. — https://kubernetes.io/docs/reference/kubectl/generated/kubectl_wait/ 41. AKO's own docs check readiness by polling `kubectl get atlasdatabaseusers <name> -o=jsonpath='{.status.conditions[?(@.type=="Ready")].status}'` "until you receive a `True` response". — https://www.mongodb.com/docs/atlas/operator/current/ak8so-quick-start/ — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/mechanism.md#j-waiting-on-readiness-correction-to-the-inherited-claim`
- **In scope:** how an operator diagnoses Atlas Kubernetes Operator (AKO) resources that do not reach `Ready`. This covers status conditions, logs and log level, controller metrics, events, dry-run as a diagnostic, Atlas API errors seen through AKO (including rate limits), known reconcile-loop bugs, and the trade-offs of each diagnostic route. — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/practice.md#scope`
- The parent facts (custom condition types, controller-runtime loop, `references/troubleshooting.md` topics) are not repeated here. This report adds only child-specific deltas. — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/practice.md#scope`
- - AKO Helm installation and upgrade paths - AKO independent CRDs - AKO reconciliation-skip annotation - AKO deletion protection (as its own concept) - Atlas Admin API rate limiting (as its own concept) — source: `~/.global-ai-hub/research-tests/mongodb-full-frontier-20261002/full-frontier-run/ako-troubleshooting-c6bc60094d/reports/practice.md#handoffs-not-researched-separate-frontier-items`

## Related concepts

- AKO — is a part of AKO Troubleshooting
- Troubleshooting — is a part of AKO Troubleshooting
