<!-- llms-explorer concept facts · https://llms-explorer.com/tree/causal-discovery-and-structure-learning/ · pack 2026-09-08 · ~7385 tokens -->

# Causal Discovery and Structure Learning

> Causal discovery (a.k.a. structure learning) learns the causal graph itself

Parent: [Data Analysis](https://llms-explorer.com/tree/data-analysis/) · 18 facets · 104 facts · page: https://llms-explorer.com/tree/causal-discovery-and-structure-learning/

## Overview

- Causal discovery (a.k.a. structure learning) learns the causal graph itself from observational and/or interventional data - the edges and their directions - rather than assuming the graph and estimating an effect. This is the upstream problem to causal inference. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#overview)
  - This skill (discovery): "What is the causal structure? Which variables cause which?" → output is a graph (DAG, CPDAG, or PAG). — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#overview)
  - da-12 (inference): "Given this DAG, what is the effect of X on Y?" → DiD, RDD, IV, propensity scores, synthetic control, backdoor/frontdoor adjustment. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#overview)
- If the user already has/assumes a DAG and wants an effect estimate, defer to da-12-ab-testing-causal-inference. Use this skill only when the structure is the unknown. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#overview)
- The hard truth of discovery: from purely observational data you usually cannot recover a single DAG - only an equivalence class of DAGs (a CPDAG or PAG). Pinning down direction requires extra assumptions (non-Gaussianity, nonlinearity), interventions, or time order. Always communicate which edges are oriented vs. undetermined. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#overview)

## 1. Markov equivalence, CPDAGs, and what is identifiable

- Two DAGs are Markov equivalent if they entail the same conditional independences - they have the same skeleton (undirected edges) and the same v-structures / colliders (A → C ← B with A, B not adjacent). Equivalent DAGs cannot be distinguished by observational independence tests alone (Verma & Pearl, 1990; Andersson, Madigan & Perlman, 1997). — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#1-markov-equivalence-cpdags-and-what-is-identifiable)
  - A CPDAG (Completed Partially Directed Acyclic Graph, a.k.a. essential graph) represents the whole Markov equivalence class: directed edges are oriented in every member, undirected edges flip across members. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#1-markov-equivalence-cpdags-and-what-is-identifiable)
  - Constraint- and score-based methods return a CPDAG, not a DAG. Reporting a single oriented DAG from such output is a common, serious error. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#1-markov-equivalence-cpdags-and-what-is-identifiable)

## 2. Foundational assumptions (state them, always)

- Causal Markov condition: each variable is independent of its non-descendants given its parents. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#2-foundational-assumptions-state-them-always)
- Faithfulness: every conditional independence in the distribution is implied by the graph structure (no exact cancellations). Near-violations cause unstable orientation in finite samples (Spirtes, Glymour & Scheines, 2000). — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#2-foundational-assumptions-state-them-always)
- Causal sufficiency: no unmeasured common causes (latent confounders). PC and GES assume this; FCI does not. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#2-foundational-assumptions-state-them-always)
- Acyclicity: most methods assume a DAG (no feedback loops). — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#2-foundational-assumptions-state-them-always)
- Identifiability hinges on these. Be explicit which the chosen method needs. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#2-foundational-assumptions-state-them-always)

## A. Constraint-based (independence-test driven)

- PC algorithm (Peter–Clark; Spirtes, Glymour & Scheines, 2000): start from a complete undirected graph, remove edges via conditional-independence (CI) tests, then orient colliders and propagate (Meek rules). Output: CPDAG. Assumes causal sufficiency + faithfulness. Order-dependence fixed by PC-stable (Colombo & Maathuis, 2014). CI tests: Fisher-Z (linear-Gaussian), G²/χ² (discrete), KCI (kernel, nonlinear). — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#a-constraint-based-independence-test-driven)
- FCI (Fast Causal Inference) and RFCI: drop causal sufficiency - handle latent confounders and selection bias. Output: a PAG (Partial Ancestral Graph) over a MAG, with edge marks ○ (unknown), → (ancestor), ↔ (latent common cause) (Spirtes et al., 2000; Zhang, 2008). — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#a-constraint-based-independence-test-driven)

## B. Score-based search

- GES (Greedy Equivalence Search; Chickering, 2002): searches over CPDAG space with a two-phase forward (edge-add) / backward (edge-delete) greedy search, scoring with a decomposable, consistent score - BIC (continuous) or BDeu (discrete). Asymptotically returns the true equivalence class. fGES is the fast/parallel variant (TETRAD). — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#b-score-based-search)
- GIES (Hauser & Bühlmann, 2012): GES extended to interventional data - searches over interventional Markov equivalence classes, exploiting experiments to orient more edges. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#b-score-based-search)

## C. Permutation / ordering search

- GRaSP and BOSS (Lam, Andrews & Ramsey, 2022): search over variable orderings; more accurate and scalable than GES on many benchmarks, available in causal-learn and TETRAD. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#c-permutation-ordering-search)

## D. Functional causal models (FCMs) — orient beyond the equivalence class

- By assuming a functional form, these identify a unique DAG, not just a CPDAG. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#d-functional-causal-models-fcms-orient-beyond-the-equivalence-class)
  - LiNGAM - Linear, Non-Gaussian, Acyclic Model (Shimizu, Hoyer, Hyvärinen & Kerminen, 2006, JMLR 7:2003–2030): linear SEM with non-Gaussian noise → full causal order is identifiable. ICA-LiNGAM uses ICA; DirectLiNGAM (Shimizu et al., 2011) is regression-based and avoids ICA local optima. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#d-functional-causal-models-fcms-orient-beyond-the-equivalence-class)
  - ANM - Additive Noise Models (Hoyer, Janzing, Mooij, Peters & Schölkopf, 2008/2009): Y = f(X) + N with N ⟂ X. Nonlinear f breaks the X↔Y symmetry → cause/effect direction identifiable. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#d-functional-causal-models-fcms-orient-beyond-the-equivalence-class)
  - Post-Nonlinear (PNL) model (Zhang & Hyvärinen, 2009): Y = g(f(X) + N) - most general identifiable FCM. In causal-learn. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#d-functional-causal-models-fcms-orient-beyond-the-equivalence-class)

## E. Continuous-optimization / gradient methods

- Reframe combinatorial DAG search as smooth optimization with a differentiable acyclicity constraint - scales and integrates with deep learning. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#e-continuous-optimization-gradient-methods)
  - NOTEARS (Zheng, Aragam, Ravikumar & Xing, NeurIPS 2018): the acyclicity breakthrough - h(W) = tr(e^{W∘W}) − d = 0 is a smooth, exact characterization of acyclicity, solved via augmented Lagrangian. Originally linear; NOTEARS-MLP extends to nonlinear. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#e-continuous-optimization-gradient-methods)
  - GOLEM (Ng, Ghassami & Zhang, NeurIPS 2020): likelihood-based score with soft acyclicity - faster and more accurate than NOTEARS in the linear-Gaussian/EV setting. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#e-continuous-optimization-gradient-methods)
  - DAG-GNN (Yu et al., ICML 2019): VAE/GNN variant for nonlinear and discrete data. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#e-continuous-optimization-gradient-methods)
  - Caveat: Reisach, Seiler & Weichwein (NeurIPS 2021, "Beware of the Simulated DAG") showed continuous-optimization methods can exploit varsortability - marginal-variance artifacts of synthetic data scaling. Standardize data and don't trust synthetic-benchmark wins blindly. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#e-continuous-optimization-gradient-methods)

## F. Time-series causal discovery

- Granger causality: X Granger-causes Y if past X improves prediction of Y beyond Y's own past. Predictive, not structural; fails with latent confounders / instantaneous effects / nonlinearity. Use only as a baseline. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#f-time-series-causal-discovery)
- PCMCI / PCMCI+ (Runge et al., Science Advances 2019; PCMCI+ in UAI 2020): two-stage - a PC-style condition-selection step, then Momentary Conditional Independence (MCI) tests controlling for autocorrelation and indirect links. PCMCI+ adds contemporaneous links. Implemented in Tigramite; pairs with any CI test (ParCorr, GPDC, CMI). LPCMCI handles latent confounders. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#f-time-series-causal-discovery)
- VAR-LiNGAM (Hyvärinen et al., 2010): combines a VAR model with LiNGAM to recover both lagged and instantaneous causal effects. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#f-time-series-causal-discovery)

## Tools / Frameworks

- causal-learn (py-why, Python; Zheng et al., 2024; docs https://causal-learn.readthedocs.io/): the reference Python toolkit - PC, FCI, GES, GRaSP, BOSS, LiNGAM family, ANM, PNL, CD-NOD, plus CI tests and graph utilities. Default first choice for general discovery. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#tools-frameworks)
- gCastle (Huawei Noah's Ark Lab; Zhang et al., 2021): gradient-based focus (NOTEARS, GOLEM, DAG-GNN, GraN-DAG, ...), PyTorch + GPU, data simulators, and a built-in metrics module (SHD, FDR, TPR, F1, NNZ). — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#tools-frameworks)
- Tigramite (Runge; https://github.com/jakobrunge/tigramite): the standard for time-series discovery (PCMCI, PCMCI+, LPCMCI, RPCMCI). — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#tools-frameworks)
- pcalg (R; Kalisch et al., JSS 2012): mature PC/FCI/RFCI/GES with IDA effect estimation. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#tools-frameworks)
- DoWhy (py-why; https://www.pywhy.org/dowhy/): primarily inference, but its GCM module and dowhy.causal_discovery wrap discovery; good for the discover-then-refute workflow. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#tools-frameworks)
- CausalNex (QuantumBlack): NOTEARS-based structure learning + Bayesian-network reasoning, with expert-knowledge constraints (tabu edges, required edges). — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#tools-frameworks)
- TETRAD / py-tetrad: large library of search algorithms and the knowledge/background-constraint framework. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#tools-frameworks)

## Practical Patterns

- Always inject background knowledge. Forbidden edges, required edges, and tiered time order (a cause can't follow its effect) dramatically reduce the equivalence class. Every major tool supports knowledge/tabu constraints - use them. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#practical-patterns)
- Match method to assumptions and data type: — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#practical-patterns)
  - Possible latent confounders → FCI / RFCI (get a PAG), not PC/GES. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#practical-patterns)
  - Linear + non-Gaussian noise → DirectLiNGAM (gets a full DAG). — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#practical-patterns)
  - Nonlinear, continuous → ANM / PNL, or NOTEARS-MLP / DAG-GNN. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#practical-patterns)
  - Discrete/categorical → score-based with BDeu, or G²-test PC. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#practical-patterns)
  - High-dim time series → PCMCI+. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#practical-patterns)
  - Have interventions/experiments → GIES or interventional NOTEARS. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#practical-patterns)
- Standardize/scale continuous variables before continuous-optimization methods to avoid varsortability artifacts. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#practical-patterns)
- Bootstrap for edge stability. Resample, re-run discovery, and report edge-presence and orientation frequencies rather than one point graph. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#practical-patterns)
- Discover → refute → estimate. Use discovery to propose a graph, validate with domain experts and refutation/sensitivity checks, then hand the validated DAG to da-12 for effect estimation. Discovery output is a hypothesis, not ground truth. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#practical-patterns)
- Evaluate with the right metric: — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#practical-patterns)
  - SHD (Structural Hamming Distance): count of edge insert/delete/reverse ops to match the truth - lower is better; compare against the CPDAG, not a DAG, when methods return equivalence classes. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#practical-patterns)
  - SID (Structural Intervention Distance; Peters & Bühlmann, 2015): counts intervention-distribution errors - closer to what matters for downstream effect estimation than SHD. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#practical-patterns)
  - Also F1 / precision / recall on the skeleton, FDR, TPR. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#practical-patterns)

## Anti-Patterns

- Reporting a single DAG when the method returns a CPDAG/PAG. Undirected / circle-marked edges are genuinely undetermined; orienting them implies assumptions you didn't make. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#anti-patterns)
- Treating Granger causality as structural causality. It's lagged prediction; silent on confounders and contemporaneous effects. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#anti-patterns)
- Trusting synthetic-benchmark performance of NOTEARS-family methods without standardizing data (varsortability - Reisach et al., 2021). — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#anti-patterns)
- Ignoring latent confounders. Running PC/GES when unmeasured common causes are plausible yields confident but wrong edges. Use FCI or sensitivity analysis. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#anti-patterns)
- Skipping faithfulness/sufficiency disclosure. Stakeholders must know the result is conditional on assumptions that can't be verified from data alone. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#anti-patterns)
- Using discovery output directly for policy. Discovery proposes; it does not prove. Validate before acting. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#anti-patterns)
- Doing effect estimation here. Backdoor adjustment, IV, DiD, propensity scores, synthetic control → da-12-ab-testing-causal-inference. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#anti-patterns)

## Troubleshooting

- Too many undirected edges in the CPDAG: expected with observational-only data. Add background knowledge, use an FCM method (LiNGAM/ANM) if assumptions hold, or collect interventional data. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#troubleshooting)
- Unstable edges across runs/bootstraps: likely faithfulness near-violations, small n, or wrong CI test. Increase data, switch CI test (e.g., KCI for nonlinearity), use PC-stable. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#troubleshooting)
- PC gives different graphs depending on variable order: use PC-stable (Colombo & Maathuis, 2014). — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#troubleshooting)
- Dense, implausible graph from NOTEARS: increase the L1 sparsity penalty, standardize data, threshold small weights; consider GOLEM. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#troubleshooting)
- Nonlinear relationships missed: linear methods (Fisher-Z PC, linear NOTEARS, LiNGAM) can't see them - use KCI tests, ANM/PNL, NOTEARS-MLP, or DAG-GNN. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#troubleshooting)
- Time-series links look confounded by autocorrelation: that's exactly what PCMCI (MCI step) controls for; plain Granger does not. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#troubleshooting)

## References

- Spirtes, Glymour & Scheines, Causation, Prediction, and Search, 2nd ed., 2000 - PC, FCI foundations. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#references)
- Andersson, Madigan & Perlman (1997) - characterization of Markov equivalence / CPDAGs. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#references)
- Chickering (2002) - Greedy Equivalence Search (GES). https://jmlr.org/papers/v3/chickering02b.html — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#references)
- Hauser & Bühlmann (2012) - GIES (interventional GES). https://jmlr.org/papers/v13/hauser12a.html — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#references)
- Shimizu, Hoyer, Hyvärinen & Kerminen (2006) - LiNGAM, JMLR. https://www.jmlr.org/papers/v7/shimizu06a.html — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#references)
- Shimizu et al. (2011) - DirectLiNGAM, JMLR. https://jmlr.org/papers/volume12/shimizu11a/shimizu11a.pdf — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#references)
- Hoyer et al. (2008/2009) - nonlinear additive noise models (ANM), NeurIPS. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#references)
- Zhang & Hyvärinen (2009) - Post-Nonlinear (PNL) model. https://arxiv.org/abs/1205.2599 — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#references)
- Zheng, Aragam, Ravikumar & Xing (2018) - NOTEARS, NeurIPS. https://arxiv.org/abs/1803.01422 — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#references)
- Ng, Ghassami & Zhang (2020) - GOLEM, NeurIPS. https://arxiv.org/abs/2006.10201 — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#references)
- Yu et al. (2019) - DAG-GNN, ICML. https://arxiv.org/abs/1904.10098 — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#references)
- Reisach, Seiler & Weichwein (2021) - "Beware of the Simulated DAG", NeurIPS. https://arxiv.org/abs/2102.13647 — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#references)
- Colombo & Maathuis (2014) - order-independent PC-stable, JMLR. https://jmlr.org/papers/v15/colombo14a.html — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#references)
- Lam, Andrews & Ramsey (2022) - GRaSP / BOSS. https://proceedings.mlr.press/v180/lam22a.html — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#references)
- Zhang (2008) - augmented FCI orientation rules for PAGs, AIJ. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#references)
- Runge et al. (2019) - PCMCI, Science Advances. https://www.science.org/doi/10.1126/sciadv.aau4996 — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#references)
- Runge (2020) - PCMCI+, UAI. https://proceedings.mlr.press/v124/runge20a.html — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#references)
- Hyvärinen et al. (2010) - VAR-LiNGAM, JMLR. https://jmlr.org/papers/v11/hyvarinen10a.html — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#references)
- Peters & Bühlmann (2015) - Structural Intervention Distance (SID). https://arxiv.org/abs/1306.1043 — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#references)
- Zheng et al. (2024) - causal-learn, JMLR; docs https://causal-learn.readthedocs.io/ — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#references)
- Zhang et al. (2021) - gCastle toolbox. https://arxiv.org/abs/2111.15155 — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#references)
- Kalisch et al. (2012) - pcalg, JSS. https://www.jstatsoft.org/article/view/v047i11 — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#references)
- Tigramite - https://github.com/jakobrunge/tigramite ; DoWhy - https://www.pywhy.org/dowhy/ ; CausalNex docs. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-32-causal-discovery/#references)

## Where this helps

- Observational data is available and the goal is to learn the causal graph itself, which variables cause which, rather than assuming a DAG and estimating one effect on it as causal inference methods do. — [source](https://llms-explorer.com/tree/causal-discovery-and-structure-learning/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Deciding which algorithm family fits the data: constraint-based methods like PC or FCI for independence-test-driven search, score-based GES for a global search over CPDAG space, or functional causal models like LiNGAM and ANM when a unique DAG is wanted rather than an equivalence class. — [source](https://llms-explorer.com/tree/causal-discovery-and-structure-learning/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Latent confounders, unmeasured common causes, might be present, which rules out causal-sufficiency-assuming methods like PC or GES and points toward FCI or RFCI instead. — [source](https://llms-explorer.com/tree/causal-discovery-and-structure-learning/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Time-series data is available and the goal is to test for Granger-style predictive causality or run PCMCI-family methods designed specifically for lagged and contemporaneous causal structure. — [source](https://llms-explorer.com/tree/causal-discovery-and-structure-learning/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*

## Project ideas

- Run the PC algorithm and FCI side by side on the same dataset with a plausible unmeasured confounder to see how the returned CPDAG differs from FCI's PAG output. — [source](https://llms-explorer.com/tree/causal-discovery-and-structure-learning/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Implement a non-Gaussian LiNGAM model on simulated data with known ground truth to confirm it recovers a fully oriented DAG where a constraint-based method would leave edges undirected. — [source](https://llms-explorer.com/tree/causal-discovery-and-structure-learning/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Build a small NOTEARS-style continuous-optimization structure learner using the smooth acyclicity constraint and compare its recovered graph on standardized versus unstandardized data. — [source](https://llms-explorer.com/tree/causal-discovery-and-structure-learning/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Apply PCMCI or PCMCI+ to a multivariate time series and compare its recovered lagged structure against naive pairwise Granger causality tests. — [source](https://llms-explorer.com/tree/causal-discovery-and-structure-learning/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*

## Common mistakes

- Reporting a single oriented DAG when the method actually returned a CPDAG or PAG — undirected or circle-marked edges are genuinely undetermined by the data, and orienting them anyway overstates what was learned. — [source](https://llms-explorer.com/tree/causal-discovery-and-structure-learning/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Treating Granger causality as if it were structural causality, when it is lagged prediction and stays silent on confounders and contemporaneous, same-timestep, effects. — [source](https://llms-explorer.com/tree/causal-discovery-and-structure-learning/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Trusting a NOTEARS-family method's strong synthetic-benchmark performance without standardizing the data first, since those benchmarks are sensitive to variable scale, the varsortability issue. — [source](https://llms-explorer.com/tree/causal-discovery-and-structure-learning/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Running PC without the PC-stable variant and being surprised the resulting graph changes depending on the order variables were entered in. — [source](https://llms-explorer.com/tree/causal-discovery-and-structure-learning/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*

## Known issues

- Constraint- and score-based methods like PC and GES only return a Markov equivalence class, a CPDAG, not a single causal DAG, so some edges are left undirected by design, not by algorithm failure. — [source](https://llms-explorer.com/tree/causal-discovery-and-structure-learning/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- PC and GES both assume causal sufficiency, no unmeasured common causes; when that assumption is false, their output is not just incomplete but can be actively misleading, which is why FCI and RFCI exist as causal-sufficiency-free alternatives. — [source](https://llms-explorer.com/tree/causal-discovery-and-structure-learning/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Functional causal models like LiNGAM and ANM can identify a unique DAG beyond the equivalence class, but only by assuming a specific functional form; a wrong functional assumption produces a confidently wrong DAG. — [source](https://llms-explorer.com/tree/causal-discovery-and-structure-learning/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Edge stability across bootstrap runs or reruns is a common practical failure mode, usually traced to near-violations of the faithfulness assumption, small sample size, or the wrong conditional-independence test for the data type. — [source](https://llms-explorer.com/tree/causal-discovery-and-structure-learning/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*

## Context files

- [Causal Discovery and Structure Learning](https://llms-explorer.com/downloads/sources/mdb-context-hub/da-32-causal-discovery.md)
