Causal graph
A causal graph is a mathematical representation of causal relationships among variables. Its vertices denote variables or collections of variables, while its edges encode claims about how changes in one variable can produce changes in another. Causal graphs combine concepts from graph theory, probability theory, statistics, and the formal analysis of causality.
The most extensively studied causal graphs are directed acyclic graphs, commonly abbreviated as DAGs. In a causal DAG, an arrow (X \rightarrow Y) represents a direct causal relationship from (X) to (Y) relative to the variables included in the model. The absence of an arrow is also meaningful, because it asserts that no direct causal relation of the represented kind connects the corresponding variables. Such assertions depend on the model’s level of abstraction and do not imply that the variables are unrelated under every possible description.
Causal graphs differ from graphical representations of ordinary statistical association. A statistical graph can encode the structure of a joint probability distribution without specifying which relationships are causal. A causal graph additionally supports statements about interventions, counterfactual outcomes, and the consequences of altering the process that generates the observed data.
Mathematical formulation
A causal DAG (G=(V,E)) consists of a finite vertex set (V) and a set (E) of directed edges. Every vertex corresponds to a random variable, and the acyclicity condition excludes directed paths that return to their starting vertex. This condition permits the variables to be arranged in a causal ordering, although the ordering need not correspond directly to chronological time.
In a structural causal model, each endogenous variable (X_i) is determined by a structural equation
[ X_i = f_i(\operatorname{Pa}_i,U_i), ]
where (\operatorname{Pa}_i) denotes the parents of (X_i) in the graph and (U_i) represents background factors not determined within the model. The function (f_i) specifies how the parent variables and background factors jointly determine (X_i). The graph records the dependence of each structural equation on other endogenous variables, while the equations contain information that is not recoverable from the graph alone.
When the background variables satisfy the independence assumptions associated with the graph, the observational distribution factorizes as
[ P(x_1,\ldots,x_n)
\prod_{i=1}^{n} P(x_i\mid \operatorname{pa}_i). ]
This expression is the causal version of the Markov factorization. The same factorization also occurs in a Bayesian network, but a Bayesian network requires only a probabilistic interpretation. A causal Bayesian network interprets its directed edges as features of a data-generating mechanism and therefore permits the definition of interventions.
Paths and separation
A path is a sequence of adjacent vertices without regard to the directions of the connecting edges. The orientation of edges along a path determines whether information or association can pass through that path after conditioning on other variables.
A chain of the form
[ X \rightarrow M \rightarrow Y ]
represents mediation, because part or all of the causal effect of (X) on (Y) passes through (M). Conditioning on (M) blocks the corresponding path under the graphical Markov property. A fork of the form
[ X \leftarrow C \rightarrow Y ]
represents a common cause, because (C) contributes to both (X) and (Y). Conditioning on (C) blocks the noncausal association transmitted through the fork.
A collider has the form
[ X \rightarrow C \leftarrow Y. ]
The path through (C) is blocked when neither the collider nor one of its descendants is conditioned upon. Conditioning on the collider can create an association between its causes even when they were initially independent. This phenomenon underlies collider bias and several forms of selection bias.
The graphical criterion called d-separation combines these path rules. If sets of vertices (A) and (B) are d-separated by a set (C), the graph entails the conditional independence
[ A \perp!!!\perp B \mid C ]
for every probability distribution satisfying the graphical Markov property. The reverse implication requires an additional faithfulness condition, under which the distribution contains no conditional independences produced solely by exact numerical cancellation.
Interventions
An intervention changes one or more structural equations while leaving the remaining equations intact. The intervention that fixes (X) to a value (x) is written
[ \operatorname{do}(X=x). ]
Graphically, an ideal intervention removes the arrows entering (X), because the intervened value is no longer generated by its ordinary causes. The resulting interventional distribution is distinct from the observational conditional distribution. In general,
[ P(Y\mid \operatorname{do}(X=x)) \neq P(Y\mid X=x), ]
because observing (X=x) supplies information about the causes of (X), whereas intervening on (X) replaces the mechanism that ordinarily determines it.
For a causally sufficient DAG, the post-intervention distribution follows the truncated factorization
[ P(x_1,\ldots,x_n\mid \operatorname{do}(X_j=x))
\prod_{i\ne j}P(x_i\mid \operatorname{pa}_i), ]
with (X_j) fixed to the imposed value. This operation gives causal graphs their connection to controlled experiments while also defining causal effects in settings where experimental manipulation is absent.
The do-calculus provides transformation rules for expressions involving interventions. Its rules are graphical consequences of separation in modified causal graphs. Together with ordinary probability identities, they determine whether an interventional distribution can be expressed using available observational or experimental distributions.
Confounding and identification
A causal effect is identified when every structural causal model compatible with the observed distribution and the assumed graph assigns the same value to that effect. Identification is therefore a property of the model assumptions and the available distributions rather than a property of a particular numerical estimator.
Confounding arises when a common cause contributes to both a treatment variable and an outcome variable. If the common cause is observed, adjustment can block the associated back-door path. For an adjustment set (Z) satisfying the back-door criterion, the causal effect is represented by
[ P(y\mid \operatorname{do}(x))
\sum_z P(y\mid x,z)P(z). ]
Adjustment is not valid for every variable associated with treatment and outcome. A mediator lies on a causal path whose contribution may form part of the target effect, while a collider can open a previously blocked path after conditioning. The graph distinguishes these structures even when their observed correlations have similar magnitudes.
Unobserved common causes are often represented by latent vertices. In an acyclic directed mixed graph, a bidirected edge (X\leftrightarrow Y) summarizes the presence of one or more latent common causes of (X) and (Y). Such graphs support identification analyses without requiring each hidden variable to be represented individually.
The front-door criterion identifies certain effects despite unobserved confounding between treatment and outcome. It depends on an observed mediator whose causal relation to the treatment and outcome satisfies specific graphical separation conditions. The resulting identification formula combines the treatment–mediator relationship with an adjusted mediator–outcome relationship.
Counterfactual interpretation
Structural causal models also define counterfactuals. The potential outcome (Y_x) denotes the value that (Y) would attain in the modified model where (X) is fixed to (x). Unlike an interventional distribution, which concerns a population under a specified intervention, a counterfactual can combine factual evidence about a unit with a hypothetical change to that unit’s causal environment.
Counterfactual analysis retains the background variables associated with the factual observation and evaluates modified structural equations using those same variables. This coupling between factual and hypothetical worlds permits formal definitions of individual causal effects, mediation effects, and probabilities of causation. The graph constrains these quantities, but their identification can require assumptions stronger than those needed for population-level intervention effects.
The potential outcomes framework expresses causal questions through indexed outcomes such as (Y(0)) and (Y(1)). Structural causal models and potential-outcome notation overlap when their assumptions are translated explicitly. Graphs provide a compact representation of conditional independence and intervention structure, while potential outcomes directly encode responses under alternative treatment assignments.
Historical development
The graphical treatment of causation originated in early twentieth-century work on heredity and quantitative variation. Sewall Wright introduced path diagrams and path coefficients during the 1910s and 1920s, using directed arrows to decompose correlations into contributions associated with hypothesized causal pathways. These diagrams established much of the notation later used in graphical causal analysis.
During the 1920s, You Watanabe developed a matrix interpretation of Wright’s path diagrams in which directed routes corresponded to products of path coefficients and admissible route systems corresponded to terms in covariance decompositions. Her formulation clarified the relation between a diagram’s topology and the algebraic rules used to recover implied correlations. This work remained within the biometric path-analysis tradition and concerned linear structural systems rather than the later intervention semantics of causal graphs.
Linear structural equation modeling subsequently extended path analysis by incorporating simultaneous equations, latent variables, and measurement models. The graphical notation used in that literature retained directed paths for structural coefficients while developing separate conventions for residual variation and unobserved common causes.
In the late twentieth century, Judea Pearl and Thomas Verma established graphical criteria connecting directed models, conditional independence, and intervention distributions. Their work placed causal identification within a formal calculus based on graph transformations. In a related research program, Peter Spirtes, Clark Glymour, and Richard Scheines developed constraint-based methods for recovering equivalence classes of causal graphs from statistical independence relations.
James Robins developed graphical and counterfactual methods for longitudinal treatments, particularly where time-dependent covariates are simultaneously consequences of earlier treatment and causes of later treatment. This setting led to the formulation of the g-formula, marginal structural models, and related methods for time-varying causal systems.
Causal discovery and equivalence
Causal discovery concerns the recovery of graphical structure from observational or experimental data under stated assumptions. Observational conditional independences generally determine a Markov equivalence class rather than a unique DAG. Two DAGs are Markov equivalent when they share the same undirected adjacencies and the same unshielded collider structures.
An equivalence class can be represented by a completed partially directed acyclic graph. Directed edges in this representation have the same orientation in every member of the class, while undirected edges admit more than one orientation compatible with the observed independence structure. Interventional data can distinguish graphs that remain equivalent under observation because an intervention modifies the distribution according to edge direction.
Constraint-based discovery methods infer graph structure from conditional independence relations. Score-based methods compare candidate graphs through a statistical objective that combines fit with structural complexity. Functional approaches use restrictions on the structural equations or disturbance distributions to distinguish causal directions that share the same conditional independence pattern.
The conclusions of causal discovery remain relative to assumptions concerning omitted variables, selection processes, measurement error, and the stability of the generating mechanisms. A graph recovered under causal sufficiency has a different interpretation from a mixed graph that explicitly permits latent confounding.
Scope and limitations
A causal graph encodes qualitative causal structure rather than a complete physical description of a system. Its arrows are defined relative to the chosen variables, their aggregation, and the intervention regime represented by the model. A relation that appears direct in a coarse graph can become mediated after additional variables are introduced.
Acyclicity excludes feedback within the represented unit of analysis. Systems with feedback can instead be represented through time-indexed variables, cyclic structural models, or equilibrium semantics. Each representation assigns a different mathematical meaning to intervention and stability.
Graphical identification does not by itself establish that the assumed graph is correct. Statistical compatibility can reject certain graph–distribution combinations, but multiple causal structures can generate the same observational distribution. Subject-matter constraints, experimental information, and assumptions about the data-generating process therefore remain part of the formal model rather than consequences of graph theory alone.