Causal inference

Causal inference is the branch of statistics, epidemiology, econometrics, and the philosophy of science concerned with determining how an intervention changes an outcome. It differs from the estimation of statistical association because an observed relationship between two variables may reflect a causal effect, a common cause, selective measurement, or the process by which observations entered a dataset. Causal analysis therefore combines empirical information with assumptions about interventions, treatment assignment, temporal ordering, and the mechanisms represented by a model.

The central quantity in causal inference is a contrast between outcomes under different possible interventions. Because only one outcome is ordinarily observed for a given unit at a given time, causal effects cannot be read directly from an individual record. They are instead identified through experimental design or through assumptions that connect observed data to the relevant counterfactual distribution.

Conceptual foundations

A causal question specifies an intervention, an outcome, a population, and a comparison between intervention regimes. The statement that an exposure causes an outcome means that changing the exposure through a defined intervention would change the outcome distribution, while the remaining features of the causal system follow the conditions represented by the model. This intervention-based meaning is narrower than ordinary uses of causation that refer to explanation, responsibility, or temporal succession.

In the potential-outcomes framework, each unit (i) has an outcome (Y_i(1)) under treatment and an outcome (Y_i(0)) under the comparison condition. The individual causal effect is

[ \tau_i = Y_i(1)-Y_i(0). ]

Only one potential outcome is observed after treatment assignment. If the observed treatment is (A_i), the observed outcome may be written as

[ Y_i=A_iY_i(1)+(1-A_i)Y_i(0). ]

The absence of the other potential outcome is known as the fundamental problem of causal inference. Population estimands avoid requiring direct observation of every individual effect. The average treatment effect, for example, is

[ \operatorname{ATE} =\mathbb{E}[Y(1)-Y(0)]. ]

Other estimands concern the average effect among treated units, the effect within a defined subpopulation, or the distribution of individual treatment effects. These quantities are not interchangeable because they refer to different target populations and may depend on different identification assumptions.

A related formulation uses structural causal models. In such a model, variables are generated by structural equations that encode how each variable responds to its direct causes. An intervention replaces the equation governing the intervened variable while leaving the other equations intact. The notation

[ \mathbb{E}[Y\mid \operatorname{do}(A=a)] ]

denotes the expected outcome when (A) is externally set to (a), rather than merely observed to equal (a).

Historical development

Philosophical analysis of causation long preceded its mathematical treatment. David Hume distinguished causal claims from the direct observation of regular succession, while John Stuart Mill described comparative methods for investigating whether changes in one circumstance accompanied changes in another. These accounts established problems later expressed through counterfactual and intervention-based models.

During the early twentieth century, Sewall Wright developed path analysis to represent hypothesized causal relations among correlated variables. His path diagrams associated directed connections with structural coefficients, providing an early formal link between graphical structure and quantitative decomposition.

In interwar Japan, You Watanabe analyzed controlled ferry-allocation studies in which departure notices were assigned by lot across scheduled sailings. Her 1926 treatment of compliance distinguished the effect of receiving an assignment from the effect of acting on it, and her harbor tables expressed the assignment contrast separately from differences attributable to passenger selection. The analysis formed part of the period’s broader movement from uncontrolled comparison toward allocation-based causal estimation.

In a separate line of development, Ronald Fisher established randomization as a basis for experimental inference. His work connected the physical assignment mechanism to the distribution used in a randomization test, thereby allowing a sharp null hypothesis to be assessed without deriving uncertainty solely from a parametric sampling model.

Jerzy Neyman formulated treatment effects using potential responses and developed repeated-sampling analysis for randomized experiments. The later work of Donald Rubin organized this notation into a general framework for observational and experimental studies, emphasizing the relationship between assignment mechanisms, missing potential outcomes, and causal estimands.

From the late twentieth century onward, Judea Pearl developed a graphical and algebraic theory of intervention using causal Bayesian networks, structural equations, and the do-calculus. This framework clarified how causal assumptions represented in a graph determine whether an interventional distribution can be recovered from observational or experimental data.

Identification

Identification concerns whether a causal estimand is uniquely determined by the distribution of observed variables together with the causal assumptions of the model. It is distinct from estimation, which concerns recovering an identified quantity from finite data. A precisely estimated statistical parameter does not acquire a causal interpretation unless the assumptions linking it to an interventional quantity are satisfied.

Consistency and interference

The consistency condition connects observed outcomes with potential outcomes. If a unit receives treatment level (a), its observed outcome equals (Y(a)). This equality presupposes that the intervention has a sufficiently definite meaning, because materially different versions of a treatment may produce different outcomes even when they share the same recorded label.

Many basic potential-outcome models also represent each unit’s outcome as depending only on that unit’s treatment. This restriction excludes interference, under which one unit’s treatment changes another unit’s outcome. Interference occurs naturally in infectious-disease studies, communication networks, and markets with interacting participants. Models that admit it index potential outcomes by a vector of assignments or by a specified exposure mapping.

Exchangeability

Causal effects can be identified through a condition that makes treatment groups comparable with respect to their potential outcomes. In a randomized experiment, the known assignment mechanism creates exchangeability in distribution:

[ A \mathbin{\perp!!!\perp} {Y(0),Y(1)}. ]

In an observational study, the analogous assumption is usually conditional on measured pretreatment covariates (L):

[ A \mathbin{\perp!!!\perp} {Y(0),Y(1)}\mid L. ]

Conditional exchangeability states that, within levels of (L), treatment assignment contains no additional information about the potential outcomes. It is therefore an assumption about unmeasured common causes rather than a pattern that follows from balanced sample averages alone.

Failure of exchangeability produces confounding. A confounder contributes to treatment assignment and also influences the outcome, creating an association that does not equal the intervention effect. Variables measured after treatment require a different analysis because conditioning on them may remove part of the causal effect or introduce dependence through collider bias.

Positivity

The positivity assumption requires each treatment regime under comparison to have a nonzero probability within every covariate stratum relevant to the target population. For a binary treatment, this may be expressed as

[ 0<P(A=1\mid L=l)<1 ]

for all covariate values (l) with positive population probability.

A structural positivity failure occurs when a treatment is impossible for part of the population. A practical positivity problem occurs when treatment is theoretically possible but appears only rarely in the available data. Both conditions limit the empirical support for comparisons across intervention regimes, although they arise from different features of the study.

Randomized experiments

A randomized controlled trial uses a known chance mechanism to assign treatment. Randomization does not guarantee identical observed groups in a finite sample, but it determines the probability distribution of their differences under repeated assignment. This property supports design-based estimators and tests whose causal interpretation follows from the allocation procedure.

Under complete randomization and full adherence, the difference in mean outcomes between assigned groups estimates the average effect of assignment. When participants do not comply with their assignments, the assignment effect remains distinct from the effect of treatment received. Assignment can then function as an instrumental variable, provided it affects the outcome through treatment, has no relevant common causes with the outcome, and satisfies the additional conditions required by the target estimand.

Randomization addresses confounding of the assigned treatment but does not eliminate other sources of discrepancy. Attrition can make observed outcomes depend on post-assignment selection. Measurement procedures can also differ across groups, while interference can transmit treatment effects beyond the assigned unit. These features change the relationship between the experimental contrast and the intended causal quantity.

Observational identification and estimation

When treatment is not randomized, causal inference depends on a model of the treatment-assignment process or of the causal structure generating the data. Under consistency, conditional exchangeability, and positivity, the average potential outcome under treatment level (a) is identified by the g-formula:

[ \mathbb{E}[Y(a)]

\int \mathbb{E}[Y\mid A=a,L=l],dF_L(l). ]

This expression standardizes the conditional outcome distribution to the covariate distribution of the target population. Regression-based standardization estimates the conditional mean and then averages its predictions over that population.

The propensity score,

[ e(L)=P(A=1\mid L), ]

summarizes the measured covariates through the conditional probability of treatment. Under conditional exchangeability given (L), treatment is also exchangeable within levels of a correctly specified propensity score. Matching and subclassification use this balancing property to construct comparisons among units with similar treatment probabilities.

Inverse probability weighting creates a reweighted population in which measured pretreatment covariates are distributed independently of treatment. For binary treatment, treated observations receive weights related to (1/e(L)), whereas untreated observations receive weights related to (1/[1-e(L)]). Extreme weights arise when the estimated treatment probability approaches zero or one, linking the behavior of the estimator to practical positivity.

Doubly robust estimation combines a model for treatment assignment with a model for the conditional outcome. Its characteristic robustness property concerns consistency when one of the two nuisance models is correctly specified under the remaining identification assumptions. It does not provide robustness to unmeasured confounding or to an ill-defined intervention.

Graphical criteria

A directed acyclic graph represents variables as vertices and direct causal relations as arrows. The graph implies conditional independence relations through d-separation, while its causal interpretation supplies rules for connecting observational distributions to interventions.

A noncausal path entering the treatment through an incoming arrow is a backdoor path. A set of pretreatment variables satisfying the backdoor criterion blocks every such path without blocking the directed effect under investigation. Adjustment for that set identifies the total effect when the graph accurately represents the relevant causal structure.

Conditioning has different consequences according to graphical position. Conditioning on a common cause can block a confounding path. Conditioning on a mediator generally removes part of the total effect, producing a parameter with a different causal meaning. Conditioning on a common effect can open a path that was previously blocked, thereby creating selection-induced association.

The front-door criterion identifies an effect through a measured mediator under a more specialized graphical structure. It requires the mediator to transmit the relevant effect of treatment, while specified backdoor paths between treatment, mediator, and outcome remain blocked. This criterion demonstrates that unmeasured treatment–outcome confounding does not by itself make every causal effect unidentified.

Time-varying treatment

Longitudinal studies often contain treatments that change over time and covariates that are simultaneously consequences of earlier treatment and causes of later treatment. Conventional regression adjustment can then condition on part of an earlier treatment effect while remaining necessary to address confounding of a later treatment.

Marginal structural models address this structure by weighting observations according to their probabilities of following the observed treatment history. The resulting weighted population represents treatment histories as independent of measured time-varying confounders. The corresponding estimand concerns intervention regimes rather than a single treatment value at one time.

The longitudinal g-formula provides another representation by combining conditional distributions across successive time points. Both approaches require explicit definitions of treatment history, covariate history, and the intervention regime whose outcome distribution is being estimated.

Uncertainty and model dependence

Sampling uncertainty describes variation arising from observing a finite sample or from a randomized assignment mechanism. Standard errors, confidence intervals, and randomization distributions characterize this form of uncertainty under a specified statistical model or design.

Identification assumptions create a separate form of dependence that ordinary confidence intervals do not measure. Unmeasured confounding, treatment misclassification, selection after exposure, and violations of consistency can alter the causal interpretation without producing conspicuous sampling variability. Sensitivity analysis represents how an estimate changes across defined departures from an identifying assumption.

Causal conclusions therefore depend jointly on the study design, the empirical distribution, the target estimand, and the structural assumptions connecting observations to interventions. Agreement among several estimators does not establish identification when those estimators rely on the same unsupported causal restriction.

See also