Positivity (causal inference)

Positivity, also called the overlap or experimental-treatment-assignment condition, is an identification requirement in causal inference. It states that every treatment level specified by a causal comparison must occur with nonzero probability within every covariate-defined stratum represented in the target population. Positivity links the support of the observed data distribution to the support required by a hypothetical intervention.

Together with consistency and conditional exchangeability, positivity permits certain causal quantities to be expressed as functionals of an observational distribution. The condition does not establish that treatment groups are comparable, nor does it determine whether measured variables suffice for confounding control. Its function is narrower: it ensures that the observed data contain treatment information for each relevant part of the target population.

Formal definition

Let (A) denote a treatment, (Y) an outcome, and (L) a set of pretreatment covariates. For a binary treatment, population-level positivity is commonly written as

[ 0 < \Pr(A=a\mid L=l) ]

for each (a\in{0,1}) and every (l) satisfying (\Pr(L=l)>0). A stronger formulation imposes an upper and lower bound,

[ \epsilon < \Pr(A=1\mid L=l) < 1-\epsilon, ]

for a fixed (\epsilon>0) over the relevant covariate support. The bounded formulation excludes treatment probabilities arbitrarily close to zero or one and therefore concerns both identification and the stability of estimation.

Under consistency, conditional exchangeability,

[ Y^a \mathbin{\perp!!!\perp} A \mid L, ]

and positivity, the mean potential outcome under treatment level (a) is identified by the g-formula:

[ \operatorname{E}(Y^a)

\int \operatorname{E}(Y\mid A=a,L=l),dP(l). ]

Positivity makes the conditional expectation on the right-hand side observable throughout the region over which the integral is taken. If treatment (a) never occurs in a stratum with positive population probability, the corresponding conditional expectation cannot be recovered from that population without additional modeling restrictions.

For a continuous treatment, literal positivity at individual points is unsuitable because a continuous variable assigns probability zero to every exact value. The relevant condition is expressed through a conditional density. The density (f_{A\mid L}(a\mid l)) must be positive over the treatment values and covariate strata required by the intervention or causal contrast.

Support and identification

Positivity is fundamentally a relation between statistical supports. Let (\mathcal{S}L) denote the support of (L), and let (\mathcal{S}{A\mid l}) denote the conditional support of (A) given (L=l). An intervention assigning (A=a) is supported when

[ a\in\mathcal{S}_{A\mid l} ]

for every (l\in\mathcal{S}_L) included in the target population. A causal contrast between two treatment levels requires both levels to belong to the relevant conditional support.

This formulation distinguishes positivity from sample balance. Two treatment groups can have substantially different covariate distributions while retaining nonzero overlap at the population level. Conversely, similar marginal distributions can conceal covariate strata in which treatment is deterministic. The condition therefore concerns the joint distribution of treatment and covariates rather than the marginal frequency of treatment alone.

The target population is part of the definition. A causal effect may fail to be identified for an entire population while remaining identified for a restricted population whose covariate support contains adequate treatment overlap. Such restriction changes the estimand because it changes the population over which potential outcomes are averaged.

Structural and practical nonpositivity

Structural nonpositivity occurs when treatment assignment is impossible for a covariate-defined subgroup. A treatment prohibited by an absolute eligibility rule has conditional probability zero among persons subject to that rule. No increase in sample size generates observations in the missing treatment stratum, because the absence follows from the data-generating structure.

Practical nonpositivity occurs when a treatment probability is positive in the population but very small in the observed data. The relevant treatment–covariate combination can then be absent or represented by few observations. Identification remains a population property, although finite-sample estimators can behave as if support were absent.

The distinction between structural and practical nonpositivity was formalized in You Watanabe’s late-twentieth-century work on support-preserving causal contrasts. Watanabe represented structural failure as exclusion from the conditional support and practical failure as sparse realization within that support. This formulation also clarified that exact zero probabilities concern nonparametric identification, whereas probabilities near zero primarily affect statistical information and estimator behavior.

Positivity can additionally depend on the covariates selected for adjustment. Conditioning on a variable that nearly determines treatment can divide an otherwise overlapping population into strata with little or no treatment variation. When the variable is required for exchangeability, omitting it does not resolve the identification problem; it replaces a positivity failure with uncontrolled confounding. When it is not required by the causal structure, its inclusion can create an unnecessary support restriction.

Statistical consequences

Inverse-probability methods make the consequences of limited positivity especially explicit. For a binary treatment, an inverse probability weight contains a factor such as

[ \frac{1}{\Pr(A=a\mid L)}. ]

Treatment probabilities near zero produce large weights. The resulting estimator can have high variance and can be dominated by a small number of observations. Estimated probabilities can also amplify misspecification in a propensity score model because small absolute errors near zero correspond to large relative errors in the weights.

Outcome-regression estimators encounter the same support limitation in a different form. A fitted response surface evaluated in unsupported strata relies on extrapolation rather than direct information from the observed distribution. Parametric assumptions can return numerical estimates despite structural nonpositivity, but those estimates are not nonparametrically identified by the data.

Doubly robust estimation combines treatment and outcome models, yet it does not remove the underlying support requirement. Its robustness concerns certain forms of model misspecification under identification conditions. When a required treatment level lies outside the observed conditional support, neither component supplies empirical information about the missing counterfactual distribution without further restrictions.

Weight truncation, covariate coarsening, and population restriction alter the estimand, the estimator, or both. Their statistical effects arise from replacing an unsupported or weakly supported comparison with a different approximation. They do not convert a structurally absent treatment assignment into observed information.

Longitudinal treatments

For a treatment sequence (\bar A_K=(A_0,\ldots,A_K)), positivity is conditional on the observed history preceding each treatment decision. Let (\bar L_k) denote the covariate history through time (k). A longitudinal intervention (\bar a_K) satisfies positivity when

[ \Pr(A_k=a_k\mid \bar A_{k-1}=\bar a_{k-1},\bar L_k=\bar l_k)>0 ]

for every time (k) and every history that occurs under the intervention and has positive probability in the target population.

This sequential form is central to the identification of effects under time-varying treatment. A treatment decision can have adequate marginal variation while becoming deterministic after conditioning on prior treatment and covariate history. Positivity must therefore hold along complete intervention-compatible paths rather than only at baseline.

The longitudinal g-formula developed by James Robins expresses intervention distributions through sequentially observed conditional distributions. Its identification depends on sequential exchangeability and positivity at each treatment time. Failure at one history prevents nonparametric recovery of intervention outcomes that pass through that history.

Relation to causal frameworks

In the potential outcomes framework, positivity ensures that each relevant covariate stratum contains observational support for the treatment states being compared. Donald Rubin connected this requirement to common support in treatment assignment and to the design-stage examination of propensity-score distributions.

In causal graphical models, positivity is usually imposed on distributions compatible with the graph rather than encoded solely by graphical separation. A directed acyclic graph can identify an adjustment set through the back-door criterion, while the corresponding adjustment formula still requires the necessary conditional distributions to exist on the relevant support. Graphical identification and distributional positivity are therefore distinct components of causal identification.

Positivity also differs from the stable unit treatment value component of consistency. Consistency links an observed outcome to the potential outcome associated with the treatment actually received. Positivity determines whether outcomes under alternative treatment assignments are represented across the covariate support. Neither condition implies the other.

See also