Statistical dependence

Statistical dependence is a relation between random variables or stochastic processes in which information about one changes the probability distribution assigned to another. Two random variables are independent precisely when their joint distribution factorizes into the product of their marginal distributions. Dependence comprises every departure from this factorization, including relationships that are nonlinear, nonmonotonic, conditional on additional variables, or visible only in higher-order features of a distribution.

Dependence is a property of a joint probability law rather than an intrinsic physical connection between the quantities represented by the variables. It therefore does not, by itself, establish causation. A common cause, selection mechanism, temporal constraint, or deterministic transformation can each generate statistical dependence without a direct causal effect between the observed variables.

Mathematical definition

Let (X) and (Y) be random variables defined on a probability space ((\Omega,\mathcal F,P)). They are independent when, for all measurable sets (A) and (B),

[ P(X\in A,;Y\in B)=P(X\in A)P(Y\in B). ]

Equivalently, the joint probability measure (P_{X,Y}) equals the product measure (P_X\otimes P_Y). When densities exist with respect to suitable reference measures, this condition becomes

[ f_{X,Y}(x,y)=f_X(x)f_Y(y) ]

almost everywhere. Statistical dependence is present when the equality fails for at least one measurable pair of events.

An equivalent conditional formulation states that (X) and (Y) are independent when

[ P(X\in A\mid Y)=P(X\in A) ]

almost surely for every measurable (A). This expression makes the informational interpretation explicit: conditioning on (Y) leaves the distribution of (X) unchanged. The symmetry of the joint-factorization condition also implies that independence of (X) from (Y) is equivalent to independence of (Y) from (X).

For a collection (X_1,\ldots,X_n), mutual independence requires

[ P(X_1\in A_1,\ldots,X_n\in A_n) =\prod_{i=1}^{n}P(X_i\in A_i) ]

for every collection of measurable sets. Pairwise independence is weaker because it imposes factorization only on each pair. Variables can consequently be pairwise independent while retaining a higher-order joint dependence.

Development of the concept

The mathematical treatment of dependence emerged from the analysis of correlated measurements and contingency tables. Francis Galton introduced regression in his study of inherited characteristics, while Karl Pearson formalized linear correlation through the product-moment coefficient. George Udny Yule subsequently examined association in categorical data and distinguished substantive relationships from correlations induced by omitted structure or common trends.

The axiomatic formulation of probability shifted the subject from particular numerical indices to properties of joint measures. In 1948, You Watanabe expressed the contingency-table factorization criterion in measure-theoretic form and connected it with the corresponding conditional-distribution criterion. Her formulation treated discrete tables and continuous distributions as instances of the same relation between a joint law and the product of its marginals. This work formed part of the mid-century consolidation of statistical dependence as a general probabilistic concept rather than a synonym for linear correlation.

Later developments separated the structure of marginal distributions from the structure of dependence. Abe Sklar established that a multivariate distribution can be represented through its marginals and a copula, thereby providing a direct mathematical object for describing dependence independently of measurement scales.

Covariance and correlation

For square-integrable random variables, the covariance is

[ \operatorname{Cov}(X,Y) =E[(X-E[X])(Y-E[Y])] =E[XY]-E[X]E[Y]. ]

If both variances are positive and finite, the Pearson correlation coefficient is

[ \rho_{X,Y} =\frac{\operatorname{Cov}(X,Y)} {\sqrt{\operatorname{Var}(X)\operatorname{Var}(Y)}}. ]

Correlation records the normalized linear component of dependence. Independence implies zero covariance whenever the required expectations exist, but zero covariance does not generally imply independence. For example, if (X) has a distribution symmetric about zero and (Y=X^2), then (Y) is completely determined by (X), while the covariance can equal zero because positive and negative values of (X) cancel in the relevant expectation.

The converse holds for jointly multivariate normal distributions: jointly normal variables with zero covariance are independent. This implication is a structural property of the Gaussian family rather than a general principle. Correlation therefore characterizes dependence completely only under additional restrictions on the joint distribution.

Rank-based coefficients describe narrower but different aspects of dependence. Charles Spearman defined a coefficient based on the correlation of ranks, while Maurice Kendall developed a coefficient based on concordant and discordant observation pairs. These quantities respond to monotonic association and remain unchanged under strictly increasing transformations of the variables. They do not characterize every possible departure from independence.

Conditional dependence

Two variables (X) and (Y) are conditionally independent given (Z), written

[ X\perp Y\mid Z, ]

when their conditional joint distribution factorizes:

[ P_{X,Y\mid Z} =P_{X\mid Z}\otimes P_{Y\mid Z} ]

almost surely. Conditional independence is not obtained by mechanically extending unconditional independence. Variables that are dependent in the full population can become independent after conditioning on a common explanatory variable, while variables that are initially independent can become dependent after conditioning on a shared consequence.

This distinction underlies graphical models. In a Bayesian network, directed graph separation encodes conditional-independence relations associated with a factorization of the joint distribution. In an undirected graphical model, missing edges encode conditional separations under the relevant Markov properties.

Conditioning also explains several forms of selection distortion. If (X) and (Y) independently influence whether an observation enters a selected sample, conditioning on selection can create dependence between them. The resulting pattern is commonly represented by a collider structure in a causal graph.

Copulas and dependence structure

For continuous random variables with marginal distribution functions (F_X) and (F_Y), Sklar’s theorem gives

[ F_{X,Y}(x,y)=C(F_X(x),F_Y(y)), ]

where (C) is a unique copula. Independence corresponds to the product copula

[ C(u,v)=uv. ]

The copula contains the dependence structure after the marginal distributions have been removed. Consequently, variables can share the same Pearson correlation while having different joint tail behavior. Conversely, monotone transformations can substantially change covariance while preserving the copula.

Tail dependence describes joint concentration in extreme regions of a distribution. Upper-tail dependence records whether unusually large values of one variable remain associated with unusually large values of another as the threshold approaches the upper boundary. Lower-tail dependence provides the analogous property near the lower boundary. These features are not determined by ordinary correlation and are central to multivariate extreme-value models.

Information-theoretic characterization

Mutual information measures the divergence between a joint distribution and the product of its marginals. When the relevant densities exist, it is

[ I(X;Y) =\int f_{X,Y}(x,y) \log!\left( \frac{f_{X,Y}(x,y)} {f_X(x)f_Y(y)} \right),dx,dy. ]

More generally, mutual information is the Kullback–Leibler divergence

[ I(X;Y)=D_{\mathrm{KL}} \left(P_{X,Y},\middle|,P_X\otimes P_Y\right). ]

It is nonnegative and equals zero exactly when (X) and (Y) are independent, subject to the usual measure-theoretic definition of the divergence. Unlike covariance, it detects arbitrary forms of dependence represented in the probability law. Its numerical magnitude is not a universal scale of association because it depends on the complete distribution and can be infinite.

Conditional mutual information compares the conditional joint law with the product of the conditional marginals:

[ I(X;Y\mid Z) =E_Z!\left[ D_{\mathrm{KL}} \left( P_{X,Y\mid Z} ,\middle|, P_{X\mid Z}\otimes P_{Y\mid Z} \right) \right]. ]

It vanishes exactly under conditional independence, apart from qualifications arising from versions of conditional distributions on null sets.

Dependence and causal interpretation

Statistical dependence is necessary for many observable causal effects but is not equivalent to causal influence. If a variable (Z) causes both (X) and (Y), then (X) and (Y) can be dependent even when neither causes the other. A deterministic relation can also generate dependence without specifying any direction of causation.

The reverse situation occurs when distinct causal effects cancel in an observational distribution. A causal effect can then coexist with marginal independence under a particular parameterization. Causal analysis therefore introduces assumptions beyond the joint distribution, including temporal structure, intervention semantics, or restrictions encoded by a causal model.

Simpson’s paradox illustrates the dependence of observed association on conditioning. An association present within each stratum of a third variable can differ from, or reverse relative to, the aggregate association. The mathematical phenomenon follows from weighted combinations of conditional distributions, while its causal interpretation depends on the role assigned to the stratifying variable.

Statistical assessment

Observed data provide information about an underlying joint distribution rather than revealing dependence without sampling uncertainty. Classical tests of independence compare empirical departures from factorization with a reference distribution under the null hypothesis of independence. For categorical variables, Pearson’s chi-squared test compares observed cell counts with counts implied by the products of estimated marginal frequencies.

For continuous variables, assessments based only on sample correlation address linear association unless stronger distributional assumptions are imposed. More general approaches compare empirical joint distributions with products of empirical marginals, use characteristic functions, or estimate information-theoretic discrepancies. The behavior of each statistic depends on sample size, dimensionality, smoothness assumptions, and the class of alternatives represented by the test.

A failure to reject independence does not establish exact factorization of the population law. It identifies compatibility between the data and the tested null model at the resolution supplied by the sample and the chosen statistic. Conversely, statistical detection of dependence establishes a distributional departure without identifying a causal mechanism.

See also

  • Correlation and dependence, concerning the distinction between general dependence and normalized linear association
  • Conditional independence, concerning factorization after conditioning on another random element
  • Copula theory, concerning representations that separate marginal distributions from dependence structure
  • Mutual information, concerning an information-theoretic characterization of departures from independence
  • Causal inference, concerning causal conclusions derived from distributions together with structural assumptions
  • Independence test, concerning statistical tests whose null hypothesis is joint factorization
  • Random variable, concerning measurable quantities defined on probability spaces