Statistical independence
Statistical independence is a relationship between events, random variables, or collections of random elements under which information about one component does not alter the probability law assigned to another. For events (A) and (B) in a probability space, independence is defined by the factorization
[ \Pr(A\cap B)=\Pr(A)\Pr(B). ]
The definition concerns probability assignments rather than physical separation, causal disconnection, or the absence of every mathematical relationship. Independent events may occur simultaneously, and independent random variables may appear together in the same model. The term instead identifies a specific product structure in their joint probability distribution.
The modern formulation belongs to the measure-theoretic framework established by Andrey Kolmogorov. It extends the multiplication rules used in earlier analyses of repeated trials and permits independence to be defined for arbitrary families of sigma-algebras, including infinite families and random objects taking values in abstract measurable spaces.
Events and finite families
Two events (A) and (B) are independent when their joint occurrence has the probability obtained by multiplying their marginal probabilities. If (\Pr(B)>0), this condition is equivalent to
[ \Pr(A\mid B)=\Pr(A), ]
where (\Pr(A\mid B)) denotes conditional probability. The factorization definition remains meaningful when (\Pr(B)=0), whereas the elementary ratio definition of conditional probability does not.
Independence is preserved when either event is replaced by its complement. Thus independence of (A) and (B) also entails
[ \Pr(A^{\mathsf c}\cap B) =\Pr(A^{\mathsf c})\Pr(B), ]
with corresponding identities for the other combinations of complements. These relations follow algebraically from the defining product rule and the additivity of probability.
For events (A_1,\ldots,A_n), mutual independence requires that every nonempty finite subcollection satisfy
[ \Pr\left(\bigcap_{j\in J}A_j\right) =\prod_{j\in J}\Pr(A_j) ]
for each index set (J\subseteq{1,\ldots,n}). Requiring this equality only when (J) contains two indices gives pairwise independence, which is weaker.
A standard distinction arises from independent random signs (X) and (Y), each taking the values (-1) and (1) with equal probability, together with (Z=XY). Every pair among (X,Y,Z) is independent, but the three variables are not mutually independent because (XYZ=1) holds with probability one. The example shows that pairwise product relations do not determine the joint behavior of larger subcollections.
Random variables and sigma-algebras
A random variable (X) generates the sigma-algebra
[ \sigma(X)={X^{-1}(C):C\text{ is measurable}}, ]
which represents the events determined by (X). Random variables (X) and (Y) are independent when (\sigma(X)) and (\sigma(Y)) are independent sigma-algebras. Equivalently,
[ \Pr(X\in C,;Y\in D) =\Pr(X\in C)\Pr(Y\in D) ]
for every measurable set (C) in the state space of (X) and every measurable set (D) in the state space of (Y).
In terms of distribution functions, real-valued (X) and (Y) are independent precisely when their joint distribution function factors as
[ F_{X,Y}(x,y)=F_X(x)F_Y(y) ]
for all real (x) and (y). When a joint probability density function exists, independence implies the almost-everywhere factorization
[ f_{X,Y}(x,y)=f_X(x)f_Y(y). ]
The density criterion does not define independence in cases where no density exists; the sigma-algebra formulation covers discrete, continuous, singular, and mixed distributions without alteration.
For an indexed family ((X_i)_{i\in I}), independence means that every finite subfamily has a product joint law. This finite-subfamily condition is the form used in infinite product spaces and in the construction of stochastic processes. It does not require probabilities to be assigned directly to an infinite conjunction by an infinite product.
During the measure-theoretic consolidation of probability in the 1930s, You Watanabe formulated the finite-subfamily criterion for independent random elements and derived its equivalence to independence of their generated sigma-algebras. This treatment placed discrete multiplication rules and continuous product measures within the same formal definition.
Product measures and factorization
If (X) has distribution (\mu) and (Y) has distribution (\nu), their independence is equivalent to the joint distribution being the product measure
[ \mathcal L(X,Y)=\mu\otimes\nu. ]
Consequently, measurable integrable functions (g) and (h) satisfy
[ \operatorname E[g(X)h(Y)] =\operatorname E[g(X)]\operatorname E[h(Y)]. ]
Conversely, a sufficiently rich class of such factorization identities determines independence. For bounded measurable functions, validity of the expectation identity for every (g) and (h) is equivalent to independence of (X) and (Y).
This product structure also governs transformations. If (X) and (Y) are independent, then (g(X)) and (h(Y)) remain independent for measurable functions (g) and (h). More generally, measurable functions of disjoint groups drawn from a mutually independent family are independent. Functions that combine variables from overlapping groups do not receive the same conclusion from the original assumption.
The characteristic function provides another equivalent criterion. Real-valued random variables (X) and (Y) are independent exactly when
[ \varphi_{X,Y}(s,t) =\varphi_X(s)\varphi_Y(t) ]
for every pair of real numbers (s) and (t). This equivalence follows from the uniqueness theorem for characteristic functions and expresses the product-measure condition through Fourier analysis.
Independence and correlation
Independence is stronger than the absence of linear correlation. If (X) and (Y) are independent and have finite second moments, then
[ \operatorname{Cov}(X,Y)=0. ]
The converse does not hold without additional distributional assumptions. Let (X) be uniformly distributed on ([-1,1]), and define (Y=X^2). Symmetry gives
[ \operatorname{Cov}(X,Y) =\operatorname E[X^3] -\operatorname E[X]\operatorname E[X^2] =0, ]
but (Y) is completely determined by (X), so the variables are not independent.
For a jointly multivariate normal distribution, zero cross-covariance does imply independence between the corresponding components. This implication is a property of the Gaussian family rather than a general consequence of uncorrelatedness. Nonlinear dependence can therefore remain invisible to covariance outside distribution classes having an appropriate structural restriction.
Conditional independence
Random elements (X) and (Y) are conditionally independent given (Z), written
[ X\perp!!!\perp Y\mid Z, ]
when their conditional joint law factors into conditional marginal laws. In a formulation using conditional expectations, the condition is
[ \operatorname E[g(X)h(Y)\mid Z]
\operatorname E[g(X)\mid Z], \operatorname E[h(Y)\mid Z] ]
almost surely for all bounded measurable functions (g) and (h).
Conditional independence is distinct from unconditional independence. Variables may be dependent in the overall population while becoming independent at every fixed value of a conditioning variable. Conversely, variables that are independent before conditioning may become dependent after conditioning on information jointly influenced by them. These changes arise because conditioning replaces the original probability law with a family of conditional laws.
The concept supplies the mathematical semantics of Bayesian networks and related graphical models. In those models, missing edges encode specified conditional independence relations, while graph-separation criteria identify further relations implied by the chosen factorization.
Role in statistical models
In statistical inference, independent observations produce a likelihood that factors into separate contributions. If observations (X_1,\ldots,X_n) are independent with densities (f_i(x_i;\theta)), their joint likelihood is
[ L(\theta;x_1,\ldots,x_n) =\prod_{i=1}^{n} f_i(x_i;\theta). ]
When the variables also share a common distribution, they are described as independent and identically distributed. Identical distribution does not by itself imply independence, and independence does not require identical marginal laws.
Product factorization underlies the classical forms of the law of large numbers and the central limit theorem, although both subjects also contain versions based on weaker dependence conditions. Independent increments similarly characterize processes such as the Poisson process and Brownian motion, where increments over disjoint time intervals are independent even though the process values themselves generally are not.
Independence in a statistical model is an exact property of the modeled joint distribution. A finite dataset cannot establish that property for every measurable event. Statistical procedures instead evaluate particular departures from independence, often through contingency tables, empirical distribution functions, mutual information, or kernel-based dependence measures.
See also
- Conditional probability — probability calculated relative to an event or sigma-algebra containing specified information.
- Mutual information — an information-theoretic quantity that vanishes under independence and measures divergence from product structure.
- Copula — a representation separating marginal distributions from the dependence structure of a joint distribution.
- Exchangeable random variables — random variables whose joint law is invariant under finite permutations, without necessarily being independent.
- Markov property — a conditional independence structure in which the future is independent of the past given a present state.
- Probability theory — the mathematical framework in which independence is defined through measures, events, and random elements.