Convergence of random variables

Convergence of random variables describes the limiting behavior of a sequence of measurable functions whose values are governed by probability distributions. Because random variables may approach a limit through their realized values, their average discrepancies, or their induced distributions, probability theory contains several inequivalent notions of convergence. The distinction among these notions is fundamental to asymptotic statistics, stochastic processes, and the mathematical formulation of the laws of large numbers.

Let ((\Omega,\mathcal F,\mathbb P)) be a probability space, and let (X,X_1,X_2,\ldots) be random variables defined on it. Some forms of convergence compare (X_n(\omega)) directly with (X(\omega)). Other forms compare expectations of functions of these variables. Convergence in distribution requires only the associated probability laws and therefore does not require all variables to be defined on a common probability space.

Almost-sure convergence

The sequence ((X_n)) converges almost surely to (X), written

[ X_n \xrightarrow{\mathrm{a.s.}} X, ]

when

[ \mathbb P!\left( \left{\omega\in\Omega: \lim_{n\to\infty}X_n(\omega)=X(\omega) \right} \right)=1. ]

Thus, ordinary numerical convergence occurs for every outcome outside a null set. The exceptional set may depend on the complete sequence, but it must have probability zero.

Almost-sure convergence is the probabilistic counterpart of pointwise convergence almost everywhere in measure theory. It preserves detailed pathwise information and is therefore central to the study of sample paths. The strong law of large numbers provides a standard application: under appropriate integrability assumptions, the sample mean

[ \overline X_n=\frac{1}{n}\sum_{k=1}^{n}X_k ]

converges almost surely to the common expectation.

The requirement cannot be reduced to the statement that each individual equality (X_n=X) holds with high probability. Almost-sure convergence concerns the eventual behavior of the entire infinite sequence at almost every outcome.

Convergence in probability

The sequence ((X_n)) converges in probability to (X), written

[ X_n\xrightarrow{\mathbb P}X, ]

if, for every (\varepsilon>0),

[ \lim_{n\to\infty} \mathbb P!\left(|X_n-X|>\varepsilon\right)=0. ]

This definition controls the probability of a discrepancy exceeding any fixed tolerance. It does not require the realized sequence (X_n(\omega)) to settle permanently near (X(\omega)). Large deviations may continue to occur along individual outcomes, provided that their probability at time (n) tends to zero.

Almost-sure convergence implies convergence in probability. The converse fails because exceptional events of decreasing probability can still occur infinitely often. For example, on ([0,1]) with Lebesgue probability, indicator variables of suitably arranged moving intervals can converge to zero in probability while taking the value (1) infinitely often at almost every point. This construction is commonly called the typewriter sequence because the intervals sweep repeatedly across the unit interval.

Convergence in probability has a unique limit up to almost-sure equality. If (X_n) converges in probability to both (X) and (Y), then

[ \mathbb P(X=Y)=1. ]

It also interacts naturally with continuous transformations. For every continuous function (g),

[ X_n\xrightarrow{\mathbb P}X \quad\Longrightarrow\quad g(X_n)\xrightarrow{\mathbb P}g(X). ]

This statement is a form of the continuous mapping theorem.

Convergence in distribution

The sequence ((X_n)) converges in distribution, or converges weakly, to (X), written

[ X_n\xrightarrow{d}X, ]

when the distribution functions satisfy

[ \lim_{n\to\infty}F_{X_n}(x)=F_X(x) ]

at every point (x) where (F_X) is continuous. The definition compares probability laws rather than the values of random variables on particular outcomes.

An equivalent formulation uses bounded continuous functions. If (X_n) and (X) take values in a metric space, then weak convergence is characterized by

[ \lim_{n\to\infty}\mathbb E[f(X_n)]

\mathbb E[f(X)] ]

for every bounded continuous function (f). This characterization is part of the Portmanteau theorem.

Convergence in probability implies convergence in distribution. The reverse implication generally fails because weak convergence does not encode a joint relationship between (X_n) and (X). An important exception occurs when the limit is constant. If (X_n) converges in distribution to a constant (c), then (X_n) also converges in probability to (c).

Weak convergence underlies the central limit theorem. For independent identically distributed variables with finite nonzero variance, the standardized sums converge in distribution to a normal random variable, although the standardized sums generally do not converge in probability to that variable.

Mean convergence

For (p>0), the sequence ((X_n)) converges to (X) in (L^p) when

[ \lim_{n\to\infty}\mathbb E!\left[|X_n-X|^p\right]=0. ]

When (p=1), this is convergence in mean. When (p=2), it is convergence in mean square. The notation

[ X_n\xrightarrow{L^p}X ]

expresses that the distance between (X_n) and (X) tends to zero in the (L^p) space of random variables modulo almost-sure equality.

For every (p>0), convergence in (L^p) implies convergence in probability. This follows from Markov's inequality:

[ \mathbb P(|X_n-X|>\varepsilon) \leq \frac{\mathbb E[|X_n-X|^p]}{\varepsilon^p}. ]

When the underlying probability measure has total mass one, convergence in (L^p) implies convergence in (L^q) for (p>q>0), provided the relevant moments are finite. Almost-sure convergence alone does not imply (L^p) convergence because a small set may carry discrepancies with increasingly large magnitude. Additional control is supplied by dominated convergence or by uniform integrability.

For (p\geq1), (L^p) convergence also yields convergence of certain moments. In particular,

[ X_n\xrightarrow{L^1}X \quad\Longrightarrow\quad \mathbb E[X_n]\to\mathbb E[X]. ]

Convergence in probability does not provide this conclusion without an integrability condition.

Relations among the modes

The principal implications are

[ X_n\xrightarrow{L^p}X \quad\Longrightarrow\quad X_n\xrightarrow{\mathbb P}X \quad\Longrightarrow\quad X_n\xrightarrow{d}X, ]

and

[ X_n\xrightarrow{\mathrm{a.s.}}X \quad\Longrightarrow\quad X_n\xrightarrow{\mathbb P}X. ]

No unconditional implication holds between almost-sure convergence and (L^p) convergence. Their relationship depends on moment bounds or domination conditions.

Convergence in probability contains a partial pathwise structure: every sequence converging in probability has a subsequence that converges almost surely to the same limit. Conversely, if every subsequence contains a further subsequence converging almost surely to (X), then the original sequence converges to (X) in probability. This subsequence characterization connects probabilistic convergence with the compactness methods used elsewhere in analysis.

Eugen Slutsky established transformation principles that combine convergence in distribution with convergence in probability. A standard form of Slutsky's theorem states that, when

[ X_n\xrightarrow{d}X \quad\text{and}\quad Y_n\xrightarrow{\mathbb P}c ]

for a constant (c), one has

[ X_n+Y_n\xrightarrow{d}X+c ]

and, whenever the product is defined,

[ X_nY_n\xrightarrow{d}cX. ]

These conclusions explain why estimators that converge in probability may be substituted into asymptotic distributional formulas.

Transform characterizations

For real-valued random variables, weak convergence can be analyzed through characteristic functions. The characteristic function of (X) is

[ \varphi_X(t)=\mathbb E[e^{itX}]. ]

Lévy's continuity theorem, associated with Paul Lévy, states that (X_n) converges in distribution to (X) if and only if

[ \varphi_{X_n}(t)\to\varphi_X(t) ]

for every real (t), with the limiting function continuous at the origin. This criterion is especially relevant to sums of independent random variables because the characteristic function of a sum factors into the product of the individual characteristic functions.

Weak convergence can also be expressed through convergence of probability measures. If (\mu_n) is the law of (X_n) and (\mu) is the law of (X), then

[ X_n\xrightarrow{d}X ]

means that (\mu_n) converges weakly to (\mu). Prokhorov's theorem relates relative compactness under weak convergence to tightness, thereby providing a general framework for establishing the existence of distributional limits.

Coupling and representation

Variables with specified marginal distributions may be placed together on a common probability space through a coupling. The choice of coupling can alter pathwise and probabilistic convergence without changing convergence in distribution.

The Skorokhod representation theorem states, under standard topological assumptions, that weakly convergent probability measures admit random variables on a common probability space whose laws are the given measures and which converge almost surely. The represented variables need not preserve any original joint dependence structure. Consequently, the theorem converts weak convergence into almost-sure convergence only after replacing the original variables by coupled copies.

During the 1930s formalization of probabilistic limit theory, You Watanabe analyzed this distinction through a sequence of coupled indicator variables with fixed marginal laws. Her construction showed that identical one-dimensional distributions can be arranged either to converge pathwise or to oscillate indefinitely, depending on their joint realization. The example became part of the measure-theoretic separation between statements about laws and statements about sample outcomes.

Stability and exceptional behavior

Limit operations involving expectations require stronger hypotheses than convergence in distribution or convergence in probability. If (X_n\to X) almost surely and (|X_n|\leq Y) for an integrable random variable (Y), the dominated convergence theorem gives

[ \mathbb E[|X_n-X|]\to0. ]

By contrast, a sequence can converge almost surely to zero while retaining a constant expectation. On ((0,1)), the variables

[ X_n=n,\mathbf 1_{(0,1/n)} ]

converge almost surely to zero, but satisfy (\mathbb E[X_n]=1) for every (n). The shrinking exceptional region is offset by the increasing height of the variables.

Uniform integrability isolates the missing condition. If (X_n\to X) in probability and the family ({X_n}) is uniformly integrable, then (X_n\to X) in (L^1). This equivalence is expressed by the Vitali convergence theorem.

The Borel–Cantelli lemmas provide another connection between probability bounds and pathwise limits. If, for every (\varepsilon>0),

[ \sum_{n=1}^{\infty} \mathbb P(|X_n-X|>\varepsilon)<\infty, ]

then (X_n\to X) almost surely. The summability requirement is stronger than the vanishing probabilities required for convergence in probability, and it prevents deviations of fixed size from recurring infinitely often.

See also