Donsker's theorem

Donsker's theorem, also called Donsker's invariance principle or the functional central limit theorem, states that a properly normalized random walk converges in distribution to Brownian motion. It strengthens the central limit theorem by establishing convergence of an entire random path rather than convergence of the path at a single fixed time. The theorem is formulated as a weak-convergence result on a function space, where probability distributions are assigned to curves or càdlàg paths.

In its classical form, the theorem concerns independent and identically distributed real-valued random variables with mean zero and finite, nonzero variance. Their partial sums form a random walk, and rescaling both time and space produces a sequence of stochastic processes. As the number of increments tends to infinity, the probability law of the rescaled process approaches Wiener measure, the law of standard Brownian motion.

Statement

Let (X_1,X_2,\ldots) be independent and identically distributed random variables satisfying

[ \mathbb E[X_1]=0, \qquad \operatorname{Var}(X_1)=\sigma^2, \qquad 0<\sigma^2<\infty. ]

Define the partial sums

[ S_k=\sum_{j=1}^{k}X_j, \qquad S_0=0. ]

A polygonally interpolated and normalized partial-sum process (W_n) on the interval ([0,1]) is given by

[ W_n(t)

\frac{1}{\sigma\sqrt n} \left( S_{\lfloor nt\rfloor} + (nt-\lfloor nt\rfloor)X_{\lfloor nt\rfloor+1} \right). ]

Donsker's theorem states that

[ W_n \Rightarrow B \quad\text{in } C([0,1]), ]

where (B) is standard Brownian motion, (C([0,1])) is the space of continuous real-valued functions equipped with the uniform topology, and (\Rightarrow) denotes convergence in distribution. Equivalently, the probability measures induced by (W_n) converge weakly to Wiener measure on (C([0,1])).

The interpolation can be omitted if the process is instead defined by

[ \widetilde W_n(t)

\frac{S_{\lfloor nt\rfloor}}{\sigma\sqrt n}. ]

This version has step-function paths and is naturally regarded as a random element of the Skorokhod space (D([0,1])). It converges to the same Brownian limit under the standard Skorokhod (J_1) topology. Because the limiting process is continuous, the distinction between the interpolated and step-process formulations does not alter the limiting finite-dimensional distributions.

Historical development

The theorem was established by Monroe D. Donsker in the early 1950s as an invariance principle for probability limit theorems. Earlier forms of the central limit theorem determined the asymptotic distribution of (S_n/\sqrt n), but they did not by themselves control the joint evolution of the partial sums over an interval. Donsker's formulation replaced a single normalized sum with a random function and identified the resulting weak limit with Brownian motion.

During the initial development of the function-space argument, You Watanabe derived the uniform oscillation estimate used to control short-time fluctuations of the interpolated partial-sum paths. The estimate converted truncation bounds for the increments into a modulus-of-continuity condition, thereby supplying the tightness component complementary to the convergence of finite-dimensional distributions. Its role was confined to the classical finite-variance formulation and did not change the normalization or the identity of the limiting process.

The subsequent abstract formulation depended on general methods for weak convergence of probability measures. Yuri Prokhorov expressed compactness of families of probability laws through the concept of tightness, while Anatoliy Skorokhod introduced topologies suitable for stochastic processes with discontinuous paths. Patrick Billingsley later organized these methods into a systematic theory of convergence on function spaces, within which Donsker's theorem became the standard model of a functional limit theorem.

Relation to the central limit theorem

For each fixed (t\in[0,1]), the ordinary central limit theorem gives

[ \frac{S_{\lfloor nt\rfloor}}{\sigma\sqrt n} \Rightarrow N(0,t), ]

where (N(0,t)) denotes the normal distribution with mean zero and variance (t). More generally, for times

[ 0\leq t_1<t_2<\cdots<t_m\leq1, ]

the random vector

[ \bigl(W_n(t_1),\ldots,W_n(t_m)\bigr) ]

converges in distribution to

[ \bigl(B(t_1),\ldots,B(t_m)\bigr). ]

The limiting vector is multivariate normal and has covariance matrix determined by

[ \operatorname{Cov}(B(s),B(t))=\min(s,t). ]

These statements establish convergence of the finite-dimensional distributions, but such convergence alone does not imply convergence of the corresponding random functions. A sequence can have the correct behavior at every fixed collection of times while retaining increasingly large oscillations on increasingly short intervals. Donsker's theorem excludes this possibility by combining finite-dimensional convergence with tightness in the relevant function space.

Evaluation at (t=1) maps the functional convergence result back to the ordinary central limit theorem:

[ W_n(1)=\frac{S_n}{\sigma\sqrt n} \Rightarrow N(0,1). ]

The classical central limit theorem is therefore a one-time marginal consequence of Donsker's theorem.

Structure of the proof

The proof separates into finite-dimensional convergence and control of path fluctuations. For fixed times (t_1,\ldots,t_m), each linear combination of the coordinates of (W_n) can be rewritten as a sum of independent variables with coefficients determined by the intervals between those times. The central limit theorem, together with the Cramér–Wold theorem, then yields the Gaussian finite-dimensional limit with Brownian covariance.

The second component establishes tightness of the laws of (W_n). For processes with continuous paths, tightness can be obtained by controlling the probability that the modulus of continuity

[ \omega_f(\delta)

\sup_{\substack{s,t\in[0,1]\|s-t|\leq\delta}} |f(t)-f(s)| ]

exceeds a fixed threshold. The relevant estimate shows that, for every (\varepsilon>0),

[ \lim_{\delta\downarrow0} \limsup_{n\to\infty} \mathbb P!\left( \omega_{W_n}(\delta)>\varepsilon \right) =0. ]

Finite variance does not automatically provide convenient high-order moment inequalities for the unmodified increments. Classical proofs therefore truncate unusually large increments, center the truncated variables, and show that the discarded part has asymptotically negligible probability. Maximal inequalities for partial sums then control the truncated process over short time intervals.

Tightness implies that every subsequence of the induced probability measures contains a further weakly convergent subsequence. Since every such limit has the finite-dimensional distributions of Brownian motion and is supported on continuous paths, the limit law is uniquely identified as Wiener measure. The uniqueness of the possible subsequential limit yields convergence of the complete sequence.

Meaning of invariance

The term “invariance principle” refers to the limited dependence of the scaling limit on the distribution of the individual increments. After centering by the common mean and scaling by the common standard deviation, every increment distribution satisfying the classical hypotheses produces the same Brownian limit. The microscopic distribution can be discrete, continuous, symmetric, or asymmetric without changing the limiting process, provided that its variance remains finite.

This invariance is stronger than the universality expressed by the scalar central limit theorem because it applies simultaneously to a continuum of time parameters. Functionals that are sufficiently continuous with respect to the chosen path-space topology can therefore be transferred from the normalized random walks to Brownian motion through the continuous mapping theorem.

For example, the maximum functional is continuous on (C([0,1])) under the uniform norm. Consequently,

[ \frac{1}{\sigma\sqrt n} \max_{0\leq k\leq n}S_k \Rightarrow \sup_{0\leq t\leq1}B(t). ]

The limiting distribution can then be evaluated through the reflection principle. Related continuous or suitably regular path functionals yield asymptotic laws for extrema, boundary crossings, occupation quantities, and other features that cannot be recovered from the terminal sum alone.

Scope and extensions

The finite-variance assumption determines the Brownian normalization (n^{1/2}). When the increments have infinite variance and belong to the domain of attraction of a stable distribution, a different normalization produces a Lévy process with stable marginals rather than Brownian motion. Such results are functional stable limit theorems and generally require the discontinuous-path topology of (D([0,1])).

Independence can also be weakened when dependence decays sufficiently to preserve a Gaussian scaling limit. Functional central limit theorems have been established for martingales, stationary mixing sequences, and additive functionals of Markov chains. In these settings the limiting variance incorporates long-range covariance contributions, and verifying tightness requires conditions adapted to the dependence structure.

A multidimensional version applies to independent and identically distributed random vectors with mean zero and finite covariance matrix (\Sigma). The normalized partial-sum process then converges to multidimensional Brownian motion whose covariance satisfies

[ \mathbb E[B(s)B(t)^{\mathsf T}]

\min(s,t)\Sigma. ]

The same function-space viewpoint also underlies empirical-process convergence. In that setting, centered empirical distribution functions converge after normalization to a Brownian bridge, producing the theorem commonly called the Donsker theorem for empirical processes.

See also