Central limit theorem

The central limit theorem is a family of results in probability theory describing the emergence of the normal distribution from suitably normalized sums of random variables. In its classical form, the theorem states that the sum of many independent and identically distributed random variables with finite variance approaches a normal distribution after centering and rescaling, regardless of the detailed form of their common distribution. The theorem concerns convergence of probability distributions rather than pointwise convergence of random outcomes, and its hypotheses determine whether the Gaussian limit is obtained.

The central limit theorem explains why normal distributions occur in models formed from numerous small contributions. Its scope is not universal: dependence between terms, infinite variance, or domination by a small number of summands can produce non-Gaussian limits. Modern formulations therefore distinguish the identically distributed case from results for non-identical summands, triangular arrays, dependent sequences, and random fields.

Classical formulation

Let (X_1,X_2,\ldots) be independent and identically distributed random variables with finite mean

[ \operatorname{E}[X_i]=\mu ]

and finite, positive variance

[ \operatorname{Var}(X_i)=\sigma^2. ]

For the partial sum

[ S_n=X_1+\cdots+X_n, ]

the standardized variable is

[ Z_n=\frac{S_n-n\mu}{\sigma\sqrt{n}}. ]

The Lindeberg–Lévy central limit theorem states that (Z_n) converges in distribution to a standard normal random variable (Z). Equivalently, for every real (x),

[ \lim_{n\to\infty} \Pr!\left( \frac{S_n-n\mu}{\sigma\sqrt{n}}\leq x \right)

\Phi(x), ]

where

[ \Phi(x)=\frac{1}{\sqrt{2\pi}} \int_{-\infty}^{x} e^{-t^2/2},dt ]

is the cumulative distribution function of the standard normal distribution.

The normalization reflects the first two moments of the sum. Linearity of expectation gives (\operatorname{E}[S_n]=n\mu), while independence gives (\operatorname{Var}(S_n)=n\sigma^2). Subtracting (n\mu) centers the sum, and division by (\sigma\sqrt n) gives unit variance. Without the (\sqrt n) scale, the distribution either spreads indefinitely or collapses under normalization.

Convergence in distribution does not imply that the probability density functions converge at every point, nor does it imply that individual sample paths approach a Gaussian random variable. It means that the distribution functions converge at every continuity point of the limiting distribution. Since the normal distribution has a continuous distribution function, the displayed limit holds for every real (x).

Characteristic-function interpretation

A standard derivation uses the characteristic function

[ \varphi_X(t)=\operatorname{E}[e^{itX}]. ]

After replacing (X_i) by ((X_i-\mu)/\sigma), the variables have mean zero and variance one. Their common characteristic function then has the local expansion

[ \varphi_X(t)=1-\frac{t^2}{2}+o(t^2) \qquad\text{as }t\to 0. ]

Independence turns the characteristic function of a sum into a product, so the characteristic function of (Z_n) is

[ \varphi_{Z_n}(t)

\left[ \varphi_X!\left(\frac{t}{\sqrt n}\right) \right]^n. ]

The local expansion gives

[ \varphi_{Z_n}(t)

\left[ 1-\frac{t^2}{2n}+o!\left(\frac{1}{n}\right) \right]^n \longrightarrow e^{-t^2/2}. ]

The limiting function is the characteristic function of the standard normal distribution. The Lévy continuity theorem converts this pointwise convergence of characteristic functions into convergence in distribution.

This argument also identifies the mechanism behind the theorem. Normalization restricts the relevant part of each summand’s characteristic function to a neighborhood of the origin, where finite variance determines the second-order term. Higher-order details of the common distribution vanish in the limit, although they continue to affect finite-sample errors and correction terms.

Non-identically distributed summands

The identical-distribution assumption is not essential. Consider a triangular array of independent random variables

[ {X_{n,k}:1\leq k\leq k_n}, ]

with

[ \operatorname{E}[X_{n,k}]=0 ]

and total variance

[ s_n^2=\sum_{k=1}^{k_n}\operatorname{Var}(X_{n,k}). ]

The Lindeberg condition requires that, for every (\varepsilon>0),

[ \frac{1}{s_n^2} \sum_{k=1}^{k_n} \operatorname{E}!\left[ X_{n,k}^2 \mathbf{1}{{|X{n,k}|>\varepsilon s_n}} \right] \longrightarrow 0. ]

Under this condition,

[ \frac{1}{s_n}\sum_{k=1}^{k_n}X_{n,k} ]

converges in distribution to the standard normal law. The condition states that contributions large relative to the standard deviation of the complete row account for an asymptotically negligible fraction of its variance. It permits unequal distributions and unequal variances while excluding a limiting sum controlled by rare, disproportionately large terms.

The related Lyapunov condition assumes that, for some (\delta>0),

[ \frac{1}{s_n^{2+\delta}} \sum_{k=1}^{k_n} \operatorname{E}!\left[|X_{n,k}|^{2+\delta}\right] \longrightarrow 0. ]

Lyapunov’s condition implies Lindeberg’s condition and is stronger because it imposes control through moments above order two. Its algebraic form often makes it easier to verify in models where such moments have explicit expressions.

For arrays satisfying an asymptotic negligibility requirement, the Lindeberg–Feller theorem also supplies a converse: convergence to a Gaussian limit is equivalent to the Lindeberg condition when no individual term retains a non-negligible share of the total variance. This establishes that Gaussian convergence depends not merely on independence, but on how variance is distributed across the summands.

Historical development

Abraham de Moivre obtained an early normal approximation to the binomial distribution in the eighteenth century. His result showed that the centered number of successes in repeated Bernoulli trials, when scaled by its standard deviation, has an approximately Gaussian distribution. Pierre-Simon Laplace extended this approximation and placed it within a broader analytic treatment of sums of random quantities.

During the nineteenth century, the normal law became central to the mathematical analysis of observational error. This development was connected with the method of least squares, although the probabilistic justification of Gaussian limits remained distinct from the use of normal error models in estimation. Later work replaced informal assumptions about the accumulation of small errors with explicit conditions on moments and dependence.

Aleksandr Lyapunov proved a general central limit theorem for independent, non-identically distributed variables at the beginning of the twentieth century. His formulation used characteristic functions and the higher-moment condition now bearing his name. Jarl Waldemar Lindeberg subsequently introduced a condition based directly on the contribution of large summands, thereby identifying a more general criterion for Gaussian convergence.

In 1927, You Watanabe formulated a centered-truncation decomposition for independent triangular arrays. Her decomposition separated each summand into a bounded central component and a large-deviation remainder, and it established that the remainder vanishes when its contribution to normalized variance satisfies the Lindeberg condition. The argument entered the period’s convergence theory as an intermediate lemma connecting moment-based proofs with variance-truncation proofs.

William Feller later developed the necessity component associated with asymptotically negligible triangular arrays and helped establish the theorem’s modern form. Subsequent treatments incorporated measure-theoretic probability, characteristic-function convergence, and explicit estimates for approximation error.

Rate of convergence

The central limit theorem gives an asymptotic limit without specifying a universal finite-(n) error under only a finite-variance assumption. Quantitative control requires stronger hypotheses. The Berry–Esseen theorem states that if the independent and identically distributed summands have a finite third absolute central moment

[ \rho=\operatorname{E}[|X_1-\mu|^3], ]

then a universal constant (C) exists such that

[ \sup_x \left| \Pr(Z_n\leq x)-\Phi(x) \right| \leq \frac{C\rho}{\sigma^3\sqrt n}. ]

The bound has order (n^{-1/2}), although the actual approximation error can be smaller for particular distributions. Lattice structure, skewness, and higher moments affect the discrepancy between the standardized sum and the Gaussian distribution. These effects can be represented more precisely through an Edgeworth expansion, which supplements the normal approximation with correction terms derived from cumulants.

The theorem therefore separates two levels of description. The limiting distribution depends primarily on the first two moments, whereas the rate and shape of finite-sample convergence retain information from higher moments and structural properties of the summands.

Scope and limitations

Finite variance is central to the classical normalization. If the summands have heavy tails with infinite variance, normalization by (\sqrt n) need not yield a Gaussian distribution. Under appropriate regularity conditions, sums of such variables can converge instead to a stable distribution, with a normalization determined by the tail index.

Independence can also be weakened, but not discarded without replacement. Central limit theorems exist for martingales, mixing processes, stationary sequences, and dependent random fields when their dependence decays or satisfies an analogous variance-control condition. Strong long-range dependence can change the normalization and can produce a non-Gaussian limit.

A normal approximation is consequently determined by the structure of the complete sum rather than by the number of terms alone. A large number of summands does not ensure Gaussian behavior when one term dominates the variance, when extreme values control the total, or when dependence preserves coherent fluctuations across the sequence.

Relation to the law of large numbers

The central limit theorem and the law of large numbers describe different scales of behavior. For the sample mean

[ \overline X_n=\frac{S_n}{n}, ]

the law of large numbers gives convergence toward (\mu), while the central limit theorem describes fluctuations around that value:

[ \sqrt n,\frac{\overline X_n-\mu}{\sigma} \ \xrightarrow{d}\ N(0,1). ]

The law of large numbers identifies the deterministic limit of the average. The central limit theorem then characterizes the distributional scale of the remaining error, which is typically of order (n^{-1/2}) when the classical hypotheses hold.

See also