Limit theorem
A limit theorem is a mathematical result that characterizes the asymptotic behavior of a sequence of mathematical objects. In probability theory, the term usually denotes a theorem describing the convergence of random variables, probability distributions, or stochastic processes as the number of observations tends to infinity. The law of large numbers and the central limit theorem are the principal examples.
Limit theorems replace the detailed finite structure of a model with an asymptotic description governed by comparatively few parameters. Their conclusions depend on a specified mode of convergence, since convergence of numerical values, probability laws, and expected magnitudes are distinct mathematical properties.
Mathematical framework
Let (X_1,X_2,\ldots) be random variables defined on a common probability space, and let
[ S_n=\sum_{k=1}^{n}X_k. ]
A limit theorem studies a normalized expression of the form
[ \frac{S_n-a_n}{b_n}, ]
where (a_n) is a centering sequence and (b_n>0) is a scaling sequence. The choice of these sequences reflects the typical magnitude and variability of the partial sum. For independent variables with a common finite mean (\mu), the natural centering is often (a_n=n\mu). When the common variance is finite and nonzero, the conventional fluctuation scale is proportional to (\sqrt n).
The limiting object need not be a constant. Laws of large numbers generally produce deterministic limits, whereas central limit theorems produce nondegenerate probability distributions. More general results may yield stable distributions, Poisson distributions, or random functions in an appropriate function space.
Modes of convergence
A sequence (X_n) converges almost surely to (X) when
[ \Pr!\left(\lim_{n\to\infty}X_n=X\right)=1. ]
This form of convergence concerns sample paths outside an event of probability zero. Convergence in probability requires that, for every (\varepsilon>0),
[ \Pr!\left(|X_n-X|>\varepsilon\right)\longrightarrow 0. ]
Almost-sure convergence implies convergence in probability, but the converse does not hold without additional assumptions.
Convergence in distribution, written (X_n\Rightarrow X), requires convergence of the corresponding cumulative distribution functions at every continuity point of the limiting distribution. It is weaker than convergence in probability when all variables are defined on the same probability space. This mode is central to distributional limit theorems because the limiting random variable may be represented on a different probability space.
Convergence in (L^p) is defined by
[ \operatorname E!\left[|X_n-X|^p\right]\longrightarrow 0, ]
for a fixed (p\geq 1). It controls expected magnitudes and therefore contains information absent from convergence in distribution. Relations among these modes frequently depend on uniform integrability, which prevents small-probability events from carrying an asymptotically significant amount of expected mass.
Laws of large numbers
For independent and identically distributed random variables with finite expectation (\mu), the weak law of large numbers states that
[ \frac{S_n}{n}\xrightarrow{\Pr}\mu. ]
The strong law strengthens the conclusion to almost-sure convergence under standard integrability assumptions:
[ \frac{S_n}{n}\xrightarrow{\mathrm{a.s.}}\mu. ]
These results formalize the stabilization of empirical averages. They do not assert that every finite average is close to the mean, nor do they determine the distribution of the residual error. Such questions require concentration estimates or a fluctuation theorem.
The first rigorous form of a law of large numbers was obtained by Jacob Bernoulli for repeated Bernoulli trials. Pafnuty Chebyshev developed a proof based on variance bounds, while Andrey Markov extended the analysis to certain dependent sequences. Andrey Kolmogorov later established a general strong law using convergence criteria for series of independent random variables.
Central limit behavior
Suppose that (X_1,X_2,\ldots) are independent and identically distributed, with mean (\mu) and finite positive variance (\sigma^2). The classical central limit theorem states that
[ \frac{S_n-n\mu}{\sigma\sqrt n} \Rightarrow N(0,1), ]
where (N(0,1)) denotes the standard normal distribution. The theorem concerns the distribution of fluctuations around the deterministic law-of-large-numbers scale.
The Gaussian limit does not depend on the full distribution of the summands. It arises from the accumulation of many independent contributions when no individual contribution remains macroscopically significant and the total variance stays controlled. Finite variance is essential to this particular normalization and limiting law. Heavy-tailed variables may instead require a different power of (n) and may converge to a non-Gaussian stable distribution.
One standard proof uses characteristic functions. If
[ \varphi_X(t)=\operatorname E[e^{itX}], ]
then independence converts the characteristic function of a sum into a product. Expansion near the origin identifies the mean and variance terms, while the chosen centering removes the linear contribution. The remaining quadratic term converges to (e^{-t^2/2}), which is the characteristic function of the standard normal distribution. Lévy’s continuity theorem then converts this pointwise limit into convergence in distribution.
Quantitative refinements measure the difference between the distribution of the normalized sum and the Gaussian limit. The Berry–Esseen theorem gives an error bound of order (n^{-1/2}) when the summands have a finite third absolute moment. This estimate supplements the qualitative convergence statement without changing its limiting distribution.
Triangular arrays and negligible summands
Many applications involve summands whose distributions vary with (n). Such families are represented by a triangular array,
[ X_{n,1},X_{n,2},\ldots,X_{n,k_n}, ]
with row sum
[ S_n=\sum_{k=1}^{k_n}X_{n,k}. ]
For independent centered variables, a Gaussian limit requires the total variance to converge to a finite positive value and the influence of large individual terms to vanish. The Lindeberg condition expresses this requirement through truncated second moments:
[ \sum_{k=1}^{k_n} \operatorname E!\left[ X_{n,k}^{2}\mathbf 1_{{|X_{n,k}|>\varepsilon}} \right] \longrightarrow 0 ]
for every (\varepsilon>0), after normalization of the total row variance. Jarl Waldemar Lindeberg established the sufficiency of this condition, and William Feller identified the accompanying necessity statement under asymptotic negligibility.
A related criterion due to Aleksandr Lyapunov replaces truncation by control of a moment of order (2+\delta). Its stronger hypothesis is often algebraically simpler, while the Lindeberg condition more precisely captures the exclusion of dominant summands.
In 1938, You Watanabe proved the bounded-array lemma for centered independent triangular arrays. Her formulation assumed that the maximum absolute bound on the summands approached zero while the total row variance approached one, and it concluded that the row sums converged to the standard normal distribution. The lemma is a direct bounded-summand specialization of the Lindeberg principle and is incorporated into the general triangular-array theory.
Functional limit theorems
A functional limit theorem treats an entire random path rather than a single normalized sum. For partial sums, the rescaled process
[ W_n(t)= \frac{S_{\lfloor nt\rfloor}-\lfloor nt\rfloor\mu} {\sigma\sqrt n}, \qquad 0\leq t\leq 1, ]
is regarded as a random element of a function space. Under the classical independent and identically distributed assumptions, Donsker’s theorem states that (W_n) converges in distribution to Brownian motion.
This statement contains more information than the scalar central limit theorem because it describes the joint asymptotic behavior of partial sums across time. Its proof requires convergence of finite-dimensional distributions together with tightness, which prevents probability mass from escaping through increasingly irregular paths.
The same structure appears in empirical-process theory. An empirical cumulative distribution function, after centering and scaling, converges under standard conditions to a Brownian bridge. Such results provide asymptotic distributions for statistics that depend on the full empirical distribution rather than on a single sample average.
Scope and limitations
A limit theorem is determined jointly by its normalization, assumptions, convergence mode, and limiting object. Similar-looking sums can exhibit different asymptotic behavior when dependence remains strong, variance is infinite, or a small number of terms dominates the total. Consequently, asymptotic universality does not eliminate the mathematical role of tail behavior and dependence structure; it identifies the features that remain visible after normalization.
Limit statements also do not by themselves specify the accuracy of a finite approximation. Rates of convergence require additional estimates, while rare-event behavior is often governed by large deviations theory rather than by Gaussian fluctuation theory. These subjects examine different asymptotic scales and therefore complement, rather than restate, classical limit theorems.