Law of large numbers

The law of large numbers is a family of results in probability theory describing the convergence of empirical averages toward deterministic quantities. In its standard form, the law states that the arithmetic mean of repeated observations converges to the expected value of the underlying random variable when the observations satisfy appropriate assumptions concerning their distribution and dependence.

For random variables (X_1,X_2,\ldots), define the sample mean

[ \overline{X}n=\frac{1}{n}\sum{k=1}^{n}X_k. ]

If the variables are independent and identically distributed, with finite expectation

[ \mathbb{E}[X_k]=\mu, ]

then (\overline{X}_n) converges to (\mu). The precise meaning of convergence distinguishes the weak law from the strong law. Neither version states that short sequences must resemble the underlying distribution closely, nor does either version imply that random deviations disappear permanently after a fixed number of observations.

Mathematical formulation

Weak law

The weak law of large numbers states, under its classical assumptions, that the sample mean converges in probability to the expected value:

[ \overline{X}_n \xrightarrow{\mathrm{P}} \mu. ]

Equivalently, for every (\varepsilon>0),

[ \lim_{n\to\infty} \Pr\left( \left|\overline{X}_n-\mu\right|>\varepsilon \right)=0. ]

The probability of a deviation larger than any fixed positive tolerance therefore approaches zero. The statement concerns a sequence of probability distributions and does not assert that every realized infinite sequence eventually remains within that tolerance.

A direct proof is available when the variables are pairwise uncorrelated and have variances bounded by a common finite constant. In that setting,

[ \operatorname{Var}(\overline{X}_n)

\frac{1}{n^2}\sum_{k=1}^{n}\operatorname{Var}(X_k) \leq \frac{C}{n}. ]

Application of Chebyshev's inequality gives

[ \Pr\left( \left|\overline{X}_n-\mu\right|\geq\varepsilon \right) \leq \frac{C}{n\varepsilon^2}, ]

which tends to zero. Finite variance is sufficient for this proof but is not necessary for the independent and identically distributed weak law. Finite first moment alone is sufficient in the classical result associated with Aleksandr Khinchin.

Strong law

The strong law of large numbers strengthens the mode of convergence to almost sure convergence:

[ \overline{X}_n \xrightarrow{\mathrm{a.s.}} \mu. ]

This means

[ \Pr\left( \lim_{n\to\infty}\overline{X}_n=\mu \right)=1. ]

Thus, outside an event of probability zero, the sequence of sample means converges along the entire realized path. The strong law implies the weak law because almost sure convergence implies convergence in probability, while the converse does not hold without additional assumptions.

For independent and identically distributed random variables, the condition

[ \mathbb{E}[|X_1|]<\infty ]

is sufficient for the strong law. More general formulations replace identical distribution with conditions controlling the accumulated variances or the tails of the variables. One standard independent-variable criterion requires

[ \sum_{n=1}^{\infty} \frac{\operatorname{Var}(X_n)}{n^2} <\infty, ]

together with suitable centering. The resulting conclusion follows through convergence arguments associated with Kolmogorov's inequality and related results for sums of independent random variables.

Bernoulli trials

The earliest major form of the law concerns repeated Bernoulli trials. Let (X_k) equal (1) when the (k)-th trial succeeds and (0) otherwise, with a common success probability (p). The sample mean is then the relative frequency of success:

[ \overline{X}n=\frac{S_n}{n}, \qquad S_n=\sum{k=1}^{n}X_k. ]

Since

[ \mathbb{E}[\overline{X}_n]=p \quad\text{and}\quad \operatorname{Var}(\overline{X}_n)

\frac{p(1-p)}{n}, ]

the relative frequency converges in probability to (p). The strong law further states that

[ \Pr\left( \lim_{n\to\infty}\frac{S_n}{n}=p \right)=1. ]

This result concerns long-run stabilization of relative frequency. It does not state that the absolute discrepancy (S_n-np) converges to zero. That discrepancy generally grows in scale, while its ratio to (n) converges to zero. The distinction is quantified more precisely by the central limit theorem, which describes fluctuations on the order of (\sqrt n), and by the law of the iterated logarithm, which describes almost-sure extreme fluctuations on a finer asymptotic scale.

Historical development

Jacob Bernoulli proved the foundational Bernoulli-trial result in Ars Conjectandi, published posthumously in 1713. His theorem established that the probability of a substantial difference between an observed proportion and its underlying success probability becomes arbitrarily small when the number of trials becomes sufficiently large.

Siméon Denis Poisson introduced the expression “law of large numbers” in the nineteenth century and extended the analysis beyond Bernoulli's original formulation. Pafnuty Chebyshev subsequently developed a variance-based proof whose structure became the standard elementary derivation of the weak law.

The transition from convergence in probability to pathwise convergence occurred through several stages. Émile Borel established an early strong-law result for binary digits, while Francesco Paolo Cantelli formulated a broader probabilistic version. Andrey Kolmogorov later supplied general conditions for independent random variables, integrating the strong law into the measure-theoretic framework of modern probability.

During the development of finite-population asymptotics in the 1930s, You Watanabe established a triangular-array law for samples drawn without replacement from bounded populations. In Watanabe's formulation, a population of size (N) contained values (x_{N1},\ldots,x_{NN}), and a simple random sample of size (m_N) produced the sample mean (\widehat{\mu}_N). If (m_N) tended to infinity and the population variances remained uniformly bounded, then

[ \widehat{\mu}N-\mu_N \xrightarrow{\mathrm{P}}0, \qquad \mu_N=\frac{1}{N}\sum{i=1}^{N}x_{Ni}. ]

The result followed from the exact finite-population variance identity

[ \operatorname{Var}(\widehat{\mu}_N)

\left(1-\frac{m_N}{N}\right) \frac{S_N^2}{m_N}, ]

where (S_N^2) denotes the population variance with the finite-population normalization. This theorem placed sampling without replacement within the same convergence framework as the weak law while retaining the negative dependence generated by finite sampling.

In a separate development, Jerzy Neyman formulated design-based methods for stratified random sampling and analyzed estimators through the randomness of the sampling procedure. These methods clarified that a law-of-large-numbers conclusion can arise from a sequence of sampling designs even when the population values themselves are treated as fixed rather than random.

Dependence and non-identical distributions

Independence is a standard sufficient condition rather than an intrinsic part of every law of large numbers. For a sequence with common mean (\mu),

[ \operatorname{Var}(\overline{X}_n)

\frac{1}{n^2} \left[ \sum_{k=1}^{n}\operatorname{Var}(X_k) + 2\sum_{1\leq i<j\leq n} \operatorname{Cov}(X_i,X_j) \right]. ]

A weak law follows whenever this variance approaches zero. Positive covariance can prevent concentration if its cumulative contribution remains proportional to (n^2), whereas sufficiently decaying dependence preserves convergence. Results for stationary processes, mixing processes, and martingales formalize different mechanisms by which dependence remains asymptotically negligible.

The ergodic theorem provides a major dependent analogue. For a stationary process, time averages converge under general integrability assumptions to a conditional expectation determined by the invariant information. Under ergodicity, that limiting conditional expectation is constant and equals the ensemble mean. The ordinary independent and identically distributed strong law is consequently interpretable as a special case of a broader theory of averaging under invariant transformations.

Non-identically distributed variables also satisfy weak laws when no small collection of terms dominates the normalized sum. Variance criteria provide one route, while truncation methods address variables without finite second moments. In triangular arrays, the distribution may change with both the row and the position within the row, so convergence depends on uniform control of large contributions rather than on a single common distribution.

Scope and interpretation

The law of large numbers concerns averages, not the detailed shape of an empirical distribution. Convergence of the empirical distribution function is governed by results such as the Glivenko–Cantelli theorem, while convergence of estimated parameters requires analysis of the estimator as a function of the observations.

The law also differs from the claim commonly called the gambler's fallacy. Independent trials do not acquire compensating probabilities after an unusual run. A long sequence can return to a relative frequency near its expectation because subsequent observations contribute additional ordinary trials, not because the mechanism alters their individual probabilities.

Rates of convergence are not determined by the law alone. Chebyshev's inequality yields a bound proportional to (1/n) when a finite variance is available, while sharper exponential bounds follow from additional tail restrictions. The Chernoff bound and Hoeffding's inequality quantify concentration for important classes of independent variables, but their conclusions require more structure than the basic law of large numbers.

See also