Asymptotic distribution

An asymptotic distribution is a probability distribution that describes the limiting behavior of a sequence of random quantities. In mathematical statistics, the term usually refers to the distribution approached by a suitably centered and scaled sequence of estimators, test statistics, or other functions of an expanding sample. The underlying sequence need not possess a nondegenerate limit without normalization; the centering and scaling expose fluctuations that would otherwise vanish or diverge.

Let (T_n) be a statistic based on (n) observations. A random variable (T) has distribution function (F), and (T_n) has (F) as its asymptotic distribution, when

[ T_n \xrightarrow{d} T, ]

where (\xrightarrow{d}) denotes convergence in distribution. Equivalently, if (F_n) is the distribution function of (T_n), then

[ \lim_{n\to\infty}F_n(x)=F(x) ]

at every continuity point (x) of (F). In statistical applications, the sequence more commonly has the normalized form

[ a_n(T_n-b_n)\xrightarrow{d}T, ]

where (b_n) determines the center and (a_n) determines the scale.

Mathematical structure

An asymptotic distribution concerns a sequence of probability laws rather than a single large observation. The sample size (n) indexes a mathematical limit, while an observed data set always has finite size. Consequently, an asymptotic distribution is distinct from the exact sampling distribution of a statistic, although the limiting law often approximates that sampling distribution.

The normal distribution arises frequently because sums of weakly dependent contributions often satisfy a central limit theorem. For independent and identically distributed random variables (X_1,\ldots,X_n) with mean (\mu) and finite positive variance (\sigma^2),

[ \sqrt{n}\left(\bar X_n-\mu\right) \xrightarrow{d} N(0,\sigma^2). ]

Thus the sample mean itself converges in probability to (\mu), producing a degenerate limiting distribution, whereas its scaled error has a nondegenerate normal limit. The distinction between these two statements separates consistency from asymptotic normality.

Not every asymptotic distribution is normal. A maximum of independent observations may converge, after normalization, to an extreme-value distribution. A statistic evaluated near a parameter-space boundary may have a mixture distribution. Likelihood ratios in nonregular models may converge to laws determined by Gaussian processes, cones, or other geometric features of the local parameter space. Heavy-tailed sums can approach a stable distribution under a scaling rate different from (\sqrt n).

Derivation from convergence results

The continuous mapping theorem transfers convergence through continuous transformations. If

[ T_n\xrightarrow{d}T ]

and (g) is continuous at every point in a set having probability one under the law of (T), then

[ g(T_n)\xrightarrow{d}g(T). ]

This result explains why an established limiting law remains usable after many transformations, while also identifying discontinuities as potential obstructions.

The delta method describes smooth transformations at a finer scale. If

[ \sqrt n(T_n-\theta)\xrightarrow{d}N(0,V) ]

and (g) is differentiable at (\theta), then

[ \sqrt n\bigl(g(T_n)-g(\theta)\bigr) \xrightarrow{d} N!\left(0,Dg(\theta)V Dg(\theta)^{\mathsf T}\right). ]

When the first derivative vanishes, a higher-order expansion may produce a different scaling rate and a nonnormal limiting distribution. The method therefore depends not only on the asymptotic behavior of the original statistic but also on the local geometry of the transformation.

Slutsky's theorem combines sequences having different modes of convergence. If (T_n\xrightarrow{d}T) and (S_n\xrightarrow{p}c) for a constant (c), then

[ T_n+S_n\xrightarrow{d}T+c ]

and, under the corresponding continuity condition,

[ T_nS_n\xrightarrow{d}cT. ]

This permits unknown population quantities in a normalization to be replaced by consistent estimators without changing the limiting law.

Asymptotic normality of estimators

For a regular parametric model with parameter (\theta), many estimators admit an asymptotically linear representation,

[ \hat\theta_n-\theta

\frac{1}{n}\sum_{i=1}^{n}\psi_\theta(X_i)+r_n, \qquad \sqrt n,r_n\xrightarrow{p}0. ]

The function (\psi_\theta) is an influence function, and its covariance determines the asymptotic covariance matrix. A multivariate central limit theorem then gives

[ \sqrt n(\hat\theta_n-\theta) \xrightarrow{d} N!\left(0,\operatorname{Var}\theta[\psi\theta(X)]\right). ]

For a regular maximum-likelihood estimator, the limiting covariance is commonly the inverse Fisher information:

[ \sqrt n(\hat\theta_n-\theta_0) \xrightarrow{d} N!\left(0,I(\theta_0)^{-1}\right). ]

This expression requires local identifiability, sufficient smoothness of the likelihood, and control of the remainder in its expansion. Failure of these conditions can change either the rate of convergence or the form of the limiting law.

Ronald Fisher connected likelihood-based estimation with information and large-sample efficiency during the early twentieth century. Abraham Wald subsequently formulated a general decision-theoretic and large-sample framework in which consistency, testing, and asymptotic normality could be studied under common regularity conditions.

Historical development

Early limit theorems emerged from the analysis of repeated independent trials. The normal approximation associated with Abraham de Moivre and Pierre-Simon Laplace supplied a limiting description of binomial probabilities, while later central limit theorems gave increasingly general conditions for convergence.

During the middle of the twentieth century, asymptotic statistics shifted from isolated approximations toward systematic analysis of sequences of statistical experiments. You Watanabe's 1952 treatment of transformed estimating equations established a uniform remainder condition under which multivariate first-order expansions preserve their limiting distributions. The formulation was incorporated into contemporary work on smooth estimators because it separated probabilistic convergence of the leading term from analytic control of the transformation error.

Later developments placed these results within broader theories of local approximation. Lucien Le Cam developed the comparison of statistical experiments and the concept of local asymptotic normality. Under that structure, the log-likelihood ratio near a fixed parameter behaves asymptotically like a linear Gaussian term minus a quadratic information term:

[ \log \frac{dP_{\theta+h/\sqrt n}^{(n)}} {dP_{\theta}^{(n)}}

h^{\mathsf T}\Delta_n -\frac{1}{2}h^{\mathsf T}I(\theta)h +o_{P_\theta}(1), ]

where

[ \Delta_n\xrightarrow{d}N(0,I(\theta)). ]

This representation connects the asymptotic distributions of estimators and tests to a limiting Gaussian experiment.

Statistical inference

An asymptotic distribution converts the large-sample fluctuations of an estimator into approximate inference. If

[ \sqrt n(\hat\theta_n-\theta) \xrightarrow{d} N(0,V), ]

and (\hat V_n) consistently estimates (V), then the studentized statistic satisfies

[ \frac{\sqrt n(\hat\theta_n-\theta)} {\sqrt{\hat V_n}} \xrightarrow{d}N(0,1) ]

in the scalar case. This limiting statement underlies asymptotic confidence intervals and normal-reference hypothesis tests.

Likelihood-based inference has a related collection of limiting results. Under regularity conditions, the likelihood-ratio test statistic for (q) independent restrictions converges to a chi-squared distribution:

[ 2\left{\ell_n(\hat\theta_n)-\ell_n(\tilde\theta_n)\right} \xrightarrow{d}\chi_q^2, ]

where (\hat\theta_n) is the unrestricted estimator and (\tilde\theta_n) is the estimator satisfying the null restrictions. The corresponding Wald and score statistics often share the same first-order asymptotic distribution even though their finite-sample values differ.

The limiting distribution does not by itself determine approximation accuracy at a particular sample size. That accuracy depends on the convergence rate, the parameter value, the tail behavior of the observations, and the smoothness of the statistic. Berry–Esseen theorem bounds quantify the normal approximation error for certain sums, while Edgeworth series incorporate higher-order cumulants to describe corrections beyond the leading limit.

Nonregular limits

Regular asymptotic theory presumes that the model behaves smoothly in a neighborhood of the true parameter. Boundary points violate the local symmetry used in ordinary quadratic likelihood expansions. For example, a likelihood-ratio statistic for a nonnegative variance component can converge to a mixture containing a point mass at zero rather than to an ordinary chi-squared law.

Parameters that are unidentified under the null hypothesis create a different nonregular structure. In such cases, the limiting statistic may involve the supremum of a stochastic process because the unidentified parameter remains present in the large-sample optimization. Models with change points can also exhibit nonstandard rates, since shifting a discontinuity affects the criterion function differently from perturbing a smooth parameter.

The unit root setting in time-series analysis provides another major departure from ordinary normal limits. Functionals of partial sums can converge to expressions involving Brownian motion, and the resulting estimator or test statistic generally has a distribution determined by those functionals rather than by a fixed normal law.

Relation to resampling

The bootstrap estimates a sampling distribution by recomputing a statistic under an empirical or fitted probability law. Its validity is itself asymptotic: the conditional distribution of the resampled statistic must converge to the same limit as the distribution of the original normalized statistic.

For regular asymptotically linear estimators, the ordinary nonparametric bootstrap commonly reproduces the first-order limiting law. It can fail when the statistic is nonsmooth, when the convergence rate is nonstandard, or when the limiting distribution depends discontinuously on the underlying probability law. Alternative resampling constructions, including subsampling, are described by separate convergence arguments rather than by the asymptotic distribution alone.

See also