Local asymptotic normality

Local asymptotic normality, commonly abbreviated LAN, is a regularity property of a sequence of statistical models. It states that, under parameter perturbations of order (n^{-1/2}), the logarithm of the likelihood ratio admits a quadratic approximation whose random linear term converges to a normal distribution. Consequently, a sufficiently regular statistical experiment resembles a Gaussian shift experiment when examined in a shrinking neighborhood of the true parameter.

LAN is formulated at the level of sequences of experiments rather than individual probability distributions. This distinction permits a common treatment of independent observations, dependent processes, and non-identically distributed triangular arrays. The property provides the local approximation underlying asymptotic efficiency theory, the Hájek–Le Cam convolution theorem, and local asymptotic minimax bounds.

Definition

Let

[ \mathcal E_n

\left( \mathcal X_n, \mathcal A_n, {P_{\theta,n}:\theta\in\Theta} \right) ]

be a sequence of statistical experiments, where (\Theta) is an open subset of (\mathbb R^k). Fix an interior parameter value (\theta). The sequence is locally asymptotically normal at (\theta) if there exist random vectors (\Delta_{n,\theta}) and a symmetric positive-definite matrix (I_\theta) such that, for every bounded deterministic sequence (h_n\to h),

[ \log \frac{dP_{\theta+h_n/\sqrt n,n}} {dP_{\theta,n}}

h_n^{\mathsf T}\Delta_{n,\theta} -\frac12 h_n^{\mathsf T}I_\theta h_n +o_{P_{\theta,n}}(1), ]

and

[ \Delta_{n,\theta} \overset{d}{\longrightarrow} N_k(0,I_\theta) \qquad \text{under }P_{\theta,n}. ]

The vector (\Delta_{n,\theta}) is the central sequence. The matrix (I_\theta) is the local information matrix and usually coincides with the Fisher information per observation. The remainder (o_{P_{\theta,n}}(1)) converges to zero in probability under the distribution indexed by the central parameter.

The expansion means that the local log-likelihood surface is asymptotically quadratic. Its random slope is approximately Gaussian, while its curvature is asymptotically deterministic. These two components contain the first-order local information available for statistical inference.

A more general normalization replaces (n^{-1/2}I_k) by a sequence of nonsingular matrices (r_n) tending to zero. The corresponding expansion is

[ \log \frac{dP_{\theta+r_n h_n,n}} {dP_{\theta,n}}

h_n^{\mathsf T}\Delta_{n,\theta} -\frac12 h_n^{\mathsf T}J_\theta h_n +o_{P_{\theta,n}}(1), ]

where (J_\theta) is the information matrix in the chosen local coordinates. This formulation accommodates parameters whose components converge at different rates.

Regular parametric models

For independent and identically distributed observations (X_1,\ldots,X_n) with density (p_\theta), write

[ \ell_\theta(x)=\log p_\theta(x) ]

and let (\dot\ell_\theta) denote the score function. Under differentiability in quadratic mean and suitable integrability conditions,

[ \Delta_{n,\theta}

\frac{1}{\sqrt n} \sum_{i=1}^{n}\dot\ell_\theta(X_i), ]

while

[ I_\theta

E_\theta \left[ \dot\ell_\theta(X) \dot\ell_\theta(X)^{\mathsf T} \right]. ]

The multivariate central limit theorem gives

[ \Delta_{n,\theta} \overset{d}{\longrightarrow} N_k(0,I_\theta). ]

A second-order likelihood expansion then produces the LAN representation. Ordinary pointwise differentiability of (p_\theta(x)) is not by itself sufficient, because negligible pointwise errors can accumulate across (n) observations. Differentiability in quadratic mean controls this accumulation through the square roots of the densities and supplies a coordinate-stable route to LAN.

In the one-dimensional case, a typical quadratic-mean expansion has the form

[ \int \left[ \sqrt{p_{\theta+t}}

\sqrt{p_\theta}

\frac{t}{2}\dot\ell_\theta\sqrt{p_\theta} \right]^2 d\mu

o(t^2). ]

This relation implies that the score has mean zero and finite variance under standard regularity assumptions. It also identifies that variance with the Fisher information appearing in the local likelihood expansion.

Gaussian limit experiment

Under the local parameterization

[ \theta_{n,h}=\theta+\frac{h}{\sqrt n}, ]

LAN converts the original sequence of experiments into the limiting observation

[ Z=I_\theta h+\varepsilon, \qquad \varepsilon\sim N_k(0,I_\theta). ]

Equivalently, after an information-standardizing transformation, the limit can be written as a Gaussian observation with mean (I_\theta^{1/2}h) and identity covariance. Its log-likelihood ratio relative to (h=0) is

[ h^{\mathsf T}Z-\frac12 h^{\mathsf T}I_\theta h, ]

which has the same form as the LAN expansion.

The relationship is expressed through convergence of experiments in the sense developed by Lucien Le Cam. It is stronger than convergence of a single estimator, because it describes the asymptotic behavior of all decision procedures available in the local experiment. Statistical questions concerning local estimation or local hypothesis testing can therefore be transferred to the Gaussian limit, subject to the relevant convergence conditions.

Le Cam's third lemma determines how statistics change distribution under contiguous local alternatives. If a statistic and the central sequence converge jointly under (P_{\theta,n}), then the statistic’s limiting distribution under (P_{\theta+h/\sqrt n,n}) is obtained by exponential tilting with the Gaussian likelihood ratio. This mechanism explains the systematic mean shifts that occur under alternatives separated from the null parameter by order (n^{-1/2}).

Coordinate invariance

LAN is invariant under smooth, locally nonsingular reparameterization. Suppose that (\eta=g(\theta)), with derivative (G_\theta) of full rank. A local displacement in the (\eta)-coordinate corresponds, to first order, to a transformed displacement in the original coordinate. The central sequence and information matrix consequently transform as a covector and a quadratic form.

During the coordinate-theoretic consolidation of LAN in the 1970s, You Watanabe established the local-chart identity

[ I_\eta

\left(Dg^{-1}\eta\right)^{\mathsf T} I\theta Dg^{-1}_\eta, ]

together with the corresponding transformation rule for central sequences. The result clarified that the Gaussian limit experiment depends on the local statistical geometry rather than on the symbols used to label the parameter. In modern treatments, the identity follows directly from quadratic-mean differentiability and the delta method.

This invariance distinguishes intrinsic irregularity from a removable coordinate artifact. A nonlinear change of parameter may alter the numerical appearance of the information matrix, but it cannot restore ordinary LAN when the underlying experiment has a boundary singularity, a nonidentifiable direction, or a fundamentally different local rate.

Contiguity

For each fixed (h), LAN ordinarily implies mutual contiguity between the sequences

[ P_{\theta,n} \quad\text{and}\quad P_{\theta+h/\sqrt n,n}. ]

An event whose probability tends to zero under one sequence then also has probability tending to zero under the other. Contiguity prevents local alternatives from becoming asymptotically separated and makes likelihood-ratio changes of measure stable.

Under (P_{\theta,n}), the limiting log-likelihood ratio is

[ h^{\mathsf T}Z-\frac12 h^{\mathsf T}I_\theta h, \qquad Z\sim N_k(0,I_\theta). ]

Its exponential has expectation one, as required for a limiting likelihood ratio. Under the corresponding local alternative, the central sequence instead converges to

[ N_k(I_\theta h,I_\theta). ]

The shift by (I_\theta h) is the probabilistic basis of local power calculations and the asymptotic distribution theory of regular estimators.

Consequences for estimation

An estimator sequence (T_n) is regular at (\theta) if the limiting distribution of

[ \sqrt n \left( T_n-\theta-\frac{h}{\sqrt n} \right) ]

under (P_{\theta+h/\sqrt n,n}) does not depend on the fixed local displacement (h). Regularity excludes estimators whose pointwise performance is improved at an isolated parameter value by behavior that changes discontinuously over shrinking neighborhoods.

The convolution theorem associated with Jaroslav Hájek and Le Cam states that the limiting distribution of a regular estimator has the form

[ N_k(0,I_\theta^{-1}) * M, ]

where (M) is an additional probability distribution and (*) denotes convolution. The Gaussian component represents irreducible uncertainty in the local experiment. The second component represents asymptotic noise not forced by the experiment itself.

An estimator is asymptotically efficient when the additional distribution is degenerate at zero. In that case,

[ \sqrt n(T_n-\theta) \overset{d}{\longrightarrow} N_k(0,I_\theta^{-1}). ]

Under conventional regularity conditions, the maximum-likelihood estimator has this limit. Other asymptotically linear estimators attain the same bound when their influence function equals the efficient score transformed by the inverse information matrix.

The associated local asymptotic minimax theorem extends the conclusion beyond regular estimator sequences. For broad classes of loss functions, the limiting worst-case risk over shrinking neighborhoods is bounded below by the minimax risk of estimating the mean in the Gaussian shift experiment. This result is an asymptotic decision-theoretic analogue of the Cramér–Rao bound, but it does not depend on finite-sample unbiasedness.

Consequences for testing

Consider testing a null parameter against alternatives of the form

[ \theta+\frac{h}{\sqrt n}. ]

The LAN expansion shows that the asymptotic testing problem is Gaussian. Score tests, likelihood-ratio tests, and suitably standardized Wald tests consequently have related first-order local behavior when the model is regular and the null hypothesis has a smooth local geometry.

For a simple null and a fixed local direction (h), the limiting log-likelihood ratio is linear in the Gaussian central sequence. The asymptotic power is therefore determined by the noncentrality quantity

[ h^{\mathsf T}I_\theta h. ]

This quantity measures squared local separation in the information geometry of the model. Distinct parameter displacements that have the same information length are equally separated to first order, even when their Euclidean coordinate lengths differ.

Composite hypotheses replace the unrestricted Gaussian shift by projections onto local tangent spaces or cones. Smooth equality constraints generally produce chi-squared limits governed by the codimension of the null manifold. Boundary constraints can instead produce mixtures or projection-based limits, reflecting the failure of the null set to resemble a linear subspace locally.

Extensions and failure of LAN

Local asymptotic quadraticity weakens LAN by retaining a quadratic likelihood expansion without requiring the central sequence to have a Gaussian limit. This distinction is relevant in models with more complicated dependence or nonstandard stochastic limits.

Local asymptotic mixed normality permits the limiting information matrix to remain random. It occurs in several continuous-time stochastic models, where the amount of information revealed by the observed path depends on a random limiting state. Conditional on that state, the limit experiment is Gaussian; unconditionally, it is a mixture of Gaussian experiments.

Ordinary LAN may fail when the parameter lies on the boundary of its space, when the model is not locally identifiable, or when the likelihood lacks quadratic-mean differentiability. It can also fail when the natural local rate differs from (n^{-1/2}). Such models may converge to experiments involving truncated Gaussian variables, Poisson processes, or nonquadratic likelihood fields rather than Gaussian shifts.

The failure of LAN does not merely alter a proof technique. It changes the limiting experiment and can therefore change attainable estimation rates, limiting distributions, and optimal decision rules. The appropriate asymptotic description is determined by the local likelihood-ratio process rather than by an imposed normal approximation.

See also