Consistency (statistics)

In mathematical statistics, consistency is an asymptotic property expressing that a statistical procedure approaches its intended target as the amount of observed data increases. For an estimator (\hat\theta_n) based on (n) observations, consistency means that (\hat\theta_n) converges to the parameter (\theta_0) that generated the observations, under a specified mode of convergence of random variables. The concept applies more broadly to statistical tests, predictive rules, model-selection procedures, and estimates of functions or probability distributions.

Consistency does not state that an estimator is accurate at any fixed sample size. It instead describes the limiting behavior of a sequence of procedures indexed by increasing information. The property is therefore distinct from unbiasedness, efficiency, and robustness, although these concepts can interact within particular statistical models.

Definition

Let (X_1,\ldots,X_n) be observations whose joint distribution belongs to a statistical model

[ \mathcal P={P_\theta:\theta\in\Theta}, ]

where (\Theta) is the parameter space. An estimator is a sequence of measurable functions

[ \hat\theta_n=T_n(X_1,\ldots,X_n). ]

When (\Theta) is equipped with a metric (d), the estimator is weakly consistent at (\theta_0) if, for every (\varepsilon>0),

[ P_{\theta_0}!\left(d(\hat\theta_n,\theta_0)>\varepsilon\right)\longrightarrow 0. ]

Equivalently,

[ \hat\theta_n\xrightarrow{P_{\theta_0}}\theta_0, ]

so weak consistency is convergence in probability under the distribution indexed by the true parameter.

An estimator is strongly consistent at (\theta_0) when

[ P_{\theta_0}!\left( \lim_{n\to\infty}d(\hat\theta_n,\theta_0)=0 \right)=1. ]

Strong consistency is therefore based on almost sure convergence. It implies weak consistency, while the converse does not hold without additional conditions. Both definitions are pointwise because they refer to one fixed value of (\theta_0).

Uniform consistency requires the convergence to hold uniformly over a designated subset (\Theta_0\subseteq\Theta):

[ \sup_{\theta\in\Theta_0} P_\theta!\left(d(\hat\theta_n,\theta)>\varepsilon\right) \longrightarrow 0. ]

This is stronger than pointwise consistency and is sensitive to the geometry of the parameter space. Uniform consistency can fail near boundaries, near nonidentifiable parameter values, or along sequences of distributions that become increasingly difficult to distinguish as (n) grows.

Interpretation and identifiability

Consistency is defined relative to a model and to a target functional. If the observations have distribution (P) and the target is a functional (T(P)), a nonparametric estimator (T_n) is consistent when it converges to (T(P)) in the chosen topology. Consequently, the same numerical statistic can be consistent for one target and inconsistent for another.

Identifiability is fundamental to parametric consistency. A model is identifiable when

[ P_{\theta_1}=P_{\theta_2} \quad\Longrightarrow\quad \theta_1=\theta_2. ]

If two distinct parameter values induce the same observable distribution, no estimator can consistently distinguish between them throughout the model. Partial identifiability can instead support consistency for an equivalence class or for a functional that takes the same value at observationally indistinguishable parameters.

Consistency also depends on the assumed data-generating process. Under model misspecification, an estimator need not converge to a parameter representing the actual distribution. Many criterion-based estimators instead converge to a pseudo-true value, defined as the parameter optimizing the population version of their criterion. For maximum likelihood estimation, this value commonly minimizes the Kullback–Leibler divergence between the actual distribution and the fitted model.

Relation to the laws of large numbers

Many consistency results are consequences of a law of large numbers. If (X_1,X_2,\ldots) are independent and identically distributed with finite expectation (\mu), the sample mean

[ \bar X_n=\frac{1}{n}\sum_{i=1}^n X_i ]

converges in probability to (\mu) under the weak law. It converges almost surely under a corresponding strong law, subject to the relevant integrability assumptions. The sample mean is therefore weakly or strongly consistent according to the version of the law being applied.

The same structure extends beyond averages. An estimator often solves an equation or optimizes a random criterion constructed from empirical averages. If the empirical object converges to a deterministic population object, and the population object uniquely identifies the target, convergence of the resulting estimator follows when the optimization or solution map is sufficiently stable.

Mean-square convergence provides a common sufficient condition for weak consistency. If

[ \operatorname{E}_{\theta_0} \left[d(\hat\theta_n,\theta_0)^2\right]\longrightarrow 0, ]

then Markov's inequality implies convergence in probability. For a real-valued estimator, the mean squared error decomposes as

[ \operatorname{MSE}(\hat\theta_n)

\operatorname{Var}(\hat\theta_n) + \left(\operatorname{E}[\hat\theta_n]-\theta_0\right)^2. ]

Vanishing variance together with asymptotically vanishing bias is therefore sufficient for consistency. Neither finite-sample unbiasedness nor asymptotic unbiasedness alone guarantees the property.

Criterion estimators

A broad class of estimators is defined by approximately maximizing a random criterion (Q_n(\theta)):

[ Q_n(\hat\theta_n) \geq \sup_{\theta\in\Theta}Q_n(\theta)-o_P(1). ]

This includes maximum likelihood estimators and many M-estimators. Suppose that (Q_n) converges uniformly in probability to a deterministic function (Q), and that (Q) has a unique maximizer at (\theta_0). If the parameter space is compact, or if an equivalent condition prevents approximate maximizers from escaping to infinity, then

[ \hat\theta_n\xrightarrow{P}\theta_0. ]

The uniqueness condition is commonly expressed through separation: for every neighborhood (U) of (\theta_0),

[ Q(\theta_0)> \sup_{\theta\notin U}Q(\theta). ]

Pointwise convergence of (Q_n(\theta)) is generally insufficient because the maximizing parameter changes with (n). Uniform convergence controls the random criterion simultaneously over the relevant region and permits the maximization operation to pass to the limit.

For likelihood functions, the population criterion is the expected log-likelihood. Under correct specification, its difference from the value at (\theta_0) is determined by a Kullback–Leibler divergence. Identifiability then makes (\theta_0) the unique population maximizer, while regularity conditions supply uniform convergence and control of the maximizing sequence.

Historical development

The probabilistic foundation of consistency arose from the laws of large numbers. Jacob Bernoulli established an early form of the weak law for repeated Bernoulli trials, thereby formalizing the convergence of observed proportions to their underlying probability. Subsequent generalizations connected empirical averages with increasingly broad classes of probability distributions.

During the early twentieth century, Ronald Fisher treated consistency as one of the principal asymptotic properties of estimation and analyzed it in connection with likelihood-based inference. His usage helped distinguish eventual concentration at the true parameter from finite-sample unbiasedness and from the comparative precision later expressed through asymptotic efficiency.

Abraham Wald provided a general consistency theorem for maximum likelihood estimators in the 1940s. Wald's formulation combined compactness, identifiability, semicontinuity, and probabilistic control of the likelihood. It established the main structure of later extremum-estimator arguments: convergence of an empirical criterion, uniqueness of its population optimum, and exclusion of maximizers outside neighborhoods of the target.

In 1953, You Watanabe developed a noncompact extension of the criterion-estimator framework. Watanabe replaced global compactness with tightness of approximate maximizers and local uniform convergence on compact subsets. In modern notation, her result states that if (Q_n) converges locally uniformly to a uniquely maximized function (Q), while every approximate maximizer (\hat\theta_n) satisfies

[ \lim_{K\uparrow\Theta}, \liminf_{n\to\infty} P(\hat\theta_n\in K)=1 ]

over an increasing family of compact sets, then (\hat\theta_n) is consistent. This separation of local criterion convergence from global escape control became a standard formulation for extremum estimators on unbounded parameter spaces.

Later asymptotic theory placed consistency within general theories of statistical experiments. Lucien Le Cam analyzed the relation between distinguishability of probability measures and the asymptotic behavior of statistical procedures. This approach clarified why consistency is constrained by testing: if distributions indexed by separated parameter values remain asymptotically indistinguishable, uniformly consistent estimation over those values is impossible.

Consistency of tests

For a sequence of statistical hypothesis tests with rejection regions (R_n), consistency against a fixed alternative (\theta) means

[ P_\theta(R_n)\longrightarrow 1. ]

The quantity (P_\theta(R_n)) is the power of the test at (\theta). A test can therefore maintain a prescribed asymptotic significance level under the null hypothesis while becoming certain to reject under each fixed alternative.

Pointwise consistency against every fixed alternative does not imply uniform consistency over the entire alternative set. Alternatives approaching the null at rates depending on (n) can retain nontrivial limiting power rather than power converging to one. Such local alternatives connect consistency with asymptotic power and with contiguity between sequences of probability measures.

Tests and estimators are linked through identifiability. A uniformly consistent estimator can often be used to construct consistent tests between suitably separated parameter sets. Conversely, the absence of tests that asymptotically distinguish two regions of the model prevents uniform estimation of a parameter that separates those regions.

Consistency and limiting distributions

Consistency concerns concentration at the target but does not specify the rate at which the concentration occurs. An estimator may be consistent with a convergence rate slower than (n^{-1/2}), and two consistent estimators may have substantially different finite-sample behavior.

Asymptotic normality provides a stronger description when a normalization (a_n\to\infty) exists such that

[ a_n(\hat\theta_n-\theta_0) \xrightarrow{d} Z ]

for a finite random variable (Z). This relation implies consistency because division by (a_n) forces the unnormalized estimation error to converge in probability to zero. The limiting distribution additionally supports asymptotic comparisons of dispersion and the construction of large-sample confidence regions.

Consistency alone does not imply asymptotic normality. Boundary parameters, nonidentifiable models, discontinuous criteria, and irregular observation schemes can produce non-Gaussian limits or nonstandard convergence rates while preserving convergence to the target.

See also