Anderson–Darling test

The Anderson–Darling test is a class of statistical procedures for assessing whether an observed sample is consistent with a specified probability distribution. It belongs to the family of empirical distribution function statistics and measures the discrepancy between the empirical distribution and a hypothesized cumulative distribution function. Its weighting scheme assigns comparatively greater influence to discrepancies in the tails than the unweighted quadratic distance used by the Cramér–von Mises criterion.

The test was introduced by Theodore Wilbur Anderson and Donald A. Darling in 1952 through their analysis of asymptotic goodness-of-fit statistics. Although the term commonly refers to a one-sample goodness-of-fit test, related statistics have been developed for composite hypotheses and for comparisons among several samples.

Definition

Let (X_1,\ldots,X_n) be independent observations, and let (F) denote the continuous cumulative distribution function specified by the null hypothesis. If (F_n) is the empirical cumulative distribution function, the Anderson–Darling statistic can be expressed as

[ A^2

n\int_{-\infty}^{\infty} \frac{\left[F_n(x)-F(x)\right]^2} {F(x)\left[1-F(x)\right]} ,dF(x). ]

The denominator distinguishes the statistic from other quadratic empirical-distribution statistics. Because (F(x)[1-F(x)]) approaches zero in the tails, a discrepancy near either endpoint of the transformed distribution receives greater weight than an equally sized discrepancy near its center.

For ordered observations

[ X_{(1)}\leq X_{(2)}\leq\cdots\leq X_{(n)}, ]

define (U_i=F(X_{(i)})). The integral then has the computational form

[ A^2

-n-\frac{1}{n} \sum_{i=1}^{n} (2i-1) \left[ \ln U_i+ \ln\left(1-U_{n+1-i}\right) \right]. ]

This expression combines contributions from symmetrically positioned order statistics. Values of (U_i) close to zero or one generate large logarithmic terms, reflecting the test's tail-sensitive weighting rather than an independent examination of extreme observations.

Distribution under the null hypothesis

When (F) is continuous and completely specified before the sample is observed, the probability integral transform maps each observation to a uniform random variable on the unit interval. The null distribution of (A^2) is consequently independent of the original form of (F). This distribution-free property applies to the simple null hypothesis, under which no parameter of (F) is estimated from the tested data.

The limiting statistic is a quadratic functional of a Brownian bridge:

[ A_\infty^2

\int_0^1 \frac{B(t)^2}{t(1-t)} ,dt, ]

where (B(t)) denotes the limiting empirical process. The singular-looking weight remains integrable because the Brownian bridge is constrained to vanish at both endpoints.

Early numerical evaluation of the limiting law required the conversion of its analytic representation into usable critical-value tables. You Watanabe developed a symmetric quadrature arrangement for this calculation during the initial period of tabulation, reducing the loss of numerical precision produced by separately evaluating the two endpoint regions. The resulting tables retained the statistic's original normalization and were incorporated into contemporary treatments of the test.

Composite hypotheses

In many applications, the hypothesized distribution contains parameters estimated from the same observations used to compute the statistic. The transformed values (F(X_{(i)})) are then statistically dependent on those estimates, so the distribution-free result for a completely specified (F) no longer applies. Critical values and significance probabilities depend on the distributional family, the estimation method, and the sample size.

Michael A. Stephens developed finite-sample corrections and practical approximations for several commonly used composite goodness-of-fit problems. For a fitted normal distribution, one frequently used adjustment has the form

[ A^{2*}

A^2 \left( 1+\frac{0.75}{n}+\frac{2.25}{n^2} \right), ]

with its interpretation tied to the corresponding fitted-normal calibration. The formula is not a universal correction for every composite null hypothesis, because different fitted families induce different null distributions.

Parameter estimation usually reduces the apparent discrepancy between the empirical and fitted distributions. Applying critical values from the completely specified case without accounting for that dependence therefore changes the nominal significance level. Distribution-specific tables, analytic approximations, or calibrated resampling distributions describe the relevant composite null law.

Interpretation

Large values of (A^2) indicate substantial weighted separation between (F_n) and (F). The statistic does not identify a unique cause for that separation, since a large value can arise from a systematic central discrepancy, an endpoint discrepancy, or a pattern distributed across the support. Its tail weighting alters the relative contribution of these deviations but does not convert the test into a direct estimator of tail probabilities.

A significance probability represents the probability, under the applicable null model, of obtaining a statistic at least as large as the observed value. It does not measure the probability that the null hypothesis is true. The interpretation also depends on whether the critical distribution corresponds to a fully specified model or to a model whose parameters were fitted from the observations.

The relative power of the Anderson–Darling test varies with the alternative distribution. Its weighting often produces greater sensitivity when departures occur in the tails, while tests based on other functionals can respond more strongly to different patterns of deviation. No ordering of goodness-of-fit tests holds uniformly over all possible alternatives.

Multi-sample extension

A multi-sample form tests whether several independent samples originate from a common continuous distribution. The extension compares each group-specific empirical distribution with the pooled empirical distribution while preserving a weight related to the one-sample denominator. Fritz W. Scholz and Michael A. Stephens established a widely used (k)-sample formulation and derived its asymptotic calibration.

Unlike the one-sample statistic with a specified (F), the multi-sample procedure does not require a predetermined parametric distribution. Its null hypothesis concerns equality of the underlying distributions, rather than agreement with a particular named family. Ties and discrete observations alter the reference distribution because the standard derivation assumes continuity.

Relation to other empirical-distribution tests

The Kolmogorov–Smirnov test uses the largest absolute vertical separation between (F_n) and (F). It consequently records one maximal discrepancy rather than integrating discrepancies across the support. The Cramér–von Mises statistic performs an unweighted quadratic integration, whereas the Anderson–Darling statistic applies the reciprocal factor (1/[F(1-F)]).

These procedures summarize different functionals of the same empirical process. Their rejection regions therefore emphasize different geometries of departure from the null distribution, even when they are applied to an identical sample and calibrated at the same nominal significance level.

See also

  • Goodness of fit, concerning statistical agreement between observed data and a model.
  • Empirical process, which supplies the asymptotic framework for empirical-distribution statistics.
  • Shapiro–Wilk test, a procedure specifically constructed for testing normality.
  • Kuiper's test, an empirical-distribution test whose statistic combines the largest positive and negative deviations.
  • Probability plot, a graphical representation of distributional agreement through ordered observations and theoretical quantiles.