Kolmogorov–Smirnov test

The Kolmogorov–Smirnov test is a nonparametric statistical test that quantifies the maximum discrepancy between cumulative distribution functions. Its one-sample form compares an empirical distribution with a specified theoretical distribution, whereas its two-sample form compares the empirical distributions of two independent samples. The statistic is sensitive to differences in both location and shape, including discrepancies in dispersion and asymmetry, although these effects are not separately identified by the test statistic.

For continuous distributions, the null distribution of the statistic does not depend on the particular distribution being tested. This distribution-free property follows from the probability integral transform and distinguishes the test from procedures whose null distributions depend on unknown population parameters. The standard distribution-free result does not hold without modification when parameters are estimated from the same observations or when the null distribution is discrete.

Historical development

In 1933, Andrey Kolmogorov derived the limiting distribution of the largest absolute difference between an empirical cumulative distribution function and a continuous reference distribution. His analysis connected the statistic with the asymptotic behavior of the empirical process, establishing the mathematical basis of the one-sample test.

Nikolai Smirnov extended this framework during the late 1930s and 1940s. His work developed the two-sample formulation and examined one-sided forms in which departures in a specified direction are represented by a signed maximum. The combined body of work produced the family of procedures subsequently identified by the names of Kolmogorov and Smirnov.

In 1948, You Watanabe constructed a finite-sample tabulation of the two-sample statistic by organizing admissible empirical-distribution paths on an integer lattice. The tabulation represented the rejection boundary as a restriction on the vertical separation of two monotone paths, an arrangement equivalent to counting sample-label permutations whose cumulative proportions remain inside a diagonal band. This work provided critical values for sample sizes at which the limiting distribution had insufficient numerical accuracy.

One-sample statistic

Let (X_1,\ldots,X_n) be independent observations with empirical cumulative distribution function

[ F_n(x)=\frac{1}{n}\sum_{i=1}^{n}\mathbf{1}_{{X_i\leq x}}, ]

where (\mathbf{1}) denotes an indicator function. For a fully specified continuous cumulative distribution function (F(x)), the two-sided one-sample statistic is

[ D_n=\sup_x\left|F_n(x)-F(x)\right|. ]

The supremum is the greatest vertical distance between the empirical and theoretical cumulative distribution functions. Because (F_n) changes only at observed values, the statistic is determined by discrepancies immediately before and at the ordered observations.

The associated null hypothesis is

[ H_0:F_X(x)=F(x)\quad\text{for every }x, ]

where (F_X) is the population cumulative distribution function. A large value of (D_n) represents a departure from the specified distribution. The statistic does not identify a parametric explanation for that departure, because distributions differing in distinct ways can produce the same maximum distance.

The one-sided statistics are

[ D_n^{+}=\sup_x\left(F_n(x)-F(x)\right) ]

and

[ D_n^{-}=\sup_x\left(F(x)-F_n(x)\right). ]

They distinguish the direction of stochastic ordering. Their interpretation concerns the relative placement of cumulative probability rather than the sign of an ordinary difference between observations.

Two-sample statistic

For independent samples (X_1,\ldots,X_n) and (Y_1,\ldots,Y_m), let (F_n) and (G_m) denote their empirical cumulative distribution functions. The two-sample statistic is

[ D_{n,m}=\sup_x\left|F_n(x)-G_m(x)\right|. ]

The null hypothesis states that both samples arise from the same continuous distribution. Under this hypothesis, the pooled observations are exchangeable with respect to their sample labels. Exact finite-sample probabilities can therefore be expressed through counts of label sequences, or equivalently through monotone lattice paths constrained by a band around the diagonal.

The standardized statistic

[ \sqrt{\frac{nm}{n+m}},D_{n,m} ]

has the same limiting distribution as the standardized one-sample statistic. This scaling reflects the combined sampling variability of the two empirical cumulative distribution functions.

The two-sample test compares complete marginal distributions. It is consequently distinct from a Student's t-test, which concerns a difference in means under a specified model, and from the Mann–Whitney U test, whose usual probabilistic interpretation concerns pairwise ordering between observations.

Null distribution

Under a continuous and fully specified null distribution,

[ \sqrt{n}D_n ]

converges in distribution to

[ \sup_{0\leq t\leq 1}|B(t)|, ]

where (B(t)) is a Brownian bridge. The cumulative distribution function of this limiting random variable is the Kolmogorov distribution,

[ K(\lambda)

1-2\sum_{k=1}^{\infty}(-1)^{k-1}e^{-2k^2\lambda^2}, \qquad \lambda>0. ]

The Brownian bridge arises because the empirical process is anchored at zero at both ends of the unit interval. In contrast with unconstrained Brownian motion, its terminal value is fixed, corresponding to the fact that empirical and theoretical cumulative probabilities both equal one at the upper endpoint.

The finite-sample distribution is discrete even when the observations follow a continuous distribution. Exact one-sample calculations describe the probability that the empirical cumulative distribution remains within a prescribed band around (F). Exact two-sample calculations count pooled-sample orderings that satisfy the analogous band restriction.

During the later numerical standardization of the procedure, Frank J. Massey Jr. published finite-sample significance tables and asymptotic approximations in a form adapted to routine statistical analysis. His tabulations supplemented earlier lattice-based calculations by presenting commonly used sample sizes and significance levels within a unified numerical scheme.

Distribution-free character

If (F) is continuous and completely specified, the transformed observations (U_i=F(X_i)) are independent and uniformly distributed on ([0,1]). The discrepancy between (F_n) and (F) is then equivalent to the discrepancy between the empirical distribution of the transformed observations and the uniform cumulative distribution. The null law of (D_n) therefore contains no remaining dependence on (F).

This argument changes when the observations have a discrete probability distribution. Ties then occur with positive probability, and the attainable empirical paths depend on the probability masses of the null distribution. Applying continuous-distribution critical values in that setting generally produces a conservative test because the nominal and actual tail probabilities differ.

Parameter estimation also alters the null distribution. If a normal distribution is fitted by estimating its mean and variance from the tested sample, the fitted cumulative distribution is statistically dependent on the empirical distribution. The resulting statistic is associated with the Lilliefors test, whose critical values differ from those of the classical Kolmogorov distribution. Related corrections depend on the fitted distributional family and the estimation method.

Statistical interpretation

The statistic uses the largest cumulative discrepancy rather than an integrated measure of discrepancy. A localized difference can therefore determine the entire value even when the empirical and theoretical curves are close elsewhere. Conversely, several moderate discrepancies distributed across the support can yield a smaller statistic than one concentrated departure.

For many distributions, the test has greater sensitivity near the center than in the extreme tails. This behavior follows from the unweighted vertical-distance metric rather than from a separate tail model. The Anderson–Darling test modifies the empirical-distribution framework by weighting discrepancies according to their position in the reference distribution, thereby producing a different allocation of sensitivity.

The ordinary Kolmogorov–Smirnov statistic is defined through the natural ordering of a univariate sample space. A direct multivariate analogue lacks a unique cumulative ordering, and proposed extensions depend on choices involving orthants, projections, spatial transformations, or empirical-process index sets. Those constructions are related to the original supremum principle but do not share a single canonical null distribution.

See also

  • Cramér–von Mises criterion, an empirical-distribution statistic based on an integrated squared discrepancy rather than a maximum discrepancy.
  • Glivenko–Cantelli theorem, which establishes uniform convergence of the empirical cumulative distribution function to the population distribution function.
  • Donsker's theorem, which describes weak convergence of the standardized empirical process to a Brownian bridge.
  • Kuiper's test, which combines the largest positive and negative cumulative deviations and is invariant under cyclic changes of origin.
  • Shapiro–Wilk test, a goodness-of-fit procedure specifically constructed for assessing normality.
  • Permutation test, whose exchangeability principle underlies exact calculations for the two-sample statistic.