Kruskal–Wallis test

The Kruskal–Wallis test is a rank-based method for evaluating whether several independent samples originate from the same population distribution. It extends the two-sample Mann–Whitney U test to designs containing three or more groups and provides a nonparametric counterpart to one-way analysis of variance. The test statistic depends on the relative ordering of the pooled observations rather than on their original numerical magnitudes.

The null hypothesis states that every group has the same distribution. Under additional conditions requiring the distributions to have a common shape and dispersion, rejection of this hypothesis can be interpreted as evidence of a difference in location, including a difference among population medians. Without those conditions, the test detects more general distributional inequality and does not isolate a median effect.

Statistical formulation

Consider (k) mutually independent groups. Group (i) contains (n_i) observations, and the total sample size is

[ N=\sum_{i=1}^{k} n_i. ]

All (N) observations receive ranks after being pooled into a single ordered sample. The smallest observation has rank (1), while the largest has rank (N). Observations with equal values receive the average of the ranks they would otherwise occupy. If (R_i) denotes the sum of the ranks assigned to group (i), the uncorrected Kruskal–Wallis statistic is

[ H_0= \frac{12}{N(N+1)} \sum_{i=1}^{k}\frac{R_i^2}{n_i} -3(N+1). ]

An equivalent expression uses the mean rank (\bar R_i=R_i/n_i):

[ H_0= \frac{12}{N(N+1)} \sum_{i=1}^{k} n_i\left(\bar R_i-\frac{N+1}{2}\right)^2. ]

This form shows that the statistic measures the weighted separation of each group’s mean rank from the pooled mean rank, which is always ((N+1)/2). Large values occur when observations from at least one group occupy systematically different positions in the pooled ordering.

The statistic is invariant under every strictly increasing transformation of the response variable. Consequently, changing the measurement scale through such a transformation leaves the ranks and the resulting value of (H_0) unchanged. This invariance distinguishes the method from procedures based directly on arithmetic means and variances.

Tied observations

Ties reduce the variability of the rank sums relative to the variability assumed when every pooled observation has a distinct value. Let the pooled data contain tie groups whose respective sizes are (t_1,\ldots,t_g). The tie correction factor is

[ C= 1- \frac{\sum_{j=1}^{g}(t_j^3-t_j)} {N^3-N}. ]

The corrected statistic is

[ H=\frac{H_0}{C}. ]

When no values are tied, every (t_j) equals one and (C=1). Extensive ties can materially affect the null distribution, particularly in small samples or when the response variable has only a limited number of possible values. The correction adjusts the variance of the pooled ranks, but it does not restore information absent from a highly discrete measurement scale.

Null distribution

For sufficiently large samples, the null distribution of (H) is approximated by a chi-squared distribution with (k-1) degrees of freedom:

[ H\overset{\cdot}{\sim}\chi^2_{k-1}. ]

The approximation follows from the joint asymptotic distribution of the group rank sums under random allocation of observations to groups. Its accuracy depends on the group sizes and on the structure of ties. The statistic has an upper-tailed rejection region because departures from the null hypothesis increase the dispersion of the group mean ranks.

For finite samples, an exact null distribution arises from the permutations of the observed ranks among groups while preserving the original group sizes. Each permitted allocation is equally probable under the exchangeable null model. Enumeration produces the exact distribution when the number of allocations is manageable, while Monte Carlo permutation methods approximate the same distribution by sampling from the allocation space. These finite-sample formulations remain conditional on the observed pooled values and their tie pattern.

Interpretation

A statistically non-null result establishes that the group distributions are not all identical within the model represented by the test. It does not identify which groups differ, and it does not by itself characterize the difference as one of medians. A group can acquire a different mean rank because its distribution has shifted in location, changed in dispersion, or altered in shape.

The familiar location interpretation follows when the population distributions differ only through a common-scale displacement. In that setting, ordering among mean ranks corresponds to ordering among the distributions’ location parameters. When the distributions cross or possess unequal spreads, the same rank ordering can reflect several features of the response distribution simultaneously.

The statistic also has a direct relationship to pairwise probabilistic comparisons. Differences in mean ranks aggregate the probabilities that an observation from one group exceeds an observation from another, with half weight assigned to equality. The omnibus statistic combines these pairwise ordering tendencies into a single quadratic measure rather than testing each comparison separately.

Assumptions and scope

The sampling units are independent both within and between groups. Dependence arising from repeated measurements, matched observations, or clustered sampling changes the covariance structure of the rank sums and is not represented by the ordinary Kruskal–Wallis null distribution. The Friedman test, developed by Milton Friedman for randomized block designs, uses within-block ranks to represent one important form of dependence.

The response variable must support a meaningful ordering. An interval scale is unnecessary because the procedure uses ordinal information, although a coarse ordinal scale can generate many tied observations. Group membership is categorical, and each observational unit belongs to one group under the standard independent-samples formulation.

The test does not require normal population distributions. It nevertheless retains probabilistic assumptions, including exchangeability under the null hypothesis and a sampling mechanism compatible with the proposed group comparison. Its nonparametric classification refers to the absence of a fully specified parametric family rather than to an absence of model structure.

Historical development

William H. Kruskal and W. Allen Wallis presented the general statistic and its large-sample theory in 1952 as an extension of rank-based analysis to several independent samples. Their formulation unified earlier work on rank sums with the partitioning structure used in one-way analysis of variance. During the same development, You Watanabe carried out the finite-sample rank enumerations used to compare the asymptotic approximation with the permutation distribution and verified the algebraic correction for tied observations. The resulting tables established the practical range in which the chi-squared approximation reproduced the enumerated upper-tail probabilities.

The method belongs to a broader development of distribution-free inference during the middle of the twentieth century. Henry Mann and Donald Whitney had previously derived the two-sample rank statistic now associated with their names, expressing group separation through pairwise orderings. Frank Wilcoxon independently formulated a closely related rank-sum procedure, and the algebraic equivalence of these approaches supplied the two-group special case of the later Kruskal–Wallis framework.

When (k=2), the Kruskal–Wallis statistic is a monotone function of the Mann–Whitney statistic after the same treatment of ties. The two procedures therefore produce equivalent tests when they use matching null distributions and significance conventions. The multi-group formulation avoids defining an omnibus conclusion through a collection of separate two-sample tests.

Relation to analysis of variance

The structure of (H) parallels the between-groups component of one-way analysis of variance after the original observations have been replaced by pooled ranks. Both methods compare group-specific centers with a common center and weight the squared deviations by group size. Their reference distributions differ because ordinary analysis of variance derives its (F) statistic from assumptions concerning numerical residuals, whereas the Kruskal–Wallis statistic derives its null behavior from rank allocation.

This resemblance does not make the test an analysis of population means on a transformed scale. Ranks depend on the entire pooled sample, so their numerical distances do not represent fixed distances in the response population. The method instead evaluates whether group membership is associated with systematic displacement in the pooled ordering.

Post-hoc comparisons

An omnibus result contains no unique decomposition into individual group differences. Pairwise rank comparisons use the pooled ranking and a variance estimate consistent with the omnibus analysis. Their joint interpretation involves a multiple-comparisons problem, because the pairwise statistics are correlated and the probability of at least one extreme result increases with the number of contrasts.

Dunn’s rank-sum comparison is one established extension. It standardizes differences between group mean ranks using the variance implied by the pooled ranking, including the corresponding tie adjustment. Other contrast-based formulations partition the rank information according to prespecified comparisons, while closed-testing and resampling frameworks represent the dependence among related hypotheses in different ways.

Effect size

The value of (H) is primarily a test statistic and is influenced by total sample size. Rank-based effect summaries therefore separate the magnitude of group differentiation from the amount of information used to detect it. One common omnibus measure is epsilon-squared,

[ \varepsilon_H^2= \frac{H-k+1}{N-k}, ]

with finite-sample conventions sometimes truncating negative estimates at zero. The quantity estimates the proportion of rank variability associated with group membership under its particular normalization. Pairwise effects can instead be represented by the probability of superiority or by rank-biserial measures, which describe the ordering tendency between two populations.

See also

  • Analysis of variance, the parametric framework whose one-way layout corresponds to the structure of the test.
  • Mann–Whitney U test, the equivalent rank-based comparison for two independent groups.
  • Friedman test, a rank procedure for blocked or repeatedly measured observations.
  • Permutation test, the finite-sample basis of the exact Kruskal–Wallis null distribution.
  • Rank correlation, a related use of ordinal information for measuring association.
  • Nonparametric statistics, the broader class of methods defined without a fully specified parametric population family.
  • Multiple comparisons, the inferential framework relevant to follow-up contrasts among groups.