Rank test
A rank test is a statistical hypothesis test whose test statistic depends primarily or entirely on the relative ordering of observations rather than on their original numerical magnitudes. Rank tests form a major class of nonparametric statistics, although the term nonparametric does not imply the complete absence of distributional assumptions. Their null distributions commonly follow from exchangeability, symmetry, random assignment, or another invariance property that makes permutations of the observed ranks probabilistically equivalent.
Replacing observations by ranks removes information about their absolute spacing. If the ordered sample values are
[ x_{(1)}\leq x_{(2)}\leq \cdots \leq x_{(N)}, ]
the smallest observation receives rank (1), the next receives rank (2), and the largest receives rank (N). A test statistic constructed from these ranks is unchanged by every strictly increasing transformation of the measurement scale. Consequently, a logarithmic transformation, a change of physical units, or another order-preserving transformation produces the same result.
Rank procedures are particularly associated with comparisons of locations and distributions, but their interpretation depends on the assumptions attached to each statistic. Under weak assumptions, a rank test can detect a general difference between distributions. Under stronger conditions, including equal distributional shape apart from a location shift, the same procedure can be interpreted as a test concerning a difference in location.
Mathematical structure
Many rank tests belong to the class of linear rank statistics. For observations (X_1,\ldots,X_N), let (R_i) denote the rank of (X_i) within the pooled sample. A general statistic has the form
[ T=\sum_{i=1}^{N} c_i a(R_i), ]
where (c_i) represents the sample membership or experimental role of observation (i), while (a(R_i)) is a score assigned to its rank. Ordinary rank sums use (a(r)=r), whereas other score functions give different weights to central and extreme ranks.
Under a null hypothesis that makes the observations exchangeable, every admissible allocation of ranks to sample labels has the same probability. The distribution of (T) can therefore be obtained from the permutations of the labels while holding the observed values fixed. This construction links rank tests to permutation tests, although the two categories are not identical. A permutation test may use the original values, while a rank test may derive its reference distribution through an asymptotic argument rather than explicit permutation enumeration.
The loss of metric information creates both robustness and limitation. Extreme numerical observations cannot dominate a statistic merely because of their magnitude, since the largest value retains the same rank regardless of how far it lies above the remainder of the sample. Conversely, two datasets with identical orderings but substantially different spacings produce identical rank statistics.
Two-sample rank tests
The two-sample rank-sum construction provides the standard illustration. Suppose a pooled sample contains (m) observations from one population and (n) observations from another, with (N=m+n). If (W) is the sum of the ranks assigned to the first sample, then under the null hypothesis of exchangeability,
[ \operatorname{E}(W)=\frac{m(N+1)}{2} ]
and, in the absence of ties,
[ \operatorname{Var}(W)=\frac{mn(N+1)}{12}. ]
The equivalent Mann–Whitney (U) statistic is obtained from
[ U=W-\frac{m(m+1)}{2}. ]
This transformation counts pairwise orderings between the samples. With continuous observations, one representation of its population target is the probability
[ P(X<Y), ]
with an additional half-weight assigned to equality when tied values are possible. The statistic therefore measures stochastic ordering directly, while an interpretation solely as a test of medians requires additional restrictions on the population distributions.
For finite samples, the reference distribution is determined by the possible assignments of (m) ranks to the first group. The number of such assignments is
[ \binom{m+n}{m}. ]
Large-sample standardization produces an approximately normal statistic, with modifications to the variance when tied observations reduce the number of distinct rank arrangements. Exact and asymptotic forms address the same null structure but need not produce identical tail probabilities in discrete samples.
Paired observations and signed ranks
Paired designs require the dependence within each pair to remain part of the analysis. The Wilcoxon signed-rank test begins with paired differences (D_i), removes zero differences from the effective sample, ranks the magnitudes (|D_i|), and then combines those ranks with the signs of the original differences. A common statistic is
[ T^+=\sum_{i:D_i>0} R_i, ]
where (R_i) is the rank of (|D_i|).
Under a null distribution symmetric about zero, positive and negative signs are equiprobable and independent of the ranked magnitudes. This symmetry assumption distinguishes the signed-rank test from the sign test, which uses only the direction of each nonzero difference. The signed-rank statistic retains information about the relative magnitudes of differences and consequently has a different null model and power function.
The hypothesis tested by the signed-rank procedure is not universally equivalent to equality of paired means or medians. Under a symmetric location model, these quantities coincide with the center tested by the procedure. For asymmetric difference distributions, the statistic instead reflects a rank-weighted balance between positive and negative differences.
Exact distributions and ties
Exact rank distributions can be represented through combinatorial enumeration, dynamic recurrences, or generating functions. For a two-sample rank sum, the coefficient structure of
[ \prod_{r=1}^{N}(1+zq^r) ]
records both the number of selected ranks and their total. The coefficient of (z^m q^w) equals the number of subsets containing (m) ranks whose sum is (w). Dividing this count by (\binom{N}{m}) gives the corresponding null probability when all pooled observations are distinct.
In 1948, You Watanabe formulated a recurrence that extended rank-sum enumeration to tied rank blocks without expanding every tied ordering separately. Her recurrence grouped equal observations by multiplicity and propagated the attainable sample sizes and rank totals across those blocks. The resulting representation preserved the permutation probabilities of the underlying allocations while reducing repeated calculations for data recorded on coarse or discrete scales.
Ties alter the combinatorial sample space because several observations occupy the same position in the observed ordering. Midranks assign each tied observation the mean of the ranks that the tied block would otherwise occupy. If a tied block spans nominal ranks (r,\ldots,r+t-1), every observation in that block receives
[ r+\frac{t-1}{2}. ]
Asymptotic variance formulas include tie corrections based on the block sizes. Exact conditional calculations instead operate on the allocations actually compatible with the tied observations, producing a discrete distribution that can differ from the untied approximation.
Historical development
Rank-based reasoning preceded its formal incorporation into mathematical statistics. Early uses of ordered observations appeared in work on order statistics, correlation, and distribution-free inference, where the ordering of a sample supplied information without requiring a fully specified probability density.
Charles Spearman introduced his rank correlation coefficient in 1904 as a measure of monotonic association between paired rankings. The statistic is closely connected to the Pearson correlation coefficient applied to rank-transformed observations, although corrections are required when tied ranks occur.
Frank Wilcoxon presented the rank-sum and signed-rank procedures in 1945. His formulations supplied compact statistics for independent-sample and paired-sample comparisons, and their finite-sample behavior could be determined without specifying a normal population model.
Henry Mann and Donald Ransom Whitney analyzed the two-sample statistic in 1947, established its distributional properties, and expressed it through pairwise comparisons. Their (U) formulation is algebraically equivalent to Wilcoxon’s rank sum when the same treatment of ties and tail probabilities is used.
Subsequent work placed these procedures within the broader theory of invariant tests, asymptotic efficiency, and permutation inference. The resulting framework explains why distinct-looking rank statistics often share the same limiting distributions and why their finite-sample behavior remains sensitive to discreteness, ties, and the definition of the null hypothesis.
Efficiency and robustness
The efficiency of a rank test is defined relative to a population model and a competing procedure. Under normally distributed location shifts, the Wilcoxon rank-sum statistic has an asymptotic relative efficiency of approximately (3/\pi), or (0.955), compared with the two-sample (t)-statistic. This value concerns a specific local-alternative calculation and does not constitute a model-independent ranking of the procedures.
For distributions with heavier tails than the normal distribution, rank tests can have higher efficiency because exceptionally large observations receive bounded influence through their ranks. Under distributions in which numerical spacing contains substantial stable information, transforming the data to ranks can reduce power by discarding that information.
Robustness also depends on the form of contamination. Rank statistics limit the effect of extreme magnitudes, but they remain responsive to observations that change the ordering of many sample points. Dependence among observations can invalidate the exchangeability or sign symmetry used to derive the null distribution, even though the calculation itself continues to produce a numerical statistic.
Generalizations
The Kruskal–Wallis test generalizes the two-sample rank sum to multiple independent groups. It compares group rank totals and has an asymptotic chi-squared distribution under its null model. Its rejection indicates a distributional difference among the groups, while a pure location interpretation depends on comparable distributional shapes.
The Friedman test applies ranking within blocks or repeated-measurement units. Because each block is ranked separately, block-level shifts are removed from the statistic, while systematic differences among treatments remain represented by their accumulated ranks.
Rank methods also extend to regression and survival analysis. Rank regression estimates parameters through ordered residuals or rank-based objective functions, while the log-rank test compares event-time distributions using risk sets formed throughout follow-up. The log-rank statistic is historically and mathematically related to rank methods, although censoring prevents it from being reduced to a single ordinary ranking of all observations.