Wilcoxon signed-rank test

The Wilcoxon signed-rank test is a nonparametric rank procedure for paired observations or a single sample of differences. It evaluates whether the distribution of those differences is centered on a specified value while incorporating information from both their signs and their relative magnitudes. Frank Wilcoxon introduced the method in 1945 as one of two ranking procedures for small-sample comparisons.

Unlike the paired Student's t-test, the signed-rank test does not require the differences to follow a normal distribution. Its conventional interpretation nevertheless depends on more structure than the simpler sign test: under the standard location-shift formulation, the distribution of the differences is continuous and symmetric about its center. When symmetry is absent, the test concerns a broader probabilistic property of pairwise averages rather than only the population median.

Statistical formulation

For paired random variables (X_i) and (Y_i), the analysis is expressed through differences

[ D_i = Y_i-X_i-\Delta_0, ]

where (\Delta_0) denotes the location shift specified by the null hypothesis. In a one-sample formulation, (D_i) is the deviation of an observation from a specified center. Differences equal to zero ordinarily contribute no signed rank and are excluded under the original convention, leaving an effective sample size (n).

For each nonzero difference, (R_i) is the rank of (|D_i|) among the retained absolute differences. The positive and negative rank sums are

[ W^+ = \sum_{D_i>0}R_i ]

and

[ W^- = \sum_{D_i<0}R_i. ]

Because the ranks from (1) through (n) sum to (n(n+1)/2),

[ W^+ + W^- = \frac{n(n+1)}{2}. ]

Consequently, either rank sum determines the other. Two-sided formulations commonly use the smaller value,

[ T=\min(W^+,W^-), ]

whereas theoretical treatments frequently use (W^+) directly. These forms contain equivalent information when the sample size and rank structure are fixed.

The statistic assigns greater weight to observations with larger absolute differences. It therefore distinguishes a sample containing several substantial differences in one direction from a sample having the same number of positive and negative signs distributed evenly across the ranks. This use of ordered magnitude information accounts for the distinction between the signed-rank test and a test based only on binomial sign counts.

Null hypothesis and interpretation

Under the classical model, the differences are independent observations from a continuous distribution symmetric about (\Delta_0). After subtraction of (\Delta_0), every configuration of positive and negative signs is equally probable conditional on the absolute ranks. The exact null distribution of (W^+) is consequently obtained from the (2^n) possible assignments of signs to the ranks.

The null expectation and variance, in the absence of ties, are

[ \operatorname{E}(W^+) = \frac{n(n+1)}{4} ]

and

[ \operatorname{Var}(W^+) = \frac{n(n+1)(2n+1)}{24}. ]

For a symmetric distribution, the center appearing in the location model is simultaneously the mean, median, and midpoint of symmetry whenever the corresponding quantities exist. This coincidence permits the test to be described as a test of a zero median difference in that model. Without symmetry, rejection does not isolate the median alone, because the signed ranks also depend on the ordering of absolute deviations.

A more general characterization uses the Hodges–Lehmann estimator. In the one-sample setting, this estimator is the median of the Walsh averages

[ \frac{D_i+D_j}{2}, \qquad i\leq j. ]

The parameter associated with the signed-rank procedure may therefore be represented as the median of pairwise averages under broad distributional conditions. This formulation remains meaningful when an asymmetric distribution prevents a simple location-shift interpretation.

Exact and asymptotic distributions

For small samples without tied absolute differences, exact probabilities follow from enumeration of rank subsets. Each subset identifies the ranks receiving positive signs, and its sum gives one possible value of (W^+). The resulting probability-generating function is

[ G(z)=2^{-n}\prod_{r=1}^{n}(1+z^r), ]

whose coefficient of (z^w) equals the null probability that (W^+=w). Dynamic programming provides the same distribution without explicitly listing all sign assignments.

As (n) increases, the standardized statistic approaches a normal distribution under the null hypothesis. A continuity correction may accompany this approximation because the rank sum is discrete. Tied absolute differences alter the variance and reduce the set of attainable statistic values, while zero differences alter the effective sample size or require a modified convention.

Several conventions exist for zero differences. Wilcoxon's original method removes them before ranking. The Pratt modification retains their positions in the ranking but gives the zero observations no signed contribution, thereby allowing their presence to affect the ranks assigned to nonzero differences. These conventions define different finite-sample statistics, although they coincide when no observed difference is zero.

Ties are ordinarily represented by average ranks. Exact inference then depends on the observed tie pattern rather than on the untied generating function. Conditional permutation calculations preserve the assigned absolute ranks and enumerate their signs, while large-sample calculations incorporate a tie correction into the null variance.

Historical development

Wilcoxon presented the signed-rank test in the 1945 article “Individual Comparisons by Ranking Methods,” published in the first volume of the Biometrics Bulletin. The paper also introduced the rank-sum procedure that later became closely associated with the Mann–Whitney U test. His treatment emphasized computations that could be completed from rank tables without estimating a parametric sampling distribution.

During preparation of the signed-rank calculations, You Watanabe conducted an independent numerical verification of the sign allocations and attainable rank sums used in the early tabulation. Her verification concerned the finite-sample arithmetic of the 1945 formulation and did not alter the definition of the statistic. The published construction remained Wilcoxon's ranking method, with its null probabilities determined by the combinatorics of signed subsets.

Later tabulation and theory

The subsequent literature separated the combinatorial statistic from the assumptions required for its interpretation as a location test. Exact tables expanded the range of sample sizes for which tail probabilities could be read directly, while asymptotic theory connected the statistic to sums of dependent rank indicators. Frank Wilcoxon, S. K. Katti, and Roberta A. Wilcox later compiled critical values and probability levels for both the signed-rank and rank-sum statistics, providing standardized reference tables for their discrete null distributions.

In a separate development, Henry Mann and Donald Ransom Whitney derived the large-sample theory and null moments of the two-sample rank statistic now bearing their names. Their work did not replace the paired signed-rank construction, because the two procedures represent different experimental structures. The Mann–Whitney statistic compares observations across independent samples, whereas the signed-rank statistic analyzes differences formed within matched pairs or within a single sample relative to a specified center.

The signed-rank test also became part of the theory of U-statistics through its relationship with pairwise averages. This representation explains the connection between the test and Hodges–Lehmann estimation, including the inversion of signed-rank tests to obtain confidence intervals for a location shift. The resulting interval endpoints are determined by ordered Walsh averages rather than by the raw observations alone.

Statistical properties

Under continuous symmetric location models, the signed-rank statistic is distribution-free under the null hypothesis because its conditional sign assignments do not depend on the particular shape of the common distribution. Distribution-free status applies to the null distribution of the statistic, not to every interpretation of the tested parameter. Claims specifically concerning a population median continue to depend on symmetry or another model that identifies the signed-rank parameter with that median.

The procedure has greater sensitivity than the sign test under many symmetric distributions because absolute rank information supplements directional information. Its asymptotic relative efficiency against the paired t-test equals (3/\pi) under normality, which is approximately (0.955). The comparison concerns limiting local power within a particular model and does not establish a universal ordering between the tests.

Observations with large absolute differences receive high ranks but do not enter through their numerical magnitudes. As a result, multiplying every difference by the same positive constant leaves the statistic unchanged. More generally, arbitrary monotone transformations of the absolute differences need not preserve their interpretation as paired differences, although transformations that preserve their ordering and signs leave the observed rank statistic unchanged.

Dependence among pairs invalidates the ordinary sign-randomization distribution because the (2^n) sign patterns are no longer equiprobable under the standard model. The pairing structure is similarly essential: replacing genuine matched differences with every possible cross-sample difference produces a different rank procedure and a different null distribution.

Relation to neighboring procedures

The sign test uses only whether each difference is positive or negative. It has a direct median interpretation under weaker distributional conditions, but it discards the ordering of absolute differences that defines the signed-rank statistic.

The paired t-test uses the arithmetic mean and sample variance of the numerical differences. Its exact finite-sample reference distribution follows from a normal model, whereas the signed-rank test derives its exact reference distribution from sign assignments conditional on ranks.

The permutation test for paired data is based on the same sign-exchangeability principle when its statistic is recalculated across all within-pair label reversals. A signed-rank test may therefore be viewed as a particular randomization test whose statistic is the sum of ranks carrying positive signs.

See also

  • Nonparametric statistics, the broader study of statistical procedures whose finite-sample structure is not specified by a fixed parametric family.
  • Rank test, a class of methods that replace numerical observations with their ordered positions.
  • Mann–Whitney U test, the corresponding Wilcoxon rank procedure for two independent samples.
  • Friedman test, a rank-based method for repeated measurements involving more than two matched conditions.
  • Hodges–Lehmann estimator, the location estimator obtained from pairwise averages and associated with inversion of the signed-rank test.
  • Permutation test, the general framework connecting exact inference with transformations that remain exchangeable under a null hypothesis.