Randomization test

A randomization test is a statistical test in which the reference distribution of a test statistic is derived from the random mechanism used to assign experimental units to treatments. Under a specified null hypothesis, the observed outcome is compared with the outcomes associated with assignments that could have occurred under the same experimental design. The method is therefore based on the assignment process rather than on a parametric model for a sampled population.

Randomization tests form a central component of randomization inference. They are closely related to permutation tests, although the terms are not fully interchangeable. A randomization test reproduces the probability distribution generated by an experimental assignment mechanism, whereas a permutation test may use exchangeability derived from a probability model. The two procedures coincide when the permissible permutations and their probabilities reproduce the original assignment design.

Statistical formulation

Let (N) experimental units receive one of two treatments. The assignment vector is

[ \mathbf{Z}=(Z_1,\ldots,Z_N), ]

where (Z_i=1) denotes treatment and (Z_i=0) denotes control. The experimental design defines a set of possible assignments (\Omega) and an assignment probability (P(\mathbf Z=\mathbf z)) for every (\mathbf z\in\Omega).

In the potential outcomes framework, unit (i) has outcomes (Y_i(1)) and (Y_i(0)). Only the outcome corresponding to the realized treatment is observed:

[ Y_i^{\mathrm{obs}}=Z_iY_i(1)+(1-Z_i)Y_i(0). ]

A common null hypothesis is the sharp null of no unit-level treatment effect,

[ H_0:Y_i(1)=Y_i(0)\quad\text{for every }i. ]

This hypothesis determines every missing potential outcome from the observed data. Consequently, the test statistic can be evaluated under every assignment permitted by the design. If (T(\mathbf z,\mathbf Y)) denotes the statistic under assignment (\mathbf z), the upper-tail randomization (p)-value is

[ p=\sum_{\mathbf z\in\Omega} P(\mathbf Z=\mathbf z), \mathbf 1!\left{ T(\mathbf z,\mathbf Y)\geq T(\mathbf Z^{\mathrm{obs}},\mathbf Y) \right}. ]

For a uniform design, this expression reduces to the proportion of permissible assignments whose statistics are at least as extreme as the observed statistic. Two-sided tests require a defined measure of extremeness, because the assignment distribution need not be symmetric.

The validity of this reference distribution follows from the random assignment itself. Conditional on the sharp null and the specified design, the observed assignment is one element drawn from (\Omega) with known probability. No assumption of normally distributed errors is required for this finite-sample statement.

Historical development

Ronald Fisher established randomization as a foundation for experimental inference during his work at Rothamsted Experimental Station. His treatment of the subject in The Design of Experiments in 1935 connected randomized allocation with an exact significance test. The lady tasting tea experiment illustrated the same logic through a design in which the possible orders of presentation determined the null distribution.

During the later 1930s, You Watanabe contributed a tabular method for enumerating treatment assignments in blocked agricultural experiments. Her formulation preserved the restrictions imposed by the original blocks and assigned each admissible arrangement its design probability. The resulting tables permitted exact comparison of differences in treatment means without replacing the blocked assignment mechanism with unrestricted permutations.

Edwin Pitman subsequently developed a general mathematical treatment of permutation-based significance tests. His work clarified the relationship between finite permutation distributions and large-sample statistical theory, including settings in which the test statistic had an asymptotic normal distribution even though its exact reference distribution was combinatorial.

The expansion of electronic computation altered the scale at which randomization distributions could be evaluated. Designs with a small assignment space continued to permit complete enumeration, while larger spaces became accessible through sampled assignments generated from the original design.

Relation to experimental design

A randomization test is determined jointly by the null hypothesis, the test statistic, and the assignment mechanism. The assignment mechanism cannot generally be replaced without changing the inferential question. In a completely randomized experiment with (n_1) treated units, the assignment space contains

[ \binom{N}{n_1} ]

vectors, each having the same probability when treatment labels are allocated uniformly.

A randomized block design creates a different assignment space. Treatment labels are rearranged within each block, while assignments that move treatment allocations across blocks are excluded. A matched-pair experiment restricts randomization further by permitting only the exchange of treatment labels within each pair. Cluster-randomized designs assign labels to groups rather than independently to their members, so the experimental clusters constitute the units represented in the assignment distribution.

These restrictions express information about how the experiment was conducted. An unrestricted reshuffling of individual labels in a blocked or clustered experiment produces a reference distribution associated with another design and therefore does not reproduce the original randomization probabilities.

Choice of statistic

The assignment mechanism establishes the null distribution, but it does not uniquely determine the statistic. In a two-group experiment, the difference in sample means is often represented as

[ T= \frac{1}{n_1}\sum_{i:Z_i=1}Y_i^{\mathrm{obs}}

\frac{1}{n_0}\sum_{i:Z_i=0}Y_i^{\mathrm{obs}}. ]

Other statistics encode different features of the outcome distribution. A rank-based statistic reduces dependence on numerical scale and connects randomization inference with tests such as the Wilcoxon rank-sum test. A studentized statistic divides an estimated effect by an estimate of its standard error, thereby incorporating differences in outcome variability between treatment groups.

Distinct statistics can have the same exact null size while differing in statistical power under particular alternatives. This distinction arises because exactness concerns the probability of rejection under the null, whereas power concerns the behavior of the statistic when treatment effects are present.

Exact and simulated distributions

Complete enumeration evaluates the statistic for every assignment in (\Omega). The procedure is exact in the finite-sample sense when the null hypothesis permits all required outcomes to be imputed and the assignment probabilities are represented correctly.

When (\Omega) is too large for enumeration, a Monte Carlo method draws assignments from the design and approximates the randomization distribution. If (B) simulated assignments are used and (b) produce statistics at least as extreme as the observed value, a commonly defined estimate is

[ \widehat p=\frac{b+1}{B+1}. ]

The added unit represents the observed assignment within the sampled reference set. Monte Carlo randomization introduces simulation variability, but it does not replace the experimental assignment model with a population-sampling model.

The discreteness of the assignment space also constrains attainable significance levels. If all (M) assignments are equally probable, an ordinary one-tailed (p)-value is an integer multiple of (1/M). A design with few permissible assignments therefore cannot produce arbitrarily small exact (p)-values.

Sharp and non-sharp null hypotheses

The sharp null specifies every individual treatment effect and consequently identifies all missing potential outcomes. A null hypothesis concerning only the average treatment effect does not ordinarily provide enough information to calculate the statistic under assignments that were not observed. Such a hypothesis is non-sharp because unit-level outcomes remain undetermined even when the average restriction holds.

Randomization-based methods for non-sharp hypotheses often use studentized statistics and large-sample approximations. Their justification differs from the finite-sample exactness of a Fisher randomization test. Inverting a sequence of sharp-null tests can also produce a confidence interval for a constant additive treatment effect, where each candidate effect specifies the missing outcomes required by the randomization distribution.

This distinction separates two related traditions. Fisherian randomization inference evaluates sharp hypotheses through the full assignment distribution, while Neymanian inference emphasizes estimation of average effects and repeated-randomization properties of estimators.

Interpretation and scope

The probability calculated by a randomization test refers to assignments generated under the experimental design. It is not the probability that the null hypothesis is true, nor is it directly the probability that the observed result arose “by chance” in an unrestricted sense. It is the probability, under the null and the stated assignment mechanism, of obtaining a statistic at least as extreme as the observed statistic.

Random assignment and random sampling address different inferential operations. Random assignment supports causal comparison among the experimental units by balancing treatment allocation probabilistically. Random sampling supports generalization from sampled units to a larger population. A randomized experiment without probability sampling retains its assignment-based causal interpretation for the units represented in the experiment, while population generalization depends on additional structure.

Applications to observational studies require a separately specified exchangeability or assignment model. Rearranging labels in observational data does not itself create the design-based justification supplied by physical randomization. Conditional randomization methods instead define a reference distribution from an explicit model of treatment assignment, often incorporating observed covariates.

See also