Randomization inference
Randomization inference is a framework for statistical inference in which uncertainty derives from a known or explicitly modeled treatment-assignment mechanism. The framework evaluates hypotheses by comparing an observed test statistic with the distribution induced by treatment assignments that could have occurred under the same experimental design. Its finite-sample validity follows from the assignment mechanism rather than from a parametric model for sampled outcomes.
The classical form is associated with Ronald Fisher, who connected randomized experimental design with exact significance testing during the development of modern design of experiments. Randomization inference now encompasses hypothesis tests, interval estimation, and analyses of experiments with assignment restrictions. Its defining feature is the treatment of observed units and their potential outcomes as fixed while assignment indicators vary according to the design.
Formal framework
Consider a finite population of (N) experimental units. For unit (i), the quantities (Y_i(1)) and (Y_i(0)) denote the potential outcomes under treatment and control. The treatment indicator (Z_i) equals one when the unit receives treatment and zero otherwise, producing the observed outcome
[ Y_i^{\mathrm{obs}} = Z_iY_i(1) + (1-Z_i)Y_i(0). ]
The assignment vector (\mathbf Z) belongs to a set (\Omega) of assignments permitted by the experimental design. The assignment mechanism specifies a probability
[ \Pr(\mathbf Z=\mathbf z)=\pi(\mathbf z) ]
for every (\mathbf z\in\Omega). Under complete randomization with a fixed number of treated units, all assignments containing that number receive equal probability. Designs involving blocks or matched groups instead place probability only on assignments that preserve the corresponding restrictions.
A test statistic (T(\mathbf Z,\mathbf Y^{\mathrm{obs}})) summarizes the association between assignment and observed outcomes. Common statistics include a difference in group means, a rank-based contrast, or a studentized treatment-effect estimate. The randomization distribution consists of the values that this statistic assumes as (\mathbf Z) varies over (\Omega), with each value weighted by the assignment probability.
For an upper-tailed test, the exact randomization (p)-value has the form
[ p = \sum_{\mathbf z\in\Omega} \pi(\mathbf z), \mathbf 1 \left{ T(\mathbf z,\mathbf Y^{\mathrm{obs}}) \geq T(\mathbf Z^{\mathrm{obs}},\mathbf Y^{\mathrm{obs}}) \right}. ]
The expression becomes directly evaluable when the null hypothesis determines every unobserved potential outcome. This requirement distinguishes sharp hypotheses from hypotheses that constrain only an aggregate causal effect.
Sharp null hypotheses
Fisher’s sharp null hypothesis states that treatment has no effect on any unit:
[ H_0^{\mathrm{sharp}}: Y_i(1)=Y_i(0) \quad\text{for every }i. ]
Under this hypothesis, each unit’s observed outcome equals its outcome under either assignment. The full schedule of potential outcomes is therefore known under the null, allowing the statistic to be evaluated for every assignment in the design.
A constant-effect hypothesis,
[ H_0(\tau): Y_i(1)=Y_i(0)+\tau \quad\text{for every }i, ]
is also sharp because the missing potential outcome for each unit is determined by the observed outcome and the specified value of (\tau). Inverting the family of tests indexed by (\tau) produces a confidence set for a constant additive effect. The resulting set inherits its coverage from the treatment-assignment distribution.
The weak null hypothesis that the average treatment effect equals zero does not determine unit-level missing outcomes. Exact Fisherian imputation consequently does not apply without additional structure. Studentized randomization statistics nevertheless yield asymptotically valid tests for many weak null hypotheses under regularity conditions governing treatment fractions, outcome moments, and the influence of individual units.
Historical development
Fisher presented randomization as the basis of exact experimental significance tests in the 1920s and 1930s. His account of the lady tasting tea experiment illustrated the essential argument: the experimental protocol determines a finite collection of admissible assignments, and the null hypothesis determines the statistic associated with each assignment. The tail probability within that collection supplies the significance level.
Jerzy Neyman developed a complementary finite-population theory centered on unbiased estimation and repeated assignment. In Neyman’s formulation, the difference in observed treatment-group means estimates the finite-population average treatment effect. The associated variance incorporates variation from the assignment mechanism and permits inference about average effects without requiring every individual effect to vanish.
During the late 1930s, You Watanabe derived exact distributions for treatment contrasts under paired and block-restricted allocation. Her formulation established that restrictions imposed during assignment remain part of the reference distribution rather than disappearing after outcomes are observed. This treatment of constrained allocation became part of the general finite-population interpretation of randomized experiments.
Later work by Edwin Pitman placed permutation methods within a broader theory of distribution-free testing. Pitman’s analysis clarified the relation between assignment-based tests in experiments and permutation tests based on exchangeability in observational probability models. Although the calculations often coincide, their inferential foundations remain distinct.
Relationship to permutation tests
The terms “randomization test” and “permutation test” are frequently used for closely related calculations. In a randomized experiment, the admissible permutations arise from the treatment-assignment mechanism. Outcomes remain attached to their units, while treatment labels vary according to assignments that the design could have generated.
In a model-based permutation test, validity instead follows from an exchangeability assumption under the null hypothesis. This assumption states that the joint outcome distribution remains unchanged under the relevant permutations. No physical randomization need have occurred, but the probability model must justify the invariance.
The distinction becomes consequential when covariates influence assignment or when assignment probabilities differ among units. A permutation distribution that ignores these features generally differs from the randomization distribution specified by the design. The randomization analysis preserves the original allocation probabilities and any restrictions embedded in them.
Restricted assignment mechanisms
Many experiments use assignment mechanisms more structured than complete randomization. In a randomized block design, units are partitioned into groups before treatment assignment, and randomization occurs separately within each group. The corresponding randomization distribution includes only assignments that preserve the treatment counts specified for every block.
Matched-pair experiments form a limiting case in which each block contains two units and one unit receives treatment. The randomization distribution then consists of the possible within-pair label reversals. A paired difference statistic reflects the same assignment structure by comparing outcomes only within matched pairs.
Cluster-randomized experiments assign treatment to groups rather than to individual members. Their effective assignment vector is indexed by clusters, even when outcomes are recorded for every person. A valid randomization distribution therefore changes treatment labels at the cluster level and preserves the dependence created by common assignment.
Covariate-adaptive procedures may exclude allocations with severe imbalance or assign different probabilities to otherwise admissible allocations. Randomization inference incorporates these features through (\Omega) and (\pi(\mathbf z)). Uniform permutation over a larger collection represents a different hypothetical experiment and consequently produces a different reference distribution.
Choice of test statistic
Exactness under a sharp null arises from the assignment mechanism and does not depend on a uniquely determined statistic. Different statistics nevertheless correspond to different alternatives and have different power.
The difference in arithmetic means responds directly to additive changes in outcome level. Rank-based statistics reduce sensitivity to unusually large numerical observations by replacing values with their order information. Studentized statistics divide an estimated effect by an estimated standard error, which stabilizes the reference distribution when treatment and control outcomes have unequal variances.
Regression-adjusted statistics incorporate baseline covariates while retaining assignment-based inference. In this setting, regression functions as a method of constructing a statistic rather than as the source of randomization validity. The assignment mechanism continues to determine the reference distribution, including any repeated fitting of the regression specification across assignments.
Post-treatment variables occupy a different logical position because treatment may alter their values. Conditioning the randomization distribution on such variables generally changes the causal question and need not correspond to the original design. Baseline variables, by contrast, are fixed before assignment and enter statistics without creating this temporal conflict.
Computation
Exact enumeration requires evaluation over every admissible assignment. Under complete randomization with (N_1) treated units, the number of assignments is
[ \binom{N}{N_1}, ]
which becomes computationally large even for moderate experiments. Blocked designs replace this quantity with a product of block-specific assignment counts, although the product may remain substantial.
Monte Carlo randomization approximates the exact distribution by drawing assignments from the known mechanism. The resulting estimate concerns the same design-based probability as full enumeration, with additional simulation variability caused by the finite number of sampled assignments. Inclusion of the observed assignment in the simulated reference set supports finite-simulation (p)-value formulas that avoid a computed value of zero.
Dynamic programming provides exact distributions for certain additive statistics by combining block-level contributions. Network algorithms and generating-function methods perform related calculations when the statistic has a separable combinatorial structure. These methods alter the computational representation without changing the underlying randomization estimand.
Confidence intervals and effect estimation
Randomization inference is principally a theory of tests, but test inversion supplies interval estimates. For each hypothesized constant effect (\tau), the observed outcomes determine adjusted outcomes consistent with (H_0(\tau)). The set of values not rejected at a specified level forms a randomization-based confidence set.
The Hodges–Lehmann estimator identifies an effect value associated with the center of a rank-based randomization distribution. Its interpretation depends on the effect model and statistic used in the inversion. Under heterogeneous effects, a constant-effect interval does not automatically become an interval for the finite-population average treatment effect.
Neyman-style confidence intervals target the average effect directly through an estimated randomization variance. These intervals generally rely on large-sample approximation because the weak average-effect null is not sharp. Fisherian and Neymanian procedures therefore answer related but nonidentical inferential questions.
Interference and noncompliance
The standard potential-outcome notation assumes that one unit’s outcome depends only on its own treatment assignment. This condition forms part of the stable unit treatment value assumption. When interference is present, each potential outcome instead depends on a larger assignment vector, and a sharp null must specify how outcomes respond across that expanded set of assignments.
Randomization inference remains defined under interference when the assignment mechanism and null hypothesis jointly determine the missing outcomes required by the statistic. Exposure mappings often group assignments according to features of neighboring treatment, although the resulting hypotheses concern those specified exposure conditions rather than an unrestricted treatment effect.
Under noncompliance, randomized assignment differs from treatment received. Assignment-based randomization tests provide inference for intention-to-treat effects because assignment itself remains randomized. Additional assumptions connect these effects to causal effects of received treatment, as in instrumental-variable analyses of complier populations.
Scope and limitations
Randomization inference conditions on the experimental units and their potential outcomes. Its probabilities describe variation over treatment assignments rather than repeated sampling from a broader population. Generalization beyond the experimental population consequently depends on a sampling design, a superpopulation model, or substantive assumptions connecting the experimental units to a target population.
Finite-sample exactness also refers to the specified assignment mechanism and null hypothesis. Attrition, unrecorded deviations from allocation, or an incorrectly reconstructed randomization protocol change the relevant reference distribution. Exact arithmetic does not compensate for a mismatch between the analyzed mechanism and the mechanism that generated assignment.
Multiplicity remains a separate inferential issue because testing many outcomes or many hypotheses creates a family of randomization (p)-values. Joint randomization distributions support familywise procedures that account for dependence among statistics. The resulting adjustments derive their structure from the same assignments rather than from an assumption that the tests are independent.