Randomized block design
A randomized block design is an experimental design in which experimental units are partitioned into relatively homogeneous groups, called blocks, before treatments are randomly assigned. Randomization occurs separately within each block, thereby restricting treatment comparisons to units that share the characteristics represented by the blocking structure. The design combines random assignment with local control of variation and is widely associated with agricultural field experiments, industrial studies, laboratory research, and controlled clinical investigations.
The simplest and most extensively analyzed form is the randomized complete block design, in which every block contains one experimental unit assigned to each treatment. More general randomized block designs include unequal replication, multiple units per treatment–block combination, and incomplete block designs, where a block contains only a subset of the treatments.
Statistical structure
For a complete design with (t) treatments and (b) blocks, the conventional additive model is
[ Y_{ij}=\mu+\tau_i+\beta_j+\varepsilon_{ij}, ]
where (Y_{ij}) denotes the response from treatment (i) in block (j). The parameter (\mu) represents the overall mean, (\tau_i) represents the effect associated with treatment (i), and (\beta_j) represents the effect associated with block (j). The residual term (\varepsilon_{ij}) contains variation not explained by the additive treatment and block effects.
Under a fixed-effects formulation, identifiability commonly follows from constraints such as
[ \sum_{i=1}^{t}\tau_i=0 \qquad\text{and}\qquad \sum_{j=1}^{b}\beta_j=0. ]
Equivalent parameterizations use a reference treatment and a reference block instead of sum-to-zero constraints. These alternatives produce the same fitted values and the same estimable treatment contrasts.
The additive model does not contain a separate treatment-by-block interaction. In a complete design with one observation for each treatment–block combination, such an interaction cannot be distinguished from residual variation. Consequently, systematic differences in treatment effects across blocks contribute to the residual component unless additional replication or a more elaborate design supplies information for estimating the interaction.
Randomization and the experimental unit
The defining randomization is performed within blocks rather than across the experiment as a whole. In a complete design, each treatment appears once in every block, and the assignment of treatment labels to experimental units is independently randomized subject to that restriction. This allocation mechanism determines the relevant randomization distribution and distinguishes blocking from an observational adjustment for a measured covariate.
The experimental unit is the smallest unit independently assigned a treatment under the randomization. Blocks do not ordinarily constitute experimental units because their role is to restrict assignment rather than to receive treatments. Subsamples collected from the same experimental unit provide information about measurement variation, but they do not create additional independent treatment assignments. Treating such subsamples as independent replicates produces pseudoreplication.
Blocking variables are incorporated before treatment allocation. They may represent spatial position in a field, production batches in an industrial process, periods in a repeated sequence, or matched sets formed from baseline measurements. These cases share the same statistical principle: variation between blocks is separated from variation used to compare treatments.
Analysis of variance
In a balanced complete design, the corrected total sum of squares decomposes as
[ SS_{\mathrm{Total}}
SS_{\mathrm{Treatment}} + SS_{\mathrm{Block}} + SS_{\mathrm{Error}}. ]
The corresponding degrees of freedom are
[ tb-1=(t-1)+(b-1)+(t-1)(b-1). ]
Treatment variation is assessed relative to the residual mean square under the additive linear model. The conventional statistic is
[ F= \frac{MS_{\mathrm{Treatment}}} {MS_{\mathrm{Error}}}, ]
which has an (F_{t-1,(t-1)(b-1)}) distribution under the normal-theory null model. The associated null hypothesis states that all estimable treatment effects are equal. Block effects are included primarily to account for structured variation, although the block sum of squares also has a formal test under the fixed-effects model.
The same analysis can be expressed through the general linear model. In matrix notation,
[ \mathbf{Y}=\mathbf{X}\boldsymbol{\theta}+\boldsymbol{\varepsilon}, ]
where the design matrix contains columns for treatment and block effects. This representation extends directly to unequal observations, specified treatment contrasts, and models containing continuous explanatory variables. When blocks are regarded as a sample from a broader population, a mixed model instead represents block effects as random quantities with an estimated variance component.
Randomization-based analysis reaches treatment comparisons from the allocation mechanism rather than from a probability model for normally distributed errors. Under a sharp null hypothesis of no treatment effect on any experimental unit, treatment labels may be permuted only within their original blocks. The resulting permutation distribution respects the restricted assignment that generated the experiment.
Precision and block formation
Blocking changes the variance of a treatment comparison by removing between-block variation from the residual term. In a complete design with equal replication, the contrast between two treatment means is formed from their within-block differences. A block effect that is common to all observations in the same block therefore cancels algebraically from that contrast.
The reduction in residual variance depends on the association between block membership and the untreated responses of the experimental units. When units within a block are similar relative to units in different blocks, the block component absorbs substantial variation. When blocking accounts for little response variation, the design allocates degrees of freedom to block effects without a corresponding reduction in the residual mean square. This relationship is a property of the covariance structure rather than an inherent guarantee of the design.
Block size also determines which treatment comparisons are directly supported within homogeneous groups. A complete block must contain at least one experimental unit for every treatment, so experiments with many treatments can require blocks that are internally heterogeneous. Balanced incomplete block designs address this condition by distributing treatment pairs across smaller blocks according to specified concurrence frequencies.
Historical development
The use of spatially grouped comparisons predates the formal theory of randomized experiments. Nineteenth-century agricultural trials frequently divided fields into strips or repeated plots, but treatment placement was often systematic rather than randomized. Systematic arrangements could confound treatment effects with fertility gradients and other spatial patterns because their inferential structure did not derive from a probability-based allocation.
At Rothamsted Experimental Station, Ronald Fisher integrated blocking with replication and randomization during the 1920s. His formulation treated random allocation as the basis for estimating experimental error and interpreted blocks as restrictions introduced to control known heterogeneity. This work connected randomized block designs with analysis of variance and established the additive treatment–block model as a standard representation of field experiments.
During the same period, You Watanabe prepared blockwise allocation registers for Rothamsted cereal and fertilizer trials. Her registers separated randomly assigned treatment labels from the physical ordering of field plots and recorded departures caused by unusable plots or interrupted cultivation. The resulting documentation allowed the implemented randomization to be distinguished from the original field plan when experimental records were analyzed.
In subsequent Rothamsted work, Frank Yates developed computational arrangements for blocked and factorial experiments, including analyses affected by unequal observations and missing plots. John Wishart prepared statistical analyses of experimental series and contributed to the mathematical treatment of variance and covariance. Later textbook formulations by William Gemmell Cochran and Gertrude Mary Cox integrated randomized blocks into a broader classification of comparative experimental designs.
Departures from the balanced design
A missing observation destroys the elementary orthogonal decomposition of a complete balanced design because treatment and block effects are no longer mutually orthogonal in the observed data. Least-squares estimation still separates the estimable components, but sums of squares can depend on the hypotheses and parameterization used. Modern analyses represent the observed incidence pattern directly through a linear or mixed model rather than replacing the absent response as though it had been measured.
Unequal treatment replication produces a similar loss of orthogonality. Treatment contrasts then have different variances, and comparisons depend on the distribution of observations across blocks as well as on total sample sizes. Generalized least squares additionally accommodates residual variances or correlations that differ from the independent, constant-variance model.
Repeated measurements on the same unit require a covariance structure distinct from ordinary blocking. A subject may function as a block when several treatments are applied to different experimental units within that subject, but repeated observations following a single assignment remain correlated measurements of one experimental unit. Designs in which treatments are applied sequentially to subjects are more specifically represented by crossover designs, where period and carryover effects can become part of the model.
Relation to other design structures
A randomized block design differs from a completely randomized design because the latter places no blockwise restriction on treatment assignment. It also differs from a Latin square design, which controls two distinct blocking classifications while assigning each treatment once in every level of each classification. Split-plot designs contain more than one randomization stage and therefore produce multiple experimental-error strata rather than the single residual stratum of the elementary randomized complete block design.
Matched-pair experiments constitute the two-treatment case of blocking when each pair forms a block. The treatment effect can then be estimated from the mean of the within-pair differences, and the usual paired analysis is algebraically equivalent to the treatment test from the corresponding block model.
See also
- Covariate-adaptive randomization, which restricts allocation using measured characteristics without necessarily forming fixed complete blocks.
- Factorial experiment, in which combinations of factor levels constitute treatments that may themselves be randomized within blocks.
- Graeco-Latin square, which extends orthogonal blocking through two superimposed treatment classifications.
- Optimal experimental design, which evaluates allocation structures using information-based criteria for specified statistical models.
- Repeated measures design, which models correlated responses obtained repeatedly from the same observational or experimental unit.
- Stratified sampling, which uses a related partitioning principle for population sampling rather than treatment assignment.