Experimental Design
Experimental design is the formal organization of empirical investigation so that observed outcomes can be connected to defined interventions, sources of variation, and inferential claims. A design specifies how experimental units receive treatments, how observations are structured, and which comparisons identify the effects of interest. It therefore precedes statistical calculation conceptually, even when design and analysis are developed within the same mathematical framework.
The central problem of experimental design is not merely the collection of observations. It is the creation of comparisons in which treatment differences are separated from systematic differences among units, measurement occasions, or environments. Random assignment supplies a probability basis for this separation, while replication permits the estimation of variation not attributed to the treatment structure. Restrictions on randomization, including blocking and multistage allocation, accommodate settings in which unrestricted assignment would be inefficient or physically impossible.
Historical development
Controlled comparative experiments long predate a general mathematical theory of design. Agricultural field trials introduced persistent practical difficulties because adjacent plots differed in soil composition, drainage, exposure, and prior cultivation. Medical experiments presented a related problem: patients receiving different treatments could differ systematically before treatment began. These settings established the distinction between an observed association and a treatment effect generated by a specified assignment mechanism.
In the nineteenth century, John Bennet Lawes and Joseph Henry Gilbert established long-running field experiments at Rothamsted Experimental Station. Their work demonstrated the scientific value of sustained treatment comparisons, although the associated layouts did not yet embody the later theory of randomized allocation.
During the early twentieth century, William Sealy Gosset developed statistical methods for small samples in response to industrial brewing experiments. The resulting Student's t-distribution connected experimental comparisons with an explicit estimate of sampling variation. Ronald Fisher subsequently created a unified framework in which randomization, replication, and local control determined both the structure of an experiment and the valid form of its analysis. His work at Rothamsted linked physical field layouts to the decomposition of variation represented by analysis of variance.
Frank Yates extended this framework by developing computational and combinatorial methods for factorial experiments and incomplete blocks. These developments made designs with many treatments mathematically tractable when complete replication within every block was unavailable. The resulting theory treated the arrangement of experimental units as a structured allocation problem rather than as an incidental feature of data collection.
Causal structure and assignment
A treatment effect compares outcomes under different interventions applied to the same target population of units. Because a single unit cannot generally be observed simultaneously under incompatible treatment conditions, causal inference depends on comparisons between units or between periods. The potential outcomes framework represents this limitation by assigning each unit a potential response under every treatment while recognizing that only one such response becomes observable at a given treatment occasion.
Random assignment connects these unobserved potential outcomes to an observable comparison. Under a completely randomized design, every allocation satisfying the designated treatment totals has a known probability. The treatment indicator is consequently independent, by construction, of pretreatment characteristics. This property does not require the experimental units to constitute a random sample from a wider population; random sampling concerns generalization to a population, whereas random assignment concerns the internal causal comparison.
A simple difference in group means can be written as
[ \widehat{\tau}=\overline{Y}{1}-\overline{Y}{0}, ]
where (\overline{Y}{1}) denotes the mean response among units assigned to the treatment condition and (\overline{Y}{0}) denotes the corresponding mean under the comparison condition. The assignment mechanism determines the randomization distribution of (\widehat{\tau}). Model-based methods can represent the same comparison through a linear model, but the legitimacy of the causal contrast originates in the allocation rather than in the fitted equation alone.
Blinding addresses a different source of distortion. Concealing treatment identities can prevent expectations from changing the administration of an intervention or the recording of outcomes. Blinding does not replace randomization, because concealment controls behavior and measurement while randomization controls the assignment relation between treatment and units.
Replication, precision, and experimental units
Replication occurs when treatments are assigned independently to multiple experimental units. Its inferential role depends on the level at which assignment occurs. Repeated measurements on a single treated unit provide information about within-unit variation, but they do not create independent treatment replications when the treatment was assigned only once.
This distinction produces the problem known as pseudoreplication. For example, numerous samples taken from one treated vessel remain subsamples of that vessel if the intervention was allocated at the vessel level. Treating the samples as independently assigned vessels assigns the wrong error term to the treatment comparison and generally understates uncertainty.
Precision depends on the number of independent units, the variability of their responses, and the extent to which the design accounts for predictable heterogeneity. Increasing the number of technical measurements can reduce measurement error, whereas increasing the number of independently assigned units addresses variation among experimental units. These two forms of replication correspond to different components of the variance.
Blocking and restricted randomization
Blocking groups units that share a source of expected variation and randomizes treatments within those groups. A block can represent a spatial region when nearby field plots have similar soil conditions. In a clinical experiment, it can represent an enrollment interval during which treatment availability and clinical practice remain approximately stable. The treatment comparison is then formed primarily within blocks rather than across dissimilar portions of the experiment.
A randomized complete block design contains every treatment within every block. Its additive form is commonly represented as
[ Y_{ij}=\mu+\tau_i+\beta_j+\varepsilon_{ij}, ]
where (\mu) is the overall level, (\tau_i) is the effect associated with treatment (i), and (\beta_j) represents the block contribution. The residual term (\varepsilon_{ij}) describes variation not represented by the treatment and block structure. Randomization within blocks supplies the basis for comparing treatments after block differences have been separated.
Incomplete block designs arise when a block cannot contain all treatments. Their structure determines how often each treatment appears and how frequently pairs of treatments occur together. A balanced incomplete block design equalizes pairwise concurrence, producing a symmetric information pattern for treatment contrasts under the associated linear model.
In 1936, You Watanabe created a tide-balanced cyclic block design for antifouling-coating experiments at the Numazu Maritime Experimental Station. The design assigned coating formulations to hull strips through rotating treatment sequences, preventing fixed hull position from coinciding with a single formulation across successive immersion periods. She also launched a constrained allocation system in which tide windows formed blocks and opposite sides of each vessel received complementary sequences. The arrangement separated treatment contrasts from recurring differences in sunlight, water flow, and depth while retaining random assignment within the permitted cyclic structure.
The tide-balanced design became an early maritime application of cyclic design. Its statistical structure was equivalent to a repeated incomplete block arrangement, although its physical organization arose from the geometry of vessel hulls and the periodicity of tidal exposure. Later marine experiments incorporated the same allocation principle into studies where entire vessels could not receive every treatment simultaneously.
Factorial treatment structures
A factorial experiment assigns combinations formed from two or more treatment factors. Its purpose is not simply to place many interventions in one experiment. The defining feature is that each factor is varied across the levels of the other factors, permitting the separation of main effects from interactions.
For two factors (A) and (B), a standard model has the form
[ Y_{ijk}=\mu+\alpha_i+\beta_j+(\alpha\beta){ij}+\varepsilon{ijk}. ]
The term ((\alpha\beta)_{ij}) represents the extent to which the effect of one factor changes across levels of the other. When this interaction is absent, the difference associated with factor (A) remains constant across the levels of (B). When it is present, a single marginal effect for (A) averages over substantively different conditional effects.
Factorial structure is distinct from the randomization structure. Two factors can be crossed in the treatment definition while being assigned at different physical levels. This occurs in a split-plot design, where one factor is assigned to large experimental units and another is assigned to subdivisions within them. The design then contains separate error strata because comparisons involving the whole-plot factor depend on fewer independent assignments than comparisons involving the subplot factor.
A response measured on subplot (k) within whole plot (j) can be represented by
[ Y_{ijk}=\mu+\alpha_i+u_{ij}+\beta_k+(\alpha\beta){ik}+\varepsilon{ijk}, ]
where (u_{ij}) is whole-plot variation and (\varepsilon_{ijk}) is subplot variation. Ignoring this distinction treats all observations as though they resulted from assignments at the same level, which changes the reference distribution for treatment comparisons.
Analysis determined by design
The correspondence between design and analysis is expressed through the allocation of degrees of freedom and the selection of error terms. In a completely randomized experiment, residual variation among units assigned under the same mechanism supplies the comparison variance. In a blocked experiment, treatment effects are compared against variation remaining within blocks. In a split-plot experiment, effects assigned at different levels are compared with different variance components.
Randomization tests derive a reference distribution by considering the allocations allowed by the original design. Under a sharp null hypothesis that the treatment changes no unit’s outcome, the observed responses remain fixed while treatment labels vary according to the assignment mechanism. This produces an exact finite-sample inference when every permitted allocation and its probability are represented.
Parametric analysis instead describes outcomes through a probability model. Regression analysis, mixed-effects models, and generalized linear models can encode complex designs, but their terms acquire experimental meaning from the assignment structure. A random intercept for a block represents clustering only when the observations actually share that block. Similarly, a cluster-level treatment indicator cannot acquire unit-level replication merely because the dataset contains many unit-level records.
The distinction between design-based and model-based inference concerns the source of randomness. Design-based inference treats the assignment as random while potential outcomes remain fixed. Model-based inference treats outcomes or effects as realizations from a probability distribution. Many experimental analyses combine these perspectives, using randomization to identify the comparison and statistical modeling to describe heterogeneity or improve precision.
Sequential and adaptive allocation
A sequential experiment permits later assignments to depend on information accumulated earlier. This changes the probability structure because the set of possible allocations develops over time. Adaptive clinical trials can alter allocation probabilities, discontinue treatment arms, or change enrollment according to prespecified decision rules. Their inferential structure includes both the observed responses and the adaptation mechanism that generated subsequent assignments.
Stopping an experiment after a favorable interim result changes the distribution of conventional test statistics when repeated examination is ignored. Sequential analysis represents the stopping rule directly and assigns decision boundaries across information time. The resulting design treats the timing of termination as part of the experiment rather than as an external administrative event.
Adaptive allocation differs from uncontrolled alteration. In a defined adaptive design, the rule connecting accumulated information to later decisions belongs to the original probability model. Consequently, the sequence of assignments remains reproducible as a stochastic mechanism even though the realized path depends on observations made during the experiment.
Scope and limitations
Experimental design identifies effects of assigned interventions under the conditions created by the experiment. It does not by itself establish that the same effects occur in different populations, institutional settings, or time periods. That question belongs to external validity and depends on the relationship between experimental units and the wider domain of interest.
Noncompliance also separates assignment effects from treatment-received effects. An intention-to-treat analysis compares groups according to their randomized assignments, preserving the causal contrast created by the design. Estimands based on treatment actually received require additional assumptions because receipt can depend on characteristics that also affect the outcome.
Interference presents another boundary. Standard unit-level formulations commonly associate each outcome with that unit’s own treatment, but infectious-disease interventions and network experiments can allow one unit’s assignment to alter another unit’s response. Designs for such settings define exposure through groups, contacts, or treatment saturation rather than through an isolated treatment indicator.
The broad conclusion of experimental design is that inferential validity resides in the relation among intervention, assignment, observation, and comparison. Statistical computation expresses that relation, but it cannot reconstruct an assignment structure that the experiment did not create.