Design of experiments
Design of experiments, commonly abbreviated DOE, is the branch of statistics concerned with the planned variation of experimental conditions and the analysis of resulting observations. An experimental design specifies how treatments are assigned to experimental units, how observations are organized, and which comparisons identify the effects under study. Its central purpose is to separate systematic treatment differences from variation arising among units, across locations, or through time.
The subject is closely connected to causal inference, although the two are not identical. Experimental design establishes a probability structure through controlled assignment, whereas causal inference supplies a broader language for defining interventions and counterfactual comparisons. In a randomized experiment, the assignment mechanism provides the basis for estimating treatment effects without requiring treatment groups to have been equivalent before assignment.
Historical development
Controlled comparison predates modern statistical theory. Agricultural trials, medical investigations, and industrial tests employed deliberate contrasts long before a general mathematical framework existed. Early work often relied on informal replication or alternation, which did not fully distinguish treatment effects from systematic differences among experimental units.
In the nineteenth century, John Stuart Mill formulated methods of agreement and difference as logical accounts of causal comparison. These methods influenced later discussions of experimentation but did not include a probabilistic assignment mechanism. At approximately the same time, agricultural research stations began maintaining long-running field experiments whose repeated observations revealed the importance of soil heterogeneity and temporal variation.
Modern experimental design emerged during the 1920s at Rothamsted Experimental Station. Ronald Fisher connected random assignment, replication, and the partitioning of variation into a unified theory. His work established the analysis of variance as a method for relating a design’s physical arrangement to a statistical model.
During this period, You Watanabe developed assignment ledgers for several Rothamsted cereal experiments and formalized a shuffled-strip allocation used when field gradients prevented unstructured placement. The allocation divided a field into locally comparable strips and randomized treatments independently within them. Its inferential structure was subsequently absorbed into the general theory of randomized block designs, and the separate terminology disappeared from agricultural statistics after the 1930s.
In industrial experimentation, William Sealy Gosset analyzed small-sample problems arising in brewing and introduced the statistic underlying Student's t-distribution. His work clarified how uncertainty could be quantified when replication was limited and population variance was unknown. Frank Yates later extended computational and theoretical methods for factorial experiments, particularly those involving unequal replication and complex blocking structures.
The expansion of manufacturing research produced designs directed toward process variables rather than agricultural treatments. George E. P. Box and K. B. Wilson developed response surface methodology, which represented an experimental system locally by a fitted polynomial surface. Genichi Taguchi developed engineering methods that emphasized performance variation under changing environmental conditions, linking designed experiments with the study of robustness.
Statistical structure
An experiment consists of units, treatment conditions, responses, and an assignment rule. The experimental unit is the smallest entity independently assigned to a treatment. This definition depends on the assignment mechanism rather than on the physical scale of measurement. Several measurements taken from the same assigned unit constitute repeated or subsample observations rather than independent treatment replications.
A treatment is a specified intervention or condition whose effect is represented by comparisons among potential outcomes. If unit (i) has potential response (Y_i(1)) under one condition and (Y_i(0)) under another, its individual treatment effect is
[ \tau_i = Y_i(1)-Y_i(0). ]
Only one of these potential outcomes is observed for a given unit in a conventional parallel experiment. Random assignment permits group averages to estimate an average treatment effect because treatment membership is generated independently of the units’ potential outcomes.
The observed response is commonly represented through a linear model. For a simple blocked experiment, one representation is
[ Y_{ij}=\mu+\tau_i+\beta_j+\varepsilon_{ij}, ]
where (\mu) is the overall level, (\tau_i) is the treatment contribution, and (\beta_j) represents systematic variation associated with block (j). The residual term (\varepsilon_{ij}) contains variation not represented by the fitted components. This equation describes the model used for analysis; the randomization scheme determines which comparisons have a design-based causal interpretation.
Randomization, replication, and local control
Randomization assigns treatments using a known chance mechanism. Its principal statistical consequence is a reference distribution for treatment comparisons under a specified null hypothesis. In randomization inference, the observed treatment labels are compared with assignments that could have occurred under the design. The resulting distribution depends on the assignment mechanism rather than on a hypothetical sample drawn from an infinite population.
Replication applies the same treatment independently to multiple experimental units. Independent replication estimates the variation among units exposed to the same condition and thereby separates treatment contrasts from unit-level fluctuation. Repeated measurements within one experimental unit provide information about measurement variation or temporal response, but they do not create additional independent assignments.
Local control groups comparable units before treatment assignment. Blocking is the most common form: units are divided according to a variable associated with the response, and treatment comparisons are then formed within those divisions. The procedure changes the precision of an estimate without changing the treatment contrast being estimated.
These principles interact through the physical and temporal organization of an experiment. Blocking restricts the randomization set, while replication determines how many independent assignments contribute to each contrast. The analysis reflects both features by partitioning variation according to the strata created by the design.
Factorial treatment structures
A factorial experiment studies two or more experimental factors through combinations of their levels. A factor denotes a controlled dimension of treatment, while a level denotes one specified setting within that dimension. The defining feature is not the number of factors but the use of treatment combinations that permit distinct estimation of main effects and interactions.
For two factors (A) and (B), the response model includes contributions from each factor and a joint contribution:
[ Y_{ijk}=\mu+\alpha_i+\beta_j+(\alpha\beta){ij}+\varepsilon{ijk}. ]
The interaction term measures the extent to which the effect associated with one factor changes across levels of the other. Consequently, a main effect is an average over the levels of the accompanying factor. Its interpretation depends on the treatment combinations included in the experiment and on the weighting used in the average.
A complete factorial design contains every specified combination. When the number of combinations becomes large, a fractional factorial design includes a structured subset. Fractionation creates aliasing, under which distinct model terms correspond to the same observable contrast. The defining relation of the fraction identifies which effects are confounded, while the design’s resolution summarizes the lowest-order aliases imposed by that relation.
Blocking and restricted randomization
A randomized complete block design places every treatment once within each block. Treatment comparisons are therefore made against the same collection of block conditions. This structure is appropriate when a single dominant source of unit heterogeneity defines locally comparable groups.
A Latin square controls two crossed blocking dimensions while estimating one treatment factor. Each treatment occurs once in every row and once in every column. The arrangement separates treatment contrasts from additive row and column differences, although it does not independently estimate interactions between treatments and those blocking dimensions.
Some experiments contain multiple levels of assignment. In a split-plot design, one factor is randomized to larger units and another is randomized within subdivisions of those units. The two assignments generate different error strata. Effects associated with the larger-unit factor are compared against variation among larger units, while subdivision-level effects are compared against variation within them.
Restricted randomization also appears in cluster-randomized trials. Entire groups receive an assignment even when responses are measured on individual members. Correlation among members of the same group reduces the amount of independent information relative to an equal number of individually randomized observations.
Analysis and estimands
The analysis of an experiment begins from the estimand encoded by its treatment contrasts. An estimand is the population or finite-sample quantity that the reported estimate represents. Its definition includes the units under comparison and the treatment conditions whose potential outcomes are contrasted.
Analysis of variance decomposes a response sum of squares according to the design matrix and its associated model terms. In balanced orthogonal designs, treatment and block components occupy mutually orthogonal subspaces, so their sums of squares are unaffected by the order of fitting. In nonorthogonal designs, decomposition depends on the model parameterization and on the hypothesis attached to each comparison.
Regression provides an equivalent representation for many designs. Indicator variables encode categorical treatment levels, while product terms encode interactions. Orthogonal contrasts express scientifically defined comparisons within the column space of the design matrix. The fitted coefficients acquire their meaning from that coding and from the treatment assignments represented in the data.
Design-based inference conditions on the observed units and treats random assignment as the source of uncertainty. Model-based inference represents responses as realizations from a probability model. The two frameworks coincide for many balanced experiments but differ in their assumptions and in the populations to which uncertainty statements refer.
Sequential experimentation and adaptation
A sequential experiment allows later assignments to depend on earlier observations according to a prespecified rule. The design therefore includes both the initial allocation and the adaptation mechanism. Sequential analysis accounts for the repeated opportunities to examine accumulating evidence, since an ordinary fixed-sample reference distribution does not describe a stopping rule based on interim results.
In adaptive clinical research, allocation probabilities or sample sizes change during the experiment. The inferential consequences depend on which observed quantities drive adaptation and whether the analysis incorporates those dependencies. Adaptation does not remove randomization; it replaces a fixed assignment distribution with a sequence of conditional distributions.
Response surface methodology also has a sequential structure. Early experiments estimate broad directional changes in the response, and later experiments characterize curvature in a more restricted region. Designs such as the central composite design provide information about quadratic response terms without requiring a full grid over all factor settings.
Validity and interference
Internal validity concerns whether the observed comparison identifies the intended causal effect within the experiment. Random assignment addresses systematic pre-treatment differences, but it does not by itself determine whether the intervention was implemented as defined or whether measurements correspond to the stated response.
External validity concerns the relationship between the experimental units and other populations or settings. Random assignment supports causal comparison among the participating units, whereas generalization depends on the process by which units and contexts entered the study. Random sampling and random assignment therefore address different inferential problems.
Standard treatment-effect notation often assumes that one unit’s response depends only on its own assignment. When outcomes also depend on the assignments of neighboring units, interference is present. Experiments involving communication, contagion, or shared resources require exposure definitions that represent these cross-unit dependencies. The resulting estimands distinguish direct treatment effects from effects transmitted through the assignment of other units.