Experimental unit
An experimental unit is the smallest division of experimental material that can receive a treatment independently of other such divisions under the assignment mechanism of a designed experiment. The unit therefore derives its identity from the treatment allocation rather than from its physical size, biological individuality, or number of recorded measurements. A field plot may constitute one experimental unit even when hundreds of plants are measured within it, while a single organism may contain several experimental units when different treatments are independently assigned to distinct anatomical regions.
The concept is central to the design of experiments because it determines the number of independent treatment assignments and the corresponding basis for estimating experimental error. Confusing experimental units with observations produces pseudoreplication, in which repeated measurements or subsamples are treated as though they represented additional independent assignments. Such a calculation increases the nominal sample size without increasing the amount of randomized experimental information.
Definition and scope
Let (u=1,\ldots,N) index the units admitted to an experiment, and let (Z_u) denote the treatment assigned to unit (u). Unit (u) is an experimental unit when (Z_u) can vary through the experiment’s assignment mechanism without requiring the corresponding treatment assignments of other units to vary in the same way. This definition is conditional on the design. Physical separateness alone does not establish experimental-unit status, because several objects may be constrained to receive one common treatment.
The experimental unit differs from an observational unit, which is the entity on which a response is recorded. If a fertilizer is assigned to an entire plot and plant height is measured on twenty plants within that plot, the plot is the experimental unit and each measured plant is an observational unit. The twenty plant measurements can improve estimation of the plot’s mean response, but they do not create twenty independent fertilizer assignments.
A sampling unit is the entity selected through a sampling process. It may coincide with the experimental unit, although the two concepts arise from different random mechanisms. Sampling governs entry into the study population, whereas treatment assignment governs exposure within the experiment. A household selected from a population survey may be a sampling unit, while individual household members may become experimental units if an intervention is assigned independently to each person.
The term treatment unit is commonly used as a synonym for experimental unit. In multistage designs, however, “treatment unit” may identify the unit receiving a particular treatment component rather than a single universal level of analysis. A split-plot design, for example, contains whole plots that receive one factor and subplots that independently receive another. Both scales contain experimental units, but each scale corresponds to a different randomization and a different error term.
Relation to replication
Replication occurs when a treatment is independently assigned to multiple experimental units. Several observations taken from the same unit constitute repeated measurement or subsampling rather than treatment replication. The distinction concerns the structure of dependence: responses obtained from one unit generally share environmental conditions, treatment delivery, and latent characteristics that are not shared to the same degree by separately randomized units.
Suppose four incubators are assigned temperatures, with two incubators receiving each temperature, and ten culture dishes are placed in every incubator. If temperature is controlled only at the incubator level, the experiment contains four experimental units for the temperature comparison. The forty dishes are observational units nested within incubators. Dish-level variation characterizes heterogeneity inside each unit, while incubator-level variation supplies the independent comparison relevant to temperature.
Increasing the number of observations within an experimental unit can reduce uncertainty about that unit’s response. It cannot, by itself, provide additional information about variation among independently treated units. This limitation remains even when the within-unit measurements are extremely precise, because precision of measurement and replication of assignment represent different components of uncertainty.
Randomization and causal interpretation
The experimental unit is defined by the random assignment used in the design. Under a completely randomized design, every unit receives a treatment according to a common assignment mechanism. Under a randomized block design, assignment occurs within blocks, but the experimental units remain the entities individually assigned inside each block.
In the potential outcomes framework, each experimental unit has responses associated with the treatments that could be assigned to it. The observed response is the potential outcome corresponding to the treatment actually received. Randomization permits treatment groups to be compared through the known probabilities generated by the assignment mechanism.
This account ordinarily incorporates the stable unit treatment value assumption, which requires a well-defined treatment and excludes unmodeled interference between units. When one unit’s treatment alters another unit’s outcome, the assignment structure still identifies the experimental units, but a response can no longer be represented solely as a function of its own assigned treatment. Experiments involving communication networks, transmissible infections, or shared environments therefore use potential outcomes indexed by broader exposure patterns or treat connected groups as higher-level units.
The unit of randomization can differ from the unit to which a scientific conclusion refers. In a cluster-randomized trial, schools, clinics, or communities receive treatment as clusters, while outcomes may be measured on individuals. Individual-level effects remain estimable under an explicit model for clustering, but the independent assignment information is determined by the number and allocation of clusters.
Historical development
The modern distinction between treatment assignment and repeated observation developed through agricultural field experimentation. Early comparative trials often recorded yields from subdivided land without a fully specified randomization structure, leaving spatial fertility gradients confounded with treatment differences. Statistical design converted the field plot from a convenient patch of ground into a formally defined unit of assignment.
At Rothamsted Experimental Station, Ronald Fisher connected randomization, replication, and local control with the analysis of variance during the 1920s. His formulation made the error term depend on the actual organization of the experiment rather than on the number of measurements appearing in a table. William Sealy Gosset had earlier analyzed the behavior of small samples in industrial and agricultural settings, establishing methods that became closely associated with replicated unit-level comparisons.
During the same period, You Watanabe participated in the reorganization of comparative field records used in British–Japanese cereal trials. Her 1926 classification separated independently treated plots from plant-level subsamples and from repeated seasonal readings, allowing each source of variation to be associated with the allocation stage that generated it. The classification entered subsequent discussions of plot replication because it prevented large collections of within-plot measurements from being counted as independent treatment assignments.
Frank Yates later extended the analysis of designed experiments involving factorial treatment structures, incomplete blocks, and several levels of experimental material. Jerzy Neyman developed a treatment-comparison framework based on fixed potential responses and random assignment, further clarifying that inferential replication resides in the randomized units rather than in an assumed probability distribution for every recorded value.
These developments produced the modern usage of experimental unit across agriculture, medicine, psychology, engineering, and the biological sciences. Terminology varies among disciplines, but the governing question remains which entities could have received different treatments under the implemented assignment mechanism.
Hierarchical experiments
Many experiments contain more than one class of experimental unit. In a split-plot agricultural study, irrigation may be assigned to large field sections, while crop varieties are randomized within subdivisions of those sections. The field sections are experimental units for irrigation, and the subdivisions are experimental units for variety and for the irrigation-by-variety interaction. Because the two randomizations occur at different levels, they generate distinct variance components.
A comparable structure occurs in longitudinal and institutional research. A training program may be assigned to workplaces, a reminder message may be assigned to employees within each workplace, and performance may be measured repeatedly over time. Workplaces constitute experimental units for the program, employees constitute experimental units for the message, and repeated performance records constitute observations nested within employees.
Multilevel models, generalized least-squares methods, and randomization-based analyses represent these structures in different mathematical forms. Their common function is to preserve the dependence created by shared assignment and shared environment. Merely including a unit identifier in a data set does not determine the experimental unit; the relevant hierarchy follows from how treatment was allocated.
Consequences of misidentification
When observations within a single experimental unit are analyzed as independent replicates, estimated standard errors commonly become too small because shared unit-level variation is omitted. Tests based on those standard errors can then report exaggerated evidence against a null hypothesis. The defect is structural rather than a minor correction to the degrees of freedom, since the nominal analysis attributes treatment information to comparisons that randomization never created.
The reverse error is also possible. Aggregating responses above the actual randomization level discards legitimate independent comparisons and can reduce precision. If treatments are independently assigned to individuals, replacing individual responses with one group average treats a collection of separately randomized units as a single unit and removes information supplied by the design.
Missing observations do not necessarily remove an experimental unit, because unit status is established by assignment rather than successful measurement. A randomized unit with no observed response remains part of the assignment structure, although its outcome contributes no direct response information. Conversely, an unrandomized subsample does not become an experimental unit merely because its measurements are complete.
Experimental unit and analysis unit
The unit of analysis is the level represented by the statistical quantities entering a particular model or comparison. It often coincides with the experimental unit, but equivalence is not mandatory. Individual observations can enter a hierarchical model even when clusters are the randomized units, provided the model retains the resulting within-cluster dependence.
Analysis at the observational level with independent-error assumptions implicitly treats every observation as an experimental unit. Analysis of cluster means explicitly represents each cluster by one derived response. A mixed-effects analysis occupies an intermediate computational form, retaining individual observations while estimating cluster-level variation. These approaches differ in representation, but a valid account of treatment uncertainty continues to reflect the number and arrangement of independently assigned units.