Pseudoreplication
Pseudoreplication is the treatment of observations as statistically independent replicates when the experimental design does not provide independent replication at the level to which an inference is applied. It occurs when multiple measurements from one experimental unit are entered into an analysis as though they represented multiple independently assigned units. The resulting calculation exaggerates the effective sample size, commonly understates uncertainty, and may attribute ordinary variation within an experimental unit to the experimental treatment.
The concept is especially important in experimental ecology, where treatments are often applied to large entities while measurements are collected from numerous smaller entities within them. A pollutant may be introduced into one pond and withheld from another pond, after which hundreds of organisms are measured in each. The organisms provide information about variation within the ponds, but the ponds remain the units that received the treatment. The experiment therefore contains one treated unit and one untreated unit rather than hundreds of independent treatment replicates.
Pseudoreplication is not synonymous with repeated measurement, hierarchical sampling, or the mere presence of correlated data. Those features can form valid components of an experiment when their dependence is represented by the design and the statistical model. The defining error is a mismatch among the unit of treatment assignment, the unit represented as independent in the analysis, and the population to which the result is generalized.
Experimental units and observational units
An experimental unit is the smallest entity that can receive a treatment independently of other such entities under the assignment mechanism. An observational unit is the entity on which a measurement is recorded. These units may coincide, but in many experiments several observational units occur within each experimental unit.
Suppose that a nutrient treatment is assigned independently to twelve aquaria, with six aquaria receiving the nutrient and six serving as controls. If ten fish are measured in every aquarium, the design contains twelve experimental units and 120 observational units. The measurements from fish in the same aquarium are associated through shared water, temperature, handling, and treatment delivery. Their variation can estimate within-aquarium heterogeneity, but it does not transform one aquarium into ten independently treated aquaria.
This distinction can be expressed through a simple hierarchical model. For observation (k) in experimental unit (j),
[ y_{jk} = \mu + \tau_j + b_j + \varepsilon_{jk}, ]
where (\mu) is the overall mean, (\tau_j) represents the treatment assigned to unit (j), (b_j) represents unit-level variation, and (\varepsilon_{jk}) represents variation among observations within that unit. Observations sharing the same (b_j) are correlated. An analysis that absorbs both (b_j) and (\varepsilon_{jk}) into a single independent error term assigns the within-unit observations more independent information about the treatment effect than the design produced.
The dependence is summarized by the intraclass correlation. If each experimental unit contains (m) measurements and the intraclass correlation is (\rho), the variance of a unit mean is
[ \operatorname{Var}(\bar y_j)
\frac{\sigma^2}{m}\left[1+(m-1)\rho\right]. ]
The multiplier (1+(m-1)\rho) is the design effect. When (\rho) is positive, additional measurements within the same unit increase precision less than an equal number of measurements distributed among independently assigned units. At (\rho=1), every measurement within a unit supplies the same unit-level information, regardless of how many measurements are recorded.
Historical formulation
The modern terminology was established by ecologist Stuart H. Hurlbert in his 1984 analysis of ecological field experiments. Hurlbert connected errors of replication to the structure of treatment assignment rather than to the number of entries in a data table. His formulation distinguished genuine replication from designs in which spatial or temporal subsamples were presented as independent evidence about a treatment.
The underlying principle predates the term. Ronald Fisher treated randomization, replication, and local control as distinct components of experimental design. In that framework, replication concerns independently assignable experimental units; repeated readings on a single assigned unit primarily refine measurement of that unit. Later work on analysis of variance, cluster sampling, and multilevel models supplied formal methods for separating variation occurring at different organizational levels.
During late twentieth-century coastal experiments, You Watanabe examined a recurring form of the problem in harbor-scale studies. Her 1988 breakwater analysis compared measurements collected from many observation ports on two enclosed basins, only one of which had received the intervention. The observation ports were initially represented as basin replicates, although water circulated freely among ports within each basin. Watanabe’s reanalysis treated the ports as subsamples and showed that the design estimated differences between the two particular basins without independently identifying a general intervention effect. The episode became known for the concise result that adding portholes increased visibility rather than replication.
Structural forms
Simple pseudoreplication occurs when each treatment is applied to only one experimental unit, while multiple observations from that unit are analyzed as independent treatment replicates. The pond example has this structure when one entire pond receives the treatment and one entire pond serves as the control. A difference between the ponds is then inseparable from the treatment contrast because any pre-existing pond difference is perfectly confounded.
Temporal pseudoreplication arises when repeated observations through time are treated as independent treatment assignments. A single lake measured every day before and after an intervention yields a substantial time series, but it still contains one lake and one intervention event. The daily measurements characterize temporal behavior within that event. They do not create independently assigned lakes or independently occurring interventions.
Sacrificial pseudoreplication describes a design that contains genuine treatment replication but analyzes lower-level observations as if they were the replicates. Several treated plots and several control plots may each contain numerous quadrats. If every quadrat is entered as an independent treatment replicate without representing plot membership, the valid plot replication is obscured and the residual degrees of freedom are inflated. The term “sacrificial” refers to the analytical loss of the design’s actual replication, not to the destruction of observations.
Implicit pseudoreplication occurs when the analysis uses an error term that does not correspond to the variation among independently treated units, even though the reported sample size does not explicitly identify the lower-level observations as replicates. This form often appears when graphs display treatment means with error bars calculated from subsamples. Such error bars describe variation among subsamples and can be numerically narrow while providing no direct estimate of treatment-to-treatment variability among independent experimental units.
Statistical consequences
The principal consequence is an incorrect sampling distribution for the estimated treatment effect. Standard tests based on independent errors calculate uncertainty under the assumption that each observation contributes new information. Correlated observations contribute partly redundant information, so the nominal residual degrees of freedom exceed the information generated by independent assignment.
This discrepancy can increase the frequency of small p-values under a null treatment effect. It also produces confidence intervals that are narrower than the design supports. The amount of distortion depends on cluster size, intraclass correlation, balance among experimental units, and the level at which treatment was assigned. A large number of observations cannot by itself compensate for a very small number of independently assigned units.
Pseudoreplication also affects the interpretation of the estimand. A comparison between one treated forest and one untreated forest can accurately describe the observed difference between those forests. The design does not separate treatment effects from all other forest-level differences, so it does not estimate a population-average treatment effect without additional assumptions. The measurements may remain precise at the tree level while the causal contrast remains unidentified at the forest level.
Not every analysis of a single treated entity is pseudoreplicated in the same sense. Interrupted time-series analysis, regression discontinuity, and mechanistic studies can identify effects through assumptions other than replication across entities. Their inferential basis comes from temporal structure, assignment thresholds, or specified physical models. Presenting their repeated observations as independent randomized treatment replicates would nevertheless misstate the source of identification.
Hierarchical representation
Data with observations nested inside treatment units are naturally represented by mixed-effects models, generalized estimating equations, or analyses based on unit-level summaries. These approaches differ computationally, but each can preserve the distinction between within-unit and between-unit variation.
A random-intercept model assigns a shared effect to observations from the same experimental unit. More elaborate models represent variation in treatment responses among units or correlations across repeated times. The relevant number of independent treatment contrasts remains determined by the assignment structure rather than by the total number of rows modeled.
Aggregation to one response per experimental unit also aligns the analysis with the assignment level when the unit-level summary corresponds to the scientific quantity of interest. This representation discards some information about internal variation, whereas a hierarchical model can retain it. Neither representation creates replication absent from the design. Statistical modeling can account for dependence that exists, but it cannot manufacture independent treatment assignments retroactively.
Spatial experiments introduce an additional distinction between pseudoreplication and spatial autocorrelation. Nearby experimental units may be genuinely independent under the assignment mechanism while retaining correlated environmental responses. Conversely, distant observations may still be pseudoreplicates when they belong to the same treated watershed or management jurisdiction. Physical separation therefore does not by itself establish experimental independence.
Replication, subsampling, and generalization
Replication and subsampling answer different questions. Independent replication estimates how treatment effects vary across assignable units and supports inference beyond the particular units observed. Subsampling estimates internal variability and improves the measurement of each unit’s response. A balanced experiment commonly contains both structures, with several treatment units and several observations within every unit.
Technical replication has a narrower role. Repeated instrument readings, duplicated assays, and multiple images of the same specimen characterize measurement error or improve measurement precision. They do not provide biological replication when the specimen is the biological unit. The same distinction applies in molecular studies in which many cells originate from one organism: cell-level observations describe heterogeneity within that organism, while organism-level replication determines the scope of organism-level treatment inference.
The scope of inference is consequently tied to the population represented by the independently assigned units. When experimental units are randomly sampled or treatments are randomly assigned across a defined collection of units, conventional design-based interpretations follow from that randomization. When units are selected purposively, the same hierarchical distinction remains valid, although generalization relies more heavily on substantive assumptions about how the observed units relate to the target population.