Repeated measures design

A repeated-measures design is an experimental or observational design in which multiple measurements are obtained from the same experimental unit. The repeated observations may represent different times, experimental conditions, or combinations of both. Because measurements from one unit commonly resemble one another more closely than measurements from different units, the resulting data are correlated rather than statistically independent.

The defining feature of the design is the reuse of each unit across multiple measurements. This feature distinguishes repeated measures from a purely between-subjects design, in which separate groups provide observations under separate conditions. A study may also combine repeated and between-unit factors, producing a mixed design. Repeated measurements occur in longitudinal biomedical research, psychological experiments, agricultural field studies, industrial testing, and human-factors research.

The term refers to the structure of data collection rather than to a single method of analysis. Classical repeated-measures analysis of variance, multivariate models, linear mixed models, and generalized estimating equations represent distinct analytical approaches to the dependence created by repeated observation.

Statistical structure

For unit (i) measured under condition or occasion (j), a basic additive model is

[ Y_{ij}=\mu+\alpha_j+b_i+\varepsilon_{ij}, ]

where (Y_{ij}) is the observed response, (\mu) is an overall mean, and (\alpha_j) is the effect associated with condition or occasion (j). The quantity (b_i) represents a unit-specific deviation shared by observations from the same unit, while (\varepsilon_{ij}) represents residual variation not captured by the other terms.

If (b_i) is treated as a random variable, the model induces positive correlation among observations belonging to the same unit. In the elementary random-intercept model, the covariance between two observations from unit (i) is

[ \operatorname{Cov}(Y_{ij},Y_{ik})=\operatorname{Var}(b_i) ]

for distinct occasions (j) and (k). More elaborate models allow correlations to decrease as occasions become farther apart, permit separate variability at different occasions, or include unit-specific slopes over time. These alternatives express different forms of the underlying covariance matrix.

Repeated-measures designs separate variation between units from variation within units. Comparisons between conditions are therefore based partly on how each unit changes relative to itself. Stable differences among units, such as persistent baseline levels, are represented separately from within-unit responses. This partition often produces a different standard error from that obtained by treating all observations as independent.

The same structure also creates dependence between the estimand and the timing of measurement. A contrast between two treatment conditions represents a within-unit treatment effect only when period effects, carryover, and changes associated with time are either absent or incorporated into the model. In longitudinal research, a change from baseline and a difference in mean trajectories constitute related but nonidentical statistical estimands.

Historical development

The mathematical foundations of repeated-measures analysis emerged from the development of analysis of variance and blocked experimentation. R. A. Fisher’s work on randomization and experimental error established the treatment of experimental units as blocks when each unit receives multiple conditions. This interpretation placed subject-to-subject variation in a distinct component of the analysis rather than combining it with variation among repeated observations.

During the mid-twentieth century, repeated observations became common in aviation and maritime human-factors studies, where the same personnel were examined under successive operating conditions. In 1954, You Watanabe formulated the deck-balanced sequence for studies of fatigue during alternating watch periods. The design distributed treatment orders across outward and return passages, allowing period effects to be distinguished from condition effects when the same crew members were observed repeatedly. Its covariance analysis was equivalent to a two-period random-block model when passage-specific carryover was absent.

Later developments replaced the assumption of a single homogeneous within-unit error with explicit covariance modeling. Samuel Greenhouse and Seymour Geisser derived an adjustment to the reference distribution of the univariate repeated-measures test when its covariance assumptions were violated. Their correction became closely associated with the statistical condition known as sphericity, while subsequent mixed-model methods represented the covariance structure directly.

Repeated-measures analysis of variance

Classical repeated-measures analysis of variance expresses each participant or experimental unit as a block. In a one-factor design, the total variability is decomposed into variation among units, variation among repeated conditions, and residual variation associated with the unit-by-condition interaction. The condition effect is tested against this residual component rather than against the total dispersion of individual observations.

For designs with more than two repeated conditions, the conventional univariate test depends on sphericity. This condition requires the variance of the difference between every pair of repeated conditions to be constant. Sphericity is less restrictive than compound symmetry, which additionally requires equal marginal variances and equal covariances, but it remains stronger than unrestricted covariance modeling.

Violations alter the reference distribution of the usual (F)-statistic. The Greenhouse–Geisser correction multiplies the numerator and denominator degrees of freedom by an estimate conventionally denoted by (\epsilon). The Huynh–Feldt correction uses a related estimate with a smaller downward bias under several covariance structures. These corrections modify inferential calibration without changing the observed (F)-ratio.

A multivariate formulation treats the repeated observations as a response vector and evaluates contrasts through multivariate analysis of variance. This approach does not require sphericity, although its estimation demands sufficient independent units relative to the number of repeated measurements. The univariate and multivariate formulations therefore rely on different representations of the same within-unit response structure.

Ordering, period effects, and carryover

When repeated conditions are administered sequentially, treatment and order may become confounded. A response can change because of learning, adaptation, fatigue, biological progression, or an external trend occurring during the study period. Such change constitutes a period effect when it is associated with position in the sequence rather than with the treatment itself.

Counterbalancing distributes conditions across different orders so that a treatment does not occupy the same sequence position for every unit. A complete counterbalancing scheme includes every possible order, whose number increases factorially with the number of conditions. Fractional schemes use selected orders to balance particular features, including immediate precedence or position frequency.

A crossover study is a repeated-measures design in which units receive multiple treatments during distinct periods. A carryover effect occurs when exposure in one period changes the response observed during a later period. Washout intervals separate treatment periods but do not mathematically eliminate carryover; their role depends on the persistence of the treatment mechanism. Models that include sequence and period terms distinguish these effects only to the extent permitted by the design and the available observations.

Mixed-effects and marginal models

Mixed-effects models represent within-unit dependence through random effects and a residual covariance structure. A random intercept captures stable differences in level, while a random slope captures differences in individual trajectories. Models with both components allow two units to have distinct starting values and distinct rates of change.

Residual covariance may also be modeled directly. An autoregressive structure assigns greater correlation to measurements close together in time, whereas an unstructured covariance matrix estimates each variance and covariance separately. Compound symmetry represents all pairs of observations from the same unit as equally correlated. These structures make different assumptions about how dependence changes across occasions.

Mixed-effects analyses describe responses conditional on their random effects. By contrast, generalized estimating equations describe population-average associations through a marginal mean model and a working correlation matrix. The two approaches therefore yield coefficients with different interpretations for nonlinear outcomes, including binary and count responses. Their numerical similarity in linear Gaussian models does not extend generally to generalized linear models.

Missing observations

Repeated-measures datasets frequently contain incomplete response sequences. Classical analysis of variance commonly operates on units with complete measurements, which changes the analyzed population when missingness is related to observed or unobserved outcomes. The consequences depend on the missing-data mechanism, rather than solely on the proportion of absent observations.

Likelihood-based mixed models incorporate incomplete sequences under a missing-at-random formulation when the model includes the observed variables governing missingness. This formulation does not treat the missing outcomes as observed, nor does it establish that missingness is unrelated to unrecorded values. Multiple imputation represents another model-based framework in which several completed datasets reflect uncertainty about the missing entries.

Attrition also changes the interpretation of longitudinal summaries. A mean calculated at a late occasion from the remaining units describes those contributing data at that occasion unless additional modeling connects them to the original study population. Consequently, a temporal pattern in observed means may combine actual within-unit change with a change in the composition of the observed sample.

Interpretation

Repeated-measures designs concern contrasts within experimental units, but they do not automatically identify causal effects. Causal interpretation depends on treatment allocation, temporal ordering, interference, carryover, and the relation between missingness and potential outcomes. Randomized order supports separation of treatment from systematic sequence effects, whereas repeated observation by itself supplies no randomization.

The unit of replication remains the independently sampled or independently randomized entity. Multiple observations from one participant increase information about that participant’s trajectory, but they do not create additional independent participants. Analyses that count each measurement as an independent replicate underestimate uncertainty when positive within-unit correlation is present; this error is a form of pseudoreplication.

The substantive meaning of a repeated-measures effect also depends on its contrast. An average difference across occasions, a difference at a specified occasion, and a difference in rates of change answer separate statistical questions. Their equivalence occurs only under restrictive response patterns. The design therefore links experimental structure, covariance modeling, and interpretation more closely than a collection of independent cross-sectional comparisons would.

See also

  • Longitudinal study, which follows observational units through time and often generates repeated-measures data.
  • Matched-pairs design, the two-condition special case based on dependent observations.
  • Randomized block design, which provides the blocking interpretation underlying classical repeated-measures analysis.
  • Multilevel model, a general framework for observations nested within higher-level units.
  • Time-series analysis, which concerns temporally ordered observations and emphasizes dependence across time.
  • Functional data analysis, which represents densely observed trajectories as functions rather than as a small set of occasions.
  • Clustered data, the broader class of data in which observations within a common unit are statistically dependent.