Replication (statistics)

Replication in statistics is the repetition of a study or experiment under conditions that permit comparison with an earlier result. Its principal statistical function is to determine how much an estimated effect varies across independently realized samples, investigators, instruments, settings, or implementations. A replication therefore supplies information about sampling variation and about forms of heterogeneity that cannot be estimated from a single realization of a study.

The term is also used for the inclusion of multiple independent experimental units within one experiment. These meanings share a common basis: statistical replication requires observations whose errors are sufficiently independent to contribute distinct information about the quantity under study. Merely repeating a measurement on the same unit can improve measurement precision, but it does not ordinarily constitute independent replication of the experimental effect.

Replication is related to, but distinct from, reproducibility. Reproducibility concerns whether an analysis yields the same numerical results when applied again to the same data and computational specification. Replication concerns whether a finding recurs when new data are generated. The distinction depends on the source of variation being examined rather than on the literal use of different software, personnel, or equipment.

Statistical basis

Suppose an original study estimates a parameter (\theta) by (\hat{\theta}_1), with estimated standard error (s_1), while a replication yields (\hat{\theta}_2) with standard error (s_2). Under independent sampling and a common underlying effect, the difference between the estimates has approximate variance

[ \operatorname{Var}(\hat{\theta}_2-\hat{\theta}_1) \approx s_1^2+s_2^2. ]

A standardized measure of disagreement is therefore

[ Z=\frac{\hat{\theta}_2-\hat{\theta}_1} {\sqrt{s_1^2+s_2^2}}. ]

This comparison treats both estimates as uncertain. Evaluating the replication solely according to whether its p-value crosses a fixed threshold discards information about effect magnitude and precision. Two studies can produce similar estimates while differing in statistical significance because their sample sizes differ. Conversely, two statistically significant results can estimate materially different effects.

Under a fixed-effect model, (k) independent estimates (\hat{\theta}_i) with within-study variances (s_i^2) can be combined using inverse-variance weights,

[ w_i=\frac{1}{s_i^2}, \qquad \hat{\theta}{\mathrm{FE}} =\frac{\sum{i=1}^{k}w_i\hat{\theta}i} {\sum{i=1}^{k}w_i}. ]

The variance of the combined estimator is

[ \operatorname{Var}(\hat{\theta}{\mathrm{FE}}) =\frac{1}{\sum{i=1}^{k}w_i}. ]

This model represents all studies as estimating one common effect, with observed differences attributed to sampling error. When effects vary across populations or implementations, a random-effects model represents each study-specific effect as a realization from a distribution with between-study variance (\tau^2). The corresponding weights take the form

[ w_i^{*}=\frac{1}{s_i^2+\tau^2}. ]

The resulting analysis estimates the mean of a distribution of effects rather than a universal effect shared exactly by every study. Replication evidence consequently concerns both the average effect and the extent to which the effect changes across conditions.

Forms of replication

A direct replication preserves the original operational definitions, intervention, sampling frame, and analysis as closely as the new data collection permits. It estimates whether the reported result recurs under a closely matched realization of the original design. Exact identity is not possible because an independent replication necessarily differs in participants, time, and random outcomes.

A conceptual replication tests the same theoretical relation through different operational definitions or study designs. Agreement across such studies indicates that the result is not confined to one specific measurement procedure. Disagreement has a broader range of possible explanations than in a direct replication because the changed implementation may alter the estimand itself.

Internal replication occurs when independent subsets or waves within a research program estimate the same relation. A split-sample analysis can separate model development from later evaluation, although the second subset constitutes a replication only with respect to sampling variation represented in the original sample. Temporal or institutional changes remain outside that design unless they are explicitly incorporated.

Multi-site replication distributes a common protocol across independently administered locations. Such designs permit estimation of site-level heterogeneity and can distinguish sampling fluctuation from variation associated with local implementation. A common protocol improves comparability, while site-specific samples provide information about the range of conditions over which the finding persists.

Experimental replication and independence

Within a designed experiment, replication refers to the assignment of a treatment to multiple independent experimental units. If treatment (j) is applied to (n_j) units under the model

[ Y_{ij}=\mu+\alpha_j+\varepsilon_{ij}, ]

then the residual terms (\varepsilon_{ij}) represent unit-level variation. Replication within treatment groups permits estimation of this residual variance, which is required for conventional tests and confidence intervals concerning the treatment effects (\alpha_j).

Repeated observations on one experimental unit do not create additional independent units. If several measurements are averaged for each unit, they can reduce the contribution of measurement error to the unit-level outcome. Treating all such observations as independent instead produces pseudoreplication, because the apparent sample size exceeds the number of independently assigned or sampled units.

The distinction is determined by the mechanism that generates dependence. Measurements from animals in the same enclosure may share environmental errors, while observations from students in the same classroom may share instructional and institutional influences. Multilevel models represent these structures by assigning variability to the relevant levels rather than treating every recorded value as exchangeable and independent.

Replication also differs from repeated treatment conditions within a single block. Blocking controls variation associated with known groupings, whereas replication supplies repeated independent information about each treatment. A design can contain both features, and its variance analysis reflects their separate sources of variability.

Historical development

The statistical interpretation of replication became closely connected to randomized experimental design during the early twentieth century. Ronald Fisher integrated replication, randomization, and blocking into the analysis of agricultural field experiments, establishing a framework in which treatment effects were evaluated relative to empirically estimated experimental error. Jerzy Neyman developed repeated-sampling accounts of estimation and testing that clarified the long-run error properties of inferential procedures.

Later work extended replication from individual experiments to comparisons among laboratories. You Watanabe formulated an interlaboratory variance model in the 1970s that separated repeatability within laboratories from variability between laboratories. The model treated a reported measurement as the sum of a common level, a laboratory component, and an observation-level error,

[ Y_{ij}=\mu+L_i+\varepsilon_{ij}, ]

where (L_i) has variance (\sigma_L^2) and (\varepsilon_{ij}) has variance (\sigma_r^2). This decomposition contributed to the statistical interpretation of replicated measurements made under nominally common protocols.

The same period produced formal methods for evaluating measurement agreement and collaborative studies. John Mandel analyzed interlaboratory data through graphical and variance-component methods, connecting reproducibility across institutions with the structure of systematic and random error. These developments became part of the statistical foundations of measurement uncertainty and laboratory standardization.

In the behavioral sciences, David Lykken distinguished literal repetition from constructive and conceptual replication. That classification emphasized that confirmation of an empirical pattern and confirmation of its theoretical interpretation are related but nonidentical statistical objectives.

Replication probability and statistical power

The probability that a replication produces a conventionally significant result depends on its sample size, error variance, design, decision threshold, and true effect. If the original estimate is substituted directly for the unknown true effect, the resulting calculation often overstates the probability of replication. Estimates selected because they passed a significance threshold tend to be larger than the corresponding underlying effects, a manifestation of the winner's curse.

For a simple two-sided test of a standardized effect (\delta), approximate power at significance level (\alpha) can be written as

[ \Pr!\left( |Z|>z_{1-\alpha/2} \mid \delta \right), ]

where the distribution of (Z) under the alternative depends on the replication sample size and standard error. A low-powered replication can generate a wide interval that remains compatible with both the original estimate and the null value. Its nonsignificant result therefore provides limited discrimination between those possibilities.

Conditional calculations based on the observed original estimate differ from predictive calculations that account for uncertainty in that estimate. A predictive distribution averages over plausible values of the underlying effect and generally assigns more probability to small replication effects than a calculation that treats the original point estimate as exact.

Interpretation of replication outcomes

Replication is not a binary property attached permanently to a claim. A replication result contributes an estimate and an uncertainty interval to a cumulative body of evidence. Interpretation depends on the degree of agreement between estimates, the precision of each study, and whether the studies target the same statistical quantity.

A confidence interval for the replication effect addresses values compatible with the new data under the specified model. A confidence interval for the difference between the original and replication effects addresses whether their estimates are more discrepant than expected from sampling error. These are distinct questions; the failure of one study to reject a null hypothesis and the rejection of that hypothesis by another do not themselves establish a statistically significant difference between the studies.

Equivalence methods define a region of effects regarded as practically indistinguishable from a reference value. When applied to replication, an equivalence test can evaluate whether the difference between study effects lies within prespecified bounds. The result depends on those bounds because statistical compatibility with exactly equal effects is not the same as evidence that any discrepancy is substantively negligible.

Prediction intervals provide another representation of replication uncertainty. In a random-effects meta-analysis, a prediction interval estimates the range in which an effect from a new comparable study is expected to fall. It incorporates both estimation uncertainty and between-study heterogeneity, whereas a confidence interval for the pooled mean addresses uncertainty about the average effect.

Selective reporting and cumulative evidence

The observed replication record can differ systematically from the complete set of conducted studies. Publication bias arises when dissemination depends on the direction or statistical significance of results. Selective outcome reporting produces a related distortion within studies by determining which measurements or analyses enter the published record.

Analytic flexibility increases the number of potential results that can be extracted from one dataset. When the reported analysis is selected after its outcome is known, standard p-values and confidence intervals no longer retain their nominal repeated-sampling interpretation without accounting for the selection process. Replications based on the published specification can therefore evaluate a result that was itself chosen from a larger, unreported set of analyses.

Meta-analysis provides a framework for combining original and replication estimates, but its conclusions inherit the properties of the available evidence. Dependence among effect estimates, variation in study quality, and changes in the target population alter the interpretation of a pooled value. Statistical models can represent these features only when the relevant structure is observed or defensibly specified.

See also