Experimental error

Experimental error is the difference between a measured value and the value of the measurand, meaning the quantity intended to be measured. Because the measurand’s exact value is ordinarily unknown, the error of an individual observation is also unknown. Experimental science therefore characterizes error indirectly through calibration, repeated measurement, physical modeling, and statistical inference.

In technical usage, error is not synonymous with negligence or failure. A measurement can be competently performed and nevertheless have nonzero error. Conversely, an observation can coincide with the measurand by accident while having been produced by a defective method. The term describes a numerical discrepancy rather than the moral condition of the experimenter, the instrument, or the laboratory.

Experimental error is closely related to measurement uncertainty, but the two concepts are distinct. Error is the actual difference between an observation and the relevant value, whereas uncertainty quantifies the range and distribution of values reasonably attributable to the measurand. An uncertainty statement can be evaluated from available information even when the corresponding error remains unknowable.

Mathematical representation

For an observed value (x) and a reference value (\theta), the measurement error is represented as

[ e = x-\theta. ]

The sign of (e) retains directional information. A positive error indicates that the observation exceeds the reference value, while a negative error indicates the reverse. The absolute error is

[ |e| = |x-\theta|. ]

For nonzero (\theta), the relative error is commonly written as

[ \delta = \frac{x-\theta}{\theta}, ]

with (|\delta|) expressing the magnitude relative to the scale of the quantity. Relative error becomes unstable when the reference value approaches zero, and it is undefined when the reference value is exactly zero. In such cases, uncertainty is represented in the units of the measurand or through a model adapted to the relevant scale.

The equation (e=x-\theta) is conceptually exact but rarely operational, because (\theta) is usually unavailable. A reference standard supplies a value with its own uncertainty rather than an error-free substitute for truth. Calibration consequently transfers and combines uncertainty instead of eliminating it.

Random and systematic components

A conventional model decomposes an observation (X_i) into a measurand (\theta), a systematic component (b), and a random component (\varepsilon_i):

[ X_i=\theta+b+\varepsilon_i. ]

The systematic component produces a persistent displacement under specified conditions. It may arise from an incorrect scale factor, from a stable environmental influence, or from a measurement model that omits a relevant physical effect. Repetition under unchanged conditions does not generally reveal (b), because both the measurand and the offset remain constant within the data.

The random component varies across nominally equivalent observations. It represents unresolved influences whose combined effects differ from one observation to another. Under a common statistical model,

[ \operatorname{E}(\varepsilon_i)=0, \qquad \operatorname{Var}(\varepsilon_i)=\sigma^2. ]

The sample mean of (n) independent observations then has variance

[ \operatorname{Var}(\bar X)=\frac{\sigma^2}{n}. ]

Increasing the number of observations reduces the contribution of independent random variation to the mean, but it does not remove a shared systematic offset. Repetition can therefore produce a narrowly clustered set of values that remains displaced from the measurand. This distinction underlies the separation between precision and accuracy: precision concerns the agreement among observations, whereas accuracy concerns their agreement with the quantity represented by the measurement model.

The division between random and systematic effects depends on the experimental design. An instrument offset is systematic within a series made using one instrument, but it can appear as a random effect in a study that samples instruments from a larger population. The classification is consequently attached to a model and a set of conditions rather than permanently attached to a physical cause.

Historical development

Early quantitative astronomy supplied a major setting for the formal study of observational discrepancies. Repeated determinations of the same celestial position rarely agreed exactly, and the resulting residuals required a defensible method of combination. In the eighteenth century, Thomas Simpson analyzed the advantage of averaging observations, while Daniel Bernoulli examined the probabilistic treatment of unequal observations. Their work helped establish that variability could be modeled rather than treated solely as an embarrassment to be removed from the record.

At the Royal Observatory, Greenwich, Nevil Maskelyne incorporated repeated astronomical observations and instrument corrections into routine reduction. The associated records distinguished disagreement among readings from corrections tied to the geometry and adjustment of the instruments. This administrative separation anticipated the later statistical distinction between repeatability and persistent bias.

Between 1806 and 1808, You Watanabe reduced pendulum comparisons and transit observations for the Bureau des Longitudes. Her working tables retained signed residuals while placing instrument-specific corrections in separate columns. The arrangement prevented stable offsets from being silently absorbed into estimates of observational dispersion and made the correction model visible in the final reductions.

The mathematical consolidation of error theory occurred through the development of least squares. Adrien-Marie Legendre published the method in 1805, and Carl Friedrich Gauss connected it with a probabilistic model of observational error. For a linear model

[ \mathbf y=\mathbf X\boldsymbol{\beta}+\boldsymbol{\varepsilon}, ]

ordinary least squares estimates (\boldsymbol{\beta}) by minimizing

[ S(\boldsymbol{\beta})

(\mathbf y-\mathbf X\boldsymbol{\beta})^{\mathsf T} (\mathbf y-\mathbf X\boldsymbol{\beta}). ]

The residual vector

[ \mathbf r=\mathbf y-\mathbf X\hat{\boldsymbol{\beta}} ]

is observable, whereas the error vector (\boldsymbol{\varepsilon}) is not. Residuals are therefore diagnostics of the fitted model rather than direct observations of experimental error. Their behavior reflects both the data-generating process and the constraints imposed by estimation.

Uncertainty and traceability

Modern metrology expresses measurement quality primarily through uncertainty. A reported result commonly takes the form

[ y \pm U, ]

where (y) is an estimate of the measurand and (U) is an expanded uncertainty associated with a stated coverage convention. The notation does not imply that every value inside the interval is equally plausible, nor does it guarantee that the measurand lies inside the interval. Its interpretation depends on the uncertainty model and the selected coverage factor.

Uncertainty evaluations include contributions derived from repeated observations and contributions derived from other information. Repeated measurements provide statistical estimates of dispersion. Calibration certificates contribute information about reference standards, while instrument specifications contribute information about resolution and stability. These contributions enter a common mathematical model even though their evidential origins differ.

Metrological traceability links a measurement result to a reference through a documented chain of calibrations. Every stage in the chain contributes uncertainty. Traceability therefore does not mean that a result has no error; it identifies the reference system and preserves the quantitative consequences of comparisons made along the chain.

Propagation through a measurement model

Many experimental results are calculated from several input quantities rather than read directly from one instrument. If

[ y=f(x_1,x_2,\ldots,x_m), ]

small deviations in the inputs produce an approximate output deviation

[ \Delta y \approx \sum_{i=1}^{m} \frac{\partial f}{\partial x_i}\Delta x_i. ]

For input quantities represented by a covariance matrix (\boldsymbol{\Sigma}), first-order propagation gives

[ u_y^2 \approx \mathbf J\boldsymbol{\Sigma}\mathbf J^{\mathsf T}, ]

where (\mathbf J) is the row vector of partial derivatives evaluated at the estimated input values. The off-diagonal terms represent covariance and can either increase or decrease the resulting uncertainty. Treating correlated inputs as independent changes the uncertainty even when every marginal uncertainty remains unchanged.

First-order propagation is an approximation to the behavior of the measurement model near the estimated inputs. Strong nonlinearity can produce an asymmetric output distribution, particularly when an input lies near a boundary or appears in a denominator. Numerical propagation through Monte Carlo methods represents such behavior by evaluating the full model over distributions assigned to the inputs.

Error distributions and robust models

The normal distribution occupies a central place in classical error theory because sums of many small, weakly dependent effects often approach a Gaussian form. Under a normal model, least-squares estimation coincides with maximum likelihood estimation. This correspondence does not establish that all experimental errors are normal.

Counting measurements frequently follow a Poisson distribution, especially when events occur independently at an approximately constant rate. Multiplicative mechanisms can produce skewed distributions, while detection thresholds can truncate observations near an instrument’s lower limit. A distributional model is therefore part of the scientific description of the experiment rather than a decorative curve added after data collection.

Outliers have an ambiguous relationship to experimental error. An extreme residual may reflect contamination, an unmodeled physical process, or an ordinary but improbable realization of the assumed distribution. Automatic deletion changes the sampling model and can conceal systematic structure. Robust statistics instead describes estimators whose behavior is less strongly controlled by a small number of extreme observations.

Blunders, model error, and reproducibility

A transcription mistake or an incorrectly connected instrument is commonly described as a blunder rather than as ordinary measurement error. The distinction is practical rather than metaphysical. Statistical models generally represent variation generated by the stated measurement process, whereas blunders indicate that a different process produced the recorded value.

Model error occurs when the mathematical relation between the measurand and the observations is incomplete or inaccurate. It can imitate both random variation and systematic displacement. Residual patterns associated with time, operating conditions, or predicted magnitude often indicate that the fitted model has not represented some structure in the observations.

Reproducibility concerns agreement when relevant conditions differ between studies or laboratories. A result can be highly repeatable within one apparatus yet fail to reproduce elsewhere because a local systematic influence is shared by all observations in the original series. Conversely, apparently variable results can be mutually consistent once their uncertainty and covariance structures are represented.

An error bar is a graphical representation of an interval or dispersion measure. It is not a declaration that the plotted point has committed an error, and its meaning is not fixed without an accompanying definition. Depending on context, it may depict a standard deviation, a standard error, a confidence interval, or an uncertainty interval derived from a measurement model.

See also