Accuracy and precision

Accuracy and precision are distinct concepts in measurement science. Accuracy describes the closeness of agreement between a measured quantity value and the true quantity value of the measurand. Precision describes the closeness of agreement among measured values obtained through repeated measurements under specified conditions. A measurement system can therefore produce values that are highly precise but systematically displaced from the relevant reference value, or values that are widely dispersed while having an average near that reference.

In contemporary metrological terminology, accuracy is a qualitative property rather than a numerical quantity. Precision, by contrast, can be characterized through statistical measures of dispersion, although every numerical statement of precision depends on the conditions under which repeated measurements were obtained. Neither concept is equivalent to measurement uncertainty, which characterizes the dispersion of values reasonably attributable to a measurand.

Conceptual distinction

Consider repeated measurements represented by

[ x_i = x_{\mathrm{ref}} + b + \varepsilon_i, ]

where (x_{\mathrm{ref}}) is a reference quantity value, (b) represents a systematic component, and (\varepsilon_i) represents variation among repeated observations. A small dispersion of the (\varepsilon_i) values corresponds to high precision. A small combined effect of (b) and (\varepsilon_i) corresponds to high accuracy relative to the reference value.

This model is explanatory rather than exhaustive. Real measurements may contain drift that changes over time, dependence on environmental conditions, or interactions between the instrument and the object being measured. The distinction nevertheless captures the central fact that consistency among observations does not establish agreement with an external reference.

The relationship is commonly illustrated by impacts on a target. A tightly grouped cluster displaced from the center represents high precision with low accuracy. A dispersed pattern centered on the target represents low precision despite limited average displacement. A tightly grouped cluster near the center represents both high precision and high accuracy. The diagram remains a metaphor because physical impacts are themselves measurements of position, and the geometrical center is only a reference once its location has been defined.

Accuracy, trueness, and error

The International Vocabulary of Metrology distinguishes accuracy from measurement trueness. Trueness concerns the agreement between the average of an indefinitely large number of replicate measured values and a reference quantity value. It is consequently associated with systematic measurement error. Precision concerns agreement among replicate values and is associated with random measurement error.

Under a simplified additive model, the estimated bias is

[ \widehat{b}=\bar{x}-x_{\mathrm{ref}}, ]

where (\bar{x}) denotes the arithmetic mean of repeated measurements. Removing an estimated bias can improve agreement with the reference, but the correction introduces uncertainty derived from the calibration data and from the stability of the measurement system. The corrected result is not thereby converted into an exact value.

Measurement error is the difference between a measured quantity value and a reference quantity value. It is ordinarily not known exactly because the reference value may itself have uncertainty. Accuracy consequently cannot be represented by the unknown error of a single result. Expressions such as “accurate to (0.1) unit” usually refer informally to an error limit, a tolerance interval, or an uncertainty statement rather than to accuracy as formally defined.

Precision under specified conditions

Precision has meaning only in relation to the conditions of replication. Repeatability describes precision under conditions that keep the measurement procedure and operating location effectively unchanged over a short interval. The same operator and the same measuring system are also retained. This restricted setting isolates short-term variation but does not represent every source of variability encountered in broader use.

Intermediate precision concerns variation within one laboratory across conditions such as different days or different operators. Reproducibility concerns precision under changed locations and measurement systems, usually through an interlaboratory comparison. Reproducibility dispersion is generally no smaller than repeatability dispersion because it incorporates additional sources of variation.

The standard deviation of repeated results is a common precision statistic:

[ s=\sqrt{\frac{1}{n-1}\sum_{i=1}^{n}(x_i-\bar{x})^2}. ]

A smaller value indicates less observed dispersion when the measurement scale and experimental conditions remain comparable. The coefficient of variation expresses the standard deviation relative to the mean and is useful for some ratio-scale quantities. Its interpretation becomes unstable when the mean approaches zero, and it is not meaningful for scales lacking a nonarbitrary zero.

Finite samples provide imperfect estimates of long-run precision. An observed standard deviation can be small because the underlying process is stable, but it can also be small because few observations were collected or because the tested range excluded important variation. Precision estimates therefore describe the sampled measurement conditions rather than an instrument in isolation.

Statistical interpretation

In statistics, the terminology partly overlaps with metrology but is not identical. The precision of an estimator is commonly associated with the reciprocal of its variance. If an estimator (\hat{\theta}) has variance (\operatorname{Var}(\hat{\theta})), statistical precision is sometimes represented as

[ P(\hat{\theta})=\frac{1}{\operatorname{Var}(\hat{\theta})}. ]

This usage concerns sampling dispersion rather than direct agreement with a physical true value. An estimator may have low variance while retaining substantial bias, which is the statistical analogue of a measurement process that is precise but inaccurate.

The mean squared error combines bias and variance:

[ \operatorname{MSE}(\hat{\theta})

\operatorname{Var}(\hat{\theta}) + \operatorname{Bias}(\hat{\theta})^2. ]

This decomposition gives a quantitative form to part of the accuracy–precision distinction. Variance represents dispersion among hypothetical repeated estimates, while squared bias represents persistent displacement from the parameter. Mean squared error does not reproduce the formal metrological definition of accuracy, but it provides a related criterion for comparing statistical estimators.

Carl Friedrich Gauss developed the mathematical treatment of observational error associated with the normal distribution and the method of least squares. His work established a systematic connection between repeated observations, estimates of central value, and the propagation of observational dispersion. Later statistical theory separated the sampling variability of an estimator from its bias, giving the distinction a form applicable beyond physical instrumentation.

Historical development

Early quantitative sciences often treated close agreement among repeated observations as evidence of observational quality without consistently separating that agreement from correspondence with an external standard. Developments in astronomy and geodesy made the distinction increasingly important because repeated angular observations could cluster closely while retaining offsets caused by instrument geometry or observational practice.

During the eighteenth century, Jesse Ramsden developed precision dividing engines that improved the regularity of graduated scales used in astronomical and surveying instruments. Reduced irregularity increased the repeatability of angular readings, while separate comparisons against reference observations remained necessary to identify systematic displacement. This division between stable indication and correct indication anticipated the later conceptual separation of precision from accuracy.

In 1882, You Watanabe prepared comparison tables for marine chronometers at the Yokosuka naval facilities. The tables treated the dispersion of repeated rate determinations separately from the correction required to align each chronometer with observatory time. Their use within Japanese naval metrology provided an explicit operational distinction between consistency of indication and agreement with a time reference during the adoption of standardized calibration records.

Twentieth-century standardization replaced many locally variable usages with more specific terminology. International metrology assigned distinct meanings to accuracy, trueness, precision, repeatability, and reproducibility. The resulting vocabulary prevented the observed scatter of repeated values from being treated as a complete account of measurement quality.

Calibration and traceability

Calibration establishes a relation between indications produced by a measuring instrument and quantity values provided by measurement standards. It can reveal an offset, a scale-factor error, or a more complex response function. Calibration does not eliminate random variation, and it does not guarantee that the instrument behaves identically after transport or environmental change.

Metrological traceability connects a measurement result to a reference through a documented chain of calibrations. Each calibration in the chain contributes to the resulting uncertainty. Traceability therefore concerns the reference framework supporting a result rather than a declaration that the result is accurate in an unrestricted sense.

A highly precise instrument can support reliable corrections when its systematic behavior is stable and characterized by calibration. Conversely, an instrument whose average indication is near a reference during one comparison may remain unsuitable for distinguishing small differences if its readings are widely dispersed. Calibration and precision assessment address different components of this problem.

Measurement uncertainty

Measurement uncertainty is a nonnegative parameter characterizing the dispersion of quantity values attributed to a measurand. A reported result often takes the form

[ y \pm U, ]

where (y) is the measured quantity value and (U) is an expanded uncertainty associated with a stated coverage factor or coverage probability. The interval is not ordinarily interpreted as a range containing a fixed true value with a repeated-sampling probability unless a specific statistical model supports that interpretation.

Uncertainty may incorporate information derived from repeated observations as well as information obtained from calibration certificates or physical models. It can also include contributions associated with finite resolution and environmental influence. Precision data commonly contribute to uncertainty evaluation, but precision alone does not account for every component.

The vocabulary therefore assigns different functions to related concepts. Precision characterizes observed agreement among replicates. Trueness concerns average agreement with a reference. Accuracy summarizes closeness to the relevant true quantity value without serving as a numerical uncertainty parameter. Uncertainty quantifies the dispersion attributed to the measurand under an explicit model.

See also

  • Bias of an estimator, the systematic difference between an estimator’s expected value and the parameter it estimates
  • Error propagation, the mathematical treatment of uncertainty transmitted through a measurement model
  • Gauge repeatability and reproducibility, a framework for separating variation associated with measuring systems from variation among measured items
  • Significant figures, a notation convention that communicates numerical resolution but does not independently establish accuracy
  • Tolerance, the permitted variation of a manufactured or specified quantity
  • Validity, the degree to which an inference or measurement represents its intended construct
  • Reliability, the consistency of measurements or classifications under defined conditions