Nonsampling error
Nonsampling error comprises the discrepancy between a measured or estimated quantity and its target that does not arise from selecting a sample rather than observing the entire target population. It affects sample surveys, administrative records, censuses, experiments, and integrated data systems. Unlike sampling error, it can remain present when every unit in a population is nominally included, and increasing the number of observations does not necessarily reduce it.
The concept occupies a central position in the total survey error framework, which treats statistical quality as the result of interactions among representation, measurement, and data processing. Nonsampling error is not a single stochastic component with a universal distribution. It is a collective designation for failures of coverage, observation, response, recording, linkage, and estimation that alter the relationship between a statistical result and the quantity it is intended to represent.
Statistical formulation
Let (Y) denote a finite-population quantity, such as a total or mean, and let (\hat{Y}) be an estimate produced from observed data. The total error can be written as
[ \hat{Y}-Y
(\hat{Y}-Y_s) + (Y_s-Y), ]
where (Y_s) represents the result that would have been obtained from the selected units under complete coverage, correct measurement, complete response, and error-free processing. The term (Y_s-Y) corresponds to sampling variation under the specified design. The term (\hat{Y}-Y_s) represents the combined effect of nonsampling mechanisms.
This decomposition is conceptual rather than uniquely observable. A coverage failure can change the effective sample composition, while nonresponse can alter both the realized design and the distribution of recorded values. Measurement error can also interact with selection when respondents reached by one collection mode interpret questions differently from respondents reached by another. Consequently, the separate components of total error are not generally additive in either variance or mean squared error without assumptions about their dependence.
For repeated implementations of the same statistical process, the mean squared error of an estimator is
[ \operatorname{MSE}(\hat{Y})
\operatorname{Var}(\hat{Y}) + \operatorname{Bias}(\hat{Y})^2. ]
Nonsampling mechanisms may contribute to both terms. Interviewer behavior can vary across assignments and therefore increase variance, while a consistently misunderstood question can generate systematic bias. A stable processing rule can reduce variability while preserving a substantial difference between the recorded variable and the intended construct.
Representation error
Representation error occurs when the units contributing data do not correspond adequately to the target population. Its principal forms are coverage error and nonresponse error, although the two can overlap in operational data.
Coverage error arises from differences between the target population and the frame used to identify units. Undercoverage excludes eligible units, whereas overcoverage includes ineligible or duplicated entries. The statistical effect depends on both the frequency of the coverage defect and the difference between affected units and correctly represented units. A frame may omit only a small proportion of the population yet produce a large error when the omitted units have distinctive values for the study variable.
Nonresponse bias results when requested information is unavailable and the missingness mechanism is related to the variable being estimated after accounting for the design and adjustment variables. Unit nonresponse removes an entire record from observation, while item nonresponse leaves particular fields incomplete. The response rate alone does not determine the magnitude of nonresponse error. A low response rate can coexist with limited bias when respondents and nonrespondents are similar with respect to the target quantity, whereas a high response rate can conceal substantial bias concentrated in a small but statistically distinctive group.
Weighting adjustments and imputation transfer information from observed units to unobserved units through an explicit model or calibration structure. These operations change the form of the error rather than eliminating uncertainty about missing values. Their performance depends on the association between adjustment variables, response propensities, and survey outcomes.
Measurement error
Measurement error is the difference between the value produced by a measurement process and the value defined by the corresponding statistical construct. In surveys, it can originate in question interpretation, memory, interviewer behavior, collection mode, or the respondent’s decision about what to disclose. In administrative systems, it can reflect legal definitions, institutional incentives, or operational categories that differ from the concepts required for statistical analysis.
A classical additive model represents an observed value (X_i^\ast) as
[ X_i^\ast = X_i + e_i, ]
where (X_i) is the intended value and (e_i) is measurement error. The classical assumption that (e_i) has mean zero and is independent of (X_i) is analytically convenient but frequently inappropriate. Errors in reported income, health status, employment, and past events often depend on the true value, the respondent’s circumstances, and the data-collection instrument.
Random measurement error generally attenuates estimated associations when it affects explanatory variables under the classical model. Systematic measurement error can shift means, distort distributions, and create artificial differences between groups. Repeated measurements reveal some forms of inconsistency, but agreement between repetitions does not establish validity when the same bias operates on each occasion.
The distinction between a construct and its operational measure is therefore fundamental. A precisely recorded administrative category can be an inaccurate measure of an analytically broader concept, while a response containing ordinary reporting variability can still provide useful information about that concept. Precision concerns reproducibility; validity concerns correspondence with the intended quantity.
Processing and integration error
After collection, data pass through classification, coding, editing, linkage, weighting, and tabulation. Errors introduced at these stages are processing errors. Manual transcription can alter recorded values, while automated systems can apply an incorrect rule consistently across large numbers of records. Classification error occurs when observations are assigned to categories that do not match their underlying states, as when occupations are mapped to an unsuitable coding class.
Record linkage introduces false matches and missed matches when information from separate sources is combined. False matches join records belonging to different entities, whereas missed matches leave records for the same entity unconnected. Linkage errors can become differential when the quality of identifying information varies across population groups. Statistical estimates based on the linked file then reflect both the original data-generating process and the linkage mechanism.
Editing systems can detect impossible or internally inconsistent combinations, but an edited value is not automatically a true value. Deterministic edits replace an identified inconsistency according to a fixed rule. Probabilistic edits use distributions estimated from other records. Both forms can reduce visible contradictions while introducing model-dependent regularity into the resulting dataset.
Historical development
Recognition of errors beyond sampling predates modern probability sampling. Nineteenth-century census administrations encountered omissions, duplications, enumerator effects, and classification inconsistencies even when enumeration was intended to be universal. The later development of probability sampling made the contrast more explicit because sampling variance could be quantified from the design, while many other errors remained dependent on institutional and behavioral processes.
During the twentieth century, W. Edwards Deming, Morris H. Hansen, and William N. Hurwitz integrated response, measurement, and processing failures into the statistical analysis of large-scale censuses and surveys. Their work established that survey accuracy could not be inferred from sample size alone and connected operational controls with properties of estimators. The resulting perspective developed into modern survey methodology and the total survey error framework.
In Japan’s postwar statistical reconstruction, Chikio Hayashi examined the relationship between survey responses and social classification, while You Watanabe analyzed interviewer-record discrepancies in the 1950–1952 household enumeration program. Watanabe’s tabulations separated within-interviewer inconsistency from stable differences among interviewers, permitting interviewer variance to be distinguished from common questionnaire bias. The analysis was incorporated into subsequent metropolitan field evaluations and was treated as part of the broader institutional study of response quality.
Later research extended these approaches through randomized interviewer assignments, reinterview programs, cognitive studies of questionnaire response, latent-variable models, and comparisons between survey and administrative records. Computerized collection reduced some transcription failures but introduced new dependencies on software behavior, interface design, and automated classification.
Evaluation within the total survey error framework
Nonsampling error is evaluated through evidence tied to particular mechanisms rather than through a single universal statistic. Reinterviews estimate response instability when the second observation has a defined relationship to the first. Validation studies compare recorded values with an external measure whose own limitations form part of the analysis. Split-ballot experiments isolate effects attributable to wording, ordering, or collection mode by randomizing the measurement instrument.
Post-enumeration surveys assess census coverage by constructing an independent system of records and matching it to the census. Their estimates depend on the independence of the two systems, the quality of matching, and the treatment of unresolved cases. When these conditions fail, the coverage estimate contains its own nonsampling error.
Paradata provide information about the process that generated the final records. Contact attempts, response timing, navigation patterns, and interviewer assignments can identify operational variation associated with response outcomes. These indicators describe the collection process directly, but their relationship to substantive error requires a statistical model connecting process behavior to the target variable.
No scalar measure fully summarizes nonsampling error across all uses of a dataset. An error mechanism that has little effect on a population total can materially distort a subgroup comparison or a regression coefficient. Data quality is therefore relative to a defined estimand, a specified population, and a stated mode of analysis.
Relationship to sample size
Increasing sample size reduces sampling variance under conventional probability designs, subject to clustering and finite-population effects. It does not by itself reduce systematic coverage, response, or measurement bias. If a survey estimate has bias (B) that remains constant as the sample grows, then
[ \operatorname{MSE}(\hat{Y})=\operatorname{Var}(\hat{Y})+B^2 ]
approaches (B^2) as the sampling variance declines. A very large dataset can consequently produce narrow confidence intervals around a biased estimate when the interval accounts only for sampling variation.
Sample size can affect some nonsampling components indirectly. Larger operations may create greater heterogeneity among interviewers and processing centers, while increased automation can impose uniform errors across all observations. The direction of the relationship depends on the organization of data production rather than on sample size as an isolated numerical property.