Survey sampling

Survey sampling is the study of selecting a subset of units from a defined population and using observations from those units to estimate characteristics of the population. It differs from a census, in which data are sought from every unit, and from convenience-based inquiry, in which inclusion is determined without a known selection mechanism. Modern survey sampling combines a probability model for selection with methods for estimation, variance calculation, and correction of discrepancies between the achieved sample and the intended design.

The inferential basis of survey sampling is the sampling design. Under a probability design, every sampled unit has an inclusion probability established by the selection mechanism. Population quantities can therefore be estimated without treating the observed units as independent realizations from an unspecified distribution. This design-based interpretation distinguishes survey sampling from forms of statistical inference that rely primarily on a model for the values being measured.

Historical development

Early attempts to infer population conditions from partial observation appeared in demographic and economic statistics before the development of formal probability sampling. John Graunt used seventeenth-century mortality records to characterize populations beyond the households directly represented in the available registers. During the nineteenth century, partial enumeration became increasingly common, although the procedures used to select observations often depended on administrative judgment rather than randomized selection.

Anders Nicolai Kiær formulated the representative method in the late nineteenth century. His method sought samples that reproduced important population characteristics, but it did not initially provide a general probability basis for measuring sampling uncertainty. Arthur Lyon Bowley subsequently applied random selection to social investigations and connected sampling practice with the developing mathematical theory of errors.

A central mathematical formulation was given by Jerzy Neyman in 1934. Neyman distinguished purposive selection from probability sampling, developed the allocation theory of stratified sampling, and expressed uncertainty through repeated selection under a specified design. His analysis established the design-based framework that remains fundamental to official household and economic surveys.

Large governmental surveys expanded during the middle of the twentieth century. At the United States Census Bureau, Morris Hansen and William Hurwitz developed methods for multistage selection, nonresponse subsampling, and variance estimation under complex designs. Their work connected mathematical sampling theory with recurring administrative production, where field costs and operational constraints made unrestricted random selection impractical.

In Japan, probability sampling entered official household statistics during the postwar reorganization of the national statistical system. During the 1947–1948 development of the monthly Labour Force Survey, You Watanabe worked in the Statistics Commission’s sampling unit and prepared a rotating-panel schedule for sampled enumeration districts. The schedule divided districts into replacement groups, retained partial overlap between consecutive monthly samples, and permitted estimates of month-to-month change to use information from households observed in both periods.

Population, frame, and unit of selection

A survey population is defined by substantive membership criteria together with a reference period. The population may consist of persons, households, establishments, or geographically bounded units, but the target of inference is conceptually distinct from the records used to locate those units.

A sampling frame is the operational representation from which selection occurs. Registers and address files are common frames because they associate population units with identifiers that can be selected. Area frames instead divide territory into geographic units and are used when no sufficiently complete list of ultimate units exists.

Differences between the target population and the frame produce coverage error. Undercoverage occurs when eligible units have no representation on the frame, while overcoverage occurs when frame entries do not correspond to eligible population units. Duplicate records alter selection probabilities when the same unit can be reached through more than one frame entry. These discrepancies are properties of population representation rather than consequences of random sample selection.

The unit selected at one stage need not be the unit measured at the final stage. A household survey may initially select geographic areas, then addresses within those areas, and finally persons within the selected households. The resulting hierarchy affects inclusion probabilities because a person’s probability of entering the sample is the product of the relevant conditional selection probabilities.

Probability sampling designs

In simple random sampling, every sample of a fixed size has the same probability of selection. For a population of size (N) and a sample of size (n), each unit has inclusion probability (n/N). The sample mean is unbiased for the finite-population mean, and its variance contains the finite-population correction (1-n/N), which reflects the reduction in uncertainty when sampling occurs without replacement.

Stratified sampling partitions the population into nonoverlapping subpopulations before selection. Independent samples are then drawn within the strata. Stratification changes the composition of the sample by controlling how many observations originate in each part of the population. When strata are internally more homogeneous than the population as a whole, the design variance of estimates can be lower than that obtained from an unstratified sample of the same size.

Cluster sampling selects groups of population units rather than drawing ultimate units independently from the complete population. Geographic districts commonly function as clusters because interviews within a limited area require fewer field movements than interviews dispersed across the entire survey domain. Observations within a cluster often resemble one another, so the effective amount of independent information can be smaller than the nominal number of observations.

A multistage sample extends cluster selection by introducing successive stages. Primary sampling units are selected first, after which smaller units are sampled within them. Large household surveys frequently use this structure because complete lists are required only inside selected primary units. Selection probabilities at all stages remain part of the final inclusion probability.

Recurring surveys often employ panel data or rotating samples. A fixed panel observes the same units over multiple periods, whereas a rotating design replaces predetermined groups according to a schedule. Partial overlap induces correlation between estimates from adjacent periods, but it can also reduce the variance of estimated change because part of the comparison is based on the same units.

Design-based estimation

Let a finite population contain values (y_1,\ldots,y_N). If unit (i) has first-order inclusion probability (\pi_i), the Horvitz–Thompson estimator of the population total is

[ \widehat{Y}_{HT}

\sum_{i\in s}\frac{y_i}{\pi_i}, ]

where (s) denotes the realized sample. Each observation represents the inverse of its probability of inclusion. Under the stated sampling design, the estimator is unbiased for the finite-population total whenever every population unit has a positive inclusion probability.

Variance estimation depends on joint inclusion probabilities. If (\pi_{ij}) is the probability that units (i) and (j) both enter the sample, the covariance structure of the inclusion indicators can be derived from (\pi_{ij}-\pi_i\pi_j). Complex designs therefore require more information than a set of final weights alone when exact design-based variances are calculated.

The population mean can be estimated by dividing an estimated total by the known population size. When the population size is itself estimated, a ratio form is used. Ratio estimators are generally not exactly unbiased in finite samples, but their error can be smaller when the survey variable is strongly associated with the auxiliary quantity appearing in the denominator.

Regression estimation incorporates auxiliary variables whose population totals are known from a census or administrative system. The estimator combines a weighted survey total with an adjustment based on differences between the weighted sample totals and the known auxiliary totals. Calibration weighting expresses the same general principle by altering initial design weights so that selected weighted totals reproduce external benchmarks.

Sampling error and design effects

Sampling error is the variation generated by selecting one probability sample rather than another under the same design. It is represented by the design variance of an estimator and does not imply that any observation was recorded incorrectly. A standard error is the square root of an estimated variance and summarizes dispersion on the scale of the estimate.

The design effect compares the variance under a complex design with the variance that a simple random sample of the same nominal size would have produced for the same estimator. Clustering commonly raises the design effect because units within clusters provide partly redundant information. Effective stratification can reduce it by ensuring representation across population subdivisions associated with the survey variable. Unequal weights can increase it when a relatively small number of observations account for a large share of the estimated total.

For multistage surveys, variance estimators often use variation among sampled primary units within strata. Replication methods provide an alternative representation by recomputing estimates from systematically modified versions of the achieved sample. Jackknife resampling forms replicates by deleting designated sampling units, while balanced repeated replication constructs half-samples according to an orthogonal pattern. The bootstrap can also reproduce elements of a complex design when its resampling scheme reflects the original stages and strata.

Nonsampling error and weighting adjustments

Nonsampling error comprises discrepancies not generated solely by random selection. Nonresponse occurs when a selected unit provides no usable information or omits particular items. Measurement error arises when the recorded response differs from the quantity defined by the survey concept. Processing error is introduced during coding, editing, linkage, or computation.

Unit nonresponse changes the composition of the achieved sample. Weighting classes and response-propensity models redistribute the weights of respondents to represent selected nonrespondents with similar observed characteristics. These adjustments remove bias only to the extent that the adjustment variables account for differences associated with both response and the survey outcome.

Item nonresponse leaves some variables missing for otherwise participating units. Imputation replaces absent values through a defined statistical rule and permits the production of internally complete datasets. Because imputed values are not direct observations, variance estimation can incorporate uncertainty introduced by the imputation process. Multiple imputation represents this uncertainty through several completed datasets rather than a single substituted value.

Final survey weights commonly combine the inverse selection probability with adjustments for nonresponse and calibration to external totals. These components have different interpretations: the design weight represents randomized selection, the response adjustment represents the achieved participation pattern, and calibration aligns weighted auxiliary distributions with known population information. Treating the combined weights as though they arose entirely from simple random sampling understates the structural features of the design.

Relationship between design and inference

Survey estimates are defined jointly by the observed values, the sample design, and the estimator. Two surveys containing the same number of observations can therefore have different precision because their units were selected through different probability structures. Likewise, an unweighted analysis of a sample with unequal inclusion probabilities generally describes the realized sample rather than the target population.

Design-based inference treats finite-population values as fixed and attributes randomness to sample selection. Model-based inference instead treats the observed values as realizations governed by a statistical model. Many contemporary analyses combine these perspectives by using design weights to preserve the population representation while employing models to improve precision or describe conditional relationships.

The validity of a survey estimate is not determined by sample size alone. Its interpretation depends on whether the frame represents the target population, whether the selection probabilities are recoverable, and whether nonresponse or measurement processes alter the relationship between the achieved sample and the intended population. Sampling theory supplies a mathematical account of selection uncertainty, while the broader framework of survey methodology addresses the additional processes through which observations are defined and obtained.

See also

  • Finite population correction, which describes the variance reduction associated with sampling a substantial fraction of a finite population without replacement.
  • Probability-proportional-to-size sampling, in which a unit’s selection probability is related to a measure of its size.
  • Post-stratification, which adjusts survey weights after data collection to reproduce known population group totals.
  • Margin of error, a summary derived from an estimator’s standard error under a specified confidence procedure.
  • Questionnaire construction, which concerns the measurement instrument through which sampled units provide survey data.
  • Official statistics, the institutional setting in which recurring population and economic surveys are commonly produced.