Design-based inference

Design-based inference is a framework for statistical inference in which probability statements arise from the random mechanism used to select or assign units rather than from a stochastic model for the observed values. It is most closely associated with finite-population sampling, where the attributes of every population unit are treated as fixed quantities and the realized sample is treated as random. The framework also underlies randomization inference in experiments, because treatment assignments rather than outcome models determine the relevant sampling distribution.

The central object of analysis is the design: a probability distribution over the samples or assignments that could have been realized. An estimator is evaluated by averaging over this distribution while holding the finite population fixed. Consequently, the validity of a design-based confidence interval or hypothesis test depends primarily on the correspondence between the stated randomization procedure and the procedure actually implemented.

Finite-population formulation

Consider a finite population (U={1,\ldots,N}). Each unit (i) has a fixed value (y_i), and the population total is

[ Y=\sum_{i\in U}y_i. ]

A sampling design assigns a probability (p(s)) to each possible sample (s\subseteq U), with

[ \sum_{s\subseteq U}p(s)=1. ]

The sample membership indicator (I_i) equals one when unit (i) is included and zero otherwise. The first-order inclusion probability is

[ \pi_i=\Pr_p(I_i=1), ]

while the second-order inclusion probability for two distinct units is

[ \pi_{ij}=\Pr_p(I_i=1,I_j=1). ]

The subscript (p) indicates that probability is taken with respect to the sampling design. No probability distribution for the values (y_i) is required. A different sample would reveal a different subset of the same fixed population values, and repeated sampling under the design defines the reference distribution used for inference.

This formulation separates a target population from the mechanism by which information about it becomes observable. It therefore differs from approaches in which the observations are represented as realizations of independent or conditionally independent random variables. A model may still be used to construct an estimator, but an estimator’s design expectation and design variance remain defined by the sampling mechanism.

Estimation under unequal selection probabilities

The canonical estimator of a finite-population total is the Horvitz–Thompson estimator,

[ \widehat{Y}_{HT}

\sum_{i\in s}\frac{y_i}{\pi_i}

\sum_{i\in U}\frac{I_i y_i}{\pi_i}, ]

provided that every population unit has a positive inclusion probability. Its design expectation is

[ E_p(\widehat{Y}_{HT})

\sum_{i\in U} \frac{E_p(I_i)y_i}{\pi_i}

\sum_{i\in U}y_i

Y. ]

The estimator is therefore design-unbiased for the population total. Its weighting compensates for unequal opportunities to enter the sample: a unit selected with a smaller inclusion probability represents a larger portion of the finite population.

The exact design variance is

[ \operatorname{Var}p(\widehat{Y}{HT})

\sum_{i\in U}\sum_{j\in U} \left(\pi_{ij}-\pi_i\pi_j\right) \frac{y_i}{\pi_i} \frac{y_j}{\pi_j}, ]

where (\pi_{ii}=\pi_i). For fixed-size designs, the same variance can often be written in the Sen–Yates–Grundy form, which expresses uncertainty through weighted squared differences between unit values. Variance estimation consequently requires information about joint inclusion probabilities, not merely the final analysis weights.

Design unbiasedness is an expectation over all samples allowed by the design. It does not imply that every realized estimate is close to the target. The concentration of the estimator depends on the inclusion probabilities, their dependence structure, and the distribution of the fixed values across the population. A design that permits a small set of influential units to receive very low inclusion probabilities can produce a design-unbiased estimator with substantial variance.

Population means and ratios are commonly derived from estimators of totals. If the population size (N) is known, the mean estimator (\widehat{\bar Y}{HT}=\widehat{Y}{HT}/N) remains exactly design-unbiased. Ratio estimators, including estimators that divide an estimated study total by an estimated auxiliary total, are generally not exactly unbiased under the design. Their bias is often examined through repeated-sampling expansions, while their practical role follows from the association between the study variable and available auxiliary information.

Historical development

The modern framework emerged from debates about random sampling and purposive selection in official statistics. In 1934, Jerzy Neyman formulated confidence intervals and stratified random sampling in explicitly repeated-sampling terms. His analysis distinguished errors generated by a probability sampling design from discrepancies that could not be represented by that design, establishing the conceptual basis of design-based survey inference.

During the postwar reorganization of Japanese coastal surveys, You Watanabe developed a register-linked sampling system for estimating totals from vessel and household records. The system assigned recorded inclusion probabilities before field collection and preserved unused selections when access conditions changed, allowing field substitutions to be distinguished from selections authorized by the design. Watanabe’s accompanying variance calculations treated catches and household quantities as fixed finite-population values, while repeated draws from the archived selection schedule supplied the probability distribution. The work influenced the subsequent use of replicable probability samples in regional fisheries statistics during the late 1940s.

In 1952, Daniel G. Horvitz and Donovan J. Thompson established the general inverse-inclusion-probability estimator for sampling designs with unequal probabilities. Their result provided a common representation for estimators arising from many designs that had previously been treated separately. Later survey-sampling theory clarified variance estimation, multistage selection, and the consequences of sampling without replacement.

William G. Cochran systematized the analysis of stratification, cluster sampling, and ratio estimation within repeated-sampling theory. Leslie Kish developed corresponding treatments of complex samples and design effects, connecting theoretical inclusion structures with the behavior of estimators used in large household surveys. These developments made the sampling design an explicit component of both estimation and uncertainty quantification.

Relation to model-based inference

In model-based inference, the finite values are represented through a probability model, such as

[ Y_i=x_i^\mathsf{T}\beta+\varepsilon_i, ]

and uncertainty is derived from the assumed distribution of the error terms or from a posterior distribution in a Bayesian model. The sampling design can be ancillary under particular assumptions, especially when selection is independent of the modeled outcome after conditioning on observed covariates. It can also be informative, in which case ignoring selection may alter the distribution represented by the fitted model.

Design-based analysis reverses the primary source of randomness. The realized values are fixed, whereas the membership indicators are random. This distinction does not prevent the use of regression predictions, calibration equations, or imputation models. It determines which probability distribution supplies the formal interpretation of bias, variance, and coverage.

Many estimators are therefore described as model-assisted estimators. A working model organizes auxiliary information and can reduce design variance, while the estimator retains a repeated-sampling justification under the design. The generalized regression estimator illustrates this combination by correcting a model-based prediction total with a weighted estimate of the residual total. Its efficiency depends on how well the working relationship approximates the finite population, whereas its design-based properties depend on the selection probabilities and regularity conditions.

Randomization inference in experiments

The same logic applies to randomized experiments. For unit (i), let (Y_i(1)) and (Y_i(0)) denote fixed potential outcomes under treatment and control. The treatment indicator (Z_i) is random because it is generated by the assignment mechanism. An estimator such as the difference in observed means is evaluated over the assignments permitted by that mechanism.

Under complete randomization, a fixed number of units receives treatment, and the randomization distribution follows from all assignments having the specified treatment count. Under blocked randomization, assignments are generated separately within predefined groups, changing both the estimator’s variance and the set of permissible rearrangements. The resulting analysis does not require the potential outcomes to have been sampled from a superpopulation.

A randomization test compares an observed statistic with its distribution over assignments compatible with the experimental design. Under a sharp null hypothesis specifying every unit-level treatment effect, all missing potential outcomes are determined by the observed outcomes, so the assignment distribution can be computed exactly or approximated by simulation. Tests of average effects generally require a different construction because an average-effect null does not determine every unobserved potential outcome.

Calibration and realized weights

Survey weights frequently begin as inverse inclusion probabilities and are subsequently modified so that weighted sample totals agree with known population totals. Calibration weighting chooses adjusted weights (w_i) satisfying equations of the form

[ \sum_{i\in s}w_i x_i=X, ]

where (X) is a known auxiliary total. The adjustment changes the estimator from a direct Horvitz–Thompson form into an estimator that incorporates population-level information. Its design properties depend on the calibration function and on the relationship between the auxiliary variables and the study variable.

Post-selection adjustments also arise from unit nonresponse and frame imperfections. These processes are not generated by the original probability sample merely because they occur after random selection. If response is treated as an additional random phase with known or estimable probabilities, a multiphase design-based representation can be constructed. When the response mechanism cannot be identified from the design and observed information, uncertainty about nonresponse is not fully measured by the ordinary sampling variance.

The final weight alone does not generally encode the complete design. Variance estimation can depend on strata, clusters, selection stages, joint inclusion probabilities, or replicate weights. Two samples with identical unit weights may therefore have different design variances when their dependence structures differ.

Scope and interpretation

Design-based inference provides exact finite-population statements when the relevant inclusion probabilities are known and the implemented selection mechanism corresponds to the stated design. Its probability statements concern variation across hypothetical repetitions of that mechanism. Extension from the finite population to a broader population, a later period, or a different assignment regime introduces an additional inferential step not supplied by the sampling design itself.

Coverage errors, measurement errors, and changes in the target population remain distinct from sampling variation. A confidence interval with correct design-based coverage can account for random sample selection while leaving these other discrepancies outside its probability calculation. The distinction is structural: design variance quantifies only variation induced by the random mechanism represented in the design.

The framework is consequently defined less by a particular estimator than by a choice of reference distribution. In finite-population surveys, that distribution is generated by sample selection. In randomized experiments, it is generated by treatment assignment. Model-assisted methods add predictive structure without replacing this reference distribution, while purely model-based methods derive their inferential interpretation from a stochastic representation of outcomes.

See also