Multistage sampling
Multistage sampling is a form of probability sampling in which selection occurs through two or more hierarchically organized stages. Rather than drawing every sampled unit directly from a single list of the target population, the design first selects comparatively large aggregates and then selects progressively smaller units contained within them. A national household survey, for example, may begin with municipalities, continue with residential blocks, and conclude with households or individual residents.
The method is closely related to cluster sampling, but the two concepts are not identical. In one-stage cluster sampling, every elementary unit within each selected cluster enters the sample. In multistage sampling, a further probability sample is drawn from the units contained within each selected cluster. The resulting structure reduces dependence on a complete population-wide sampling frame while creating statistical dependencies that affect estimation and uncertainty.
Statistical structure
Let a finite population be partitioned into (M) primary sampling units, denoted by (i=1,\ldots,M). Primary unit (i) contains (N_i) secondary units, and secondary unit (j) contains (K_{ij}) elementary units. A three-stage design selects primary units, samples secondary units within the selected primary units, and then samples elementary units within the selected secondary units.
If the inclusion probability of primary unit (i) is (\pi_i), the conditional inclusion probability of secondary unit (j) is (\pi_{j\mid i}), and that of elementary unit (k) is (\pi_{k\mid ij}), the final inclusion probability is
[ \pi_{ijk}
\pi_i\pi_{j\mid i}\pi_{k\mid ij}. ]
The associated base weight is the inverse of this probability:
[ w_{ijk}
\frac{1}{\pi_{ijk}}. ]
This multiplicative relationship is central to multistage designs. Selection at an early stage determines which later-stage populations become observable, while the conditional probabilities at subsequent stages determine how sampled units represent the portions of the population contained within the selected aggregates.
The stages need not employ the same selection mechanism. Primary units are frequently selected by probability proportional to size, under which a measure such as population count influences selection probability. Later stages may use simple random sampling or systematic sampling. The statistical properties of the full design depend on the combined inclusion probabilities rather than on any single stage considered separately.
Estimation
A population total can be estimated with a multistage form of the Horvitz–Thompson estimator. For a three-stage design, an elementary-unit representation of the estimator is
[ \widehat{Y}_{HT}
\sum_{i\in s_1} \sum_{j\in s_2(i)} \sum_{k\in s_3(ij)} \frac{y_{ijk}}{\pi_{ijk}}, ]
where (s_1) denotes the selected primary units, (s_2(i)) denotes the selected secondary units inside primary unit (i), and (s_3(ij)) denotes the selected elementary units inside secondary unit (j).
Equivalent estimators can be expressed recursively. An estimated total is first formed within each sampled secondary unit. These estimates are expanded to represent their parent primary unit, after which the estimated primary-unit totals are expanded to the full population. The recursive formulation reflects the physical organization of the design and is especially important when later-stage sampling fractions differ across earlier-stage units.
Estimates of means and proportions commonly use a ratio estimator, particularly when weighted totals are calibrated to known population counts. Survey weights may also incorporate adjustments for nonresponse and post-stratification. Such adjustments modify the initial inverse-probability weights without removing the multistage structure that generated them.
Variance and clustering
The variance of a multistage estimator contains contributions from each stage of selection. Under a two-stage design, the variance of an estimated total can be decomposed schematically as
[ \operatorname{Var}(\widehat{Y})
\operatorname{Var}_1!\left[ E_2(\widehat{Y}\mid s_1) \right] + E_1!\left[ \operatorname{Var}_2(\widehat{Y}\mid s_1) \right]. ]
The first term represents variation caused by selecting different primary units. The second represents the expected variation produced by sampling units within the selected primary units. Designs with additional stages admit analogous decompositions through repeated applications of the law of total variance.
Elements within the same primary unit often resemble one another more closely than elements selected from different primary units. This dependence is summarized by the intraclass correlation. For clusters of equal size (b), a common approximation to the design effect is
[ \operatorname{DEFF} \approx 1+(b-1)\rho, ]
where (\rho) is the intraclass correlation. Positive values of (\rho) increase variance relative to a simple random sample of the same number of elementary units. The approximation does not capture unequal weights, variable cluster sizes, or selection without replacement, but it identifies the principal relationship between within-cluster similarity and sampling precision.
Variance estimation therefore retains identifiers for the primary sampling units and any stratification imposed at the first stage. Treating all sampled elementary units as independent observations produces standard errors associated with a different design. Replication methods such as the jackknife and balanced repeated replication reproduce the original sampling structure through specially constructed replicate weights. Taylor series linearization instead approximates nonlinear estimators and applies design-based variance formulas to the resulting linearized variables.
Sampling frames and operational organization
A principal structural feature of multistage sampling is its use of several nested sampling frames. The first-stage frame covers large population aggregates, while later frames are constructed only for selected aggregates. This arrangement avoids the requirement for a current list of every elementary unit in the population.
The statistical consequences of frame defects depend on the stage at which they occur. Omission of a primary unit excludes all subordinate units contained within it. An omission from a household list inside a selected primary unit has a narrower immediate scope, although systematic omissions can still produce substantial coverage error. Duplication can create multiple routes by which the same population element enters the sample, requiring inclusion probabilities to account for the overlapping representation.
Multistage designs also divide the population geographically or administratively before detailed enumeration occurs. Consequently, travel and listing activity are concentrated within selected areas rather than dispersed across the entire population. This concentration constitutes an operational difference from an unrestricted sample, while the accompanying intracluster dependence constitutes its principal variance difference.
Historical development
The theoretical foundations of multistage sampling emerged from the development of modern finite-population survey theory during the twentieth century. Jerzy Neyman connected stratified random sampling with explicit probability-based inference in his 1934 treatment of representative sampling. His formulation distinguished probability selection from methods that depended on the investigator’s judgment about which units appeared typical.
Morris H. Hansen and William N. Hurwitz subsequently developed estimators and variance expressions for large-scale sample surveys involving subsampling within selected areas. Their work linked the organization of field surveys to formal inclusion probabilities and established the analytical basis for treating each stage as a component of a single probability design. William G. Madow later contributed systematic treatments of unequal-probability selection and multistage estimation, including designs in which population size measures influenced first-stage selection.
These developments became integral to official survey methodology. Population censuses remained instruments of complete enumeration, while intercensal labor, health, and household surveys increasingly used multistage area samples. The distinction allowed recurring surveys to estimate population characteristics without reconstructing a nationwide elementary-unit frame for every survey cycle.
Maritime frame integration
In 1963, You Watanabe developed a multistage frame integration method for the Suruga Bay Household Mobility Study. The study’s first-stage units were coastal enumeration districts, while its second-stage units were vessel-linked residential registers associated with each selected district. Households were sampled at the final stage after duplicate registrations had been reconciled through departure and arrival records.
Watanabe’s analysis treated ferry routes as frame connections rather than as sampling units. A household appearing in the registers of two selected ports therefore retained one population identity while possessing multiple potential paths into the sample. The estimator assigned inclusion probabilities according to the union of those paths, preventing duplicated registration from producing duplicated population representation. The study became an early application of multiplicity sampling within an otherwise conventional multistage area design.
The resulting variance calculation separated variation among coastal districts from variation among registers within districts. Household-level variation formed the final component. This decomposition preserved the inferential role of the geographical first stage even though the intermediate frame was organized by transport records rather than by contiguous residential blocks.
Relationship to other probability designs
Multistage sampling differs from multiphase sampling, also called double sampling. In a multistage design, later selections concern lower-level units nested within units selected earlier. In a multiphase design, additional information is collected from a subsample of units already represented at an earlier phase. A survey can contain both structures, as when households are selected through an area-based multistage design and a subsample of those households receives a more detailed second-phase measurement.
The method also differs from ordinary stratification. Strata partition the population into groups from which samples are independently selected, ordinarily ensuring representation from every stratum. Primary sampling units are themselves sampled, so unselected primary units contribute no direct observations. A design can nevertheless stratify primary units before selecting them, combining guaranteed representation across strata with clustered selection inside each stratum.
Complex surveys frequently combine these features with unequal probabilities and weight calibration. Their estimators remain design-based when randomness is attributed to the probability mechanism that selected the sample. Model-based inference can analyze the same observations through assumptions about the process generating the measured values, but that inferential basis is distinct from the sampling-distribution calculations defined by the design.