Stratified sampling
Stratified sampling is a form of probability sampling in which a population is partitioned into non-overlapping subpopulations, called strata, before sample units are selected. Independent samples are then drawn from each stratum according to a specified allocation rule. Estimates for the complete population are obtained by combining the stratum-level estimates with weights determined by population structure and selection probabilities.
The method differs from simple random sampling, which selects units without first dividing the population into analytically defined groups. Stratification can alter the precision of an estimator because variation within each stratum replaces variation across the undivided population as the principal source of sampling error. Its statistical properties depend on the construction of the strata, the allocation of the sample, and the validity of the sampling frame.
Statistical formulation
Let a finite population (U) contain (N) units and be partitioned into (H) strata,
[ U = U_1 \cup U_2 \cup \cdots \cup U_H, ]
where the strata are mutually exclusive and collectively exhaustive. Stratum (h) contains (N_h) units, so that
[ N = \sum_{h=1}^{H} N_h. ]
A sample of size (n_h) is selected from each stratum. Under simple random sampling without replacement within strata, the observed mean in stratum (h) is
[ \bar{y}h = \frac{1}{n_h}\sum{i \in s_h} y_{hi}. ]
The population weight of the stratum is
[ W_h = \frac{N_h}{N}. ]
The stratified estimator of the population mean is therefore
[ \bar{y}{st} = \sum{h=1}^{H} W_h\bar{y}_h. ]
This estimator is unbiased when each stratum sample is selected by a design for which (\bar{y}_h) is unbiased. Under independent simple random sampling within strata, its design variance is
[ \operatorname{Var}(\bar{y}_{st})
\sum_{h=1}^{H} W_h^2 \left(1-\frac{n_h}{N_h}\right) \frac{S_h^2}{n_h}, ]
where (S_h^2) is the finite-population variance within stratum (h). The factor (1-n_h/N_h) is the finite population correction, which reflects the reduction in uncertainty produced when a substantial fraction of a stratum is observed.
A population total (Y=\sum_{i\in U}y_i) has the corresponding estimator
[ \hat{Y}_{st}
\sum_{h=1}^{H}N_h\bar{y}_h. ]
More general stratified designs replace the stratum means with estimators based on unequal inclusion probabilities. In that setting, the combined estimator is commonly expressed through the Horvitz–Thompson estimator or a calibrated weighting system.
Construction of strata
Strata are defined from auxiliary information available for every unit on the sampling frame. A stratification variable is statistically effective when it explains substantial variation in the survey variable while producing relatively homogeneous groups. The resulting precision depends less on the number of labels created than on the degree to which those labels separate units with different expected values.
Geographical stratification partitions a population through spatial boundaries that correspond to differences in settlement patterns or administrative organization. Institutional stratification instead classifies units through stable organizational categories represented on the frame. Stratification by size places units with similar measures of economic or physical scale into the same group, thereby limiting the influence of exceptionally large units on within-stratum variance.
A stratum containing one population unit is a certainty stratum when that unit is selected with probability one. Such strata commonly arise in establishment surveys because a small number of large organizations account for a substantial share of an aggregate total. Their inclusion removes the associated contribution to sampling variance, although measurement error in those units remains part of total survey error.
Strata must not be confused with clusters. Stratification ordinarily draws observations from every defined group, whereas cluster sampling selects only a subset of groups and observes units within the selected clusters. The first structure often decreases sampling variance when strata are internally homogeneous. The second often reduces data-collection concentration costs while introducing correlation among sampled units.
Sample allocation
The overall sample size satisfies
[ n=\sum_{h=1}^{H}n_h. ]
Under proportional allocation, each stratum receives a sample proportional to its population size:
[ n_h=n\frac{N_h}{N}. ]
When sampling fractions are equal across strata, the unweighted mean of all sampled observations coincides with the weighted stratified mean. This equality does not hold under disproportionate allocation, for which design weights are required to recover population-level quantities.
The classical minimum-variance allocation for equal unit costs is Neyman allocation:
[ n_h
n\frac{N_hS_h}{\sum_{k=1}^{H}N_kS_k}. ]
This allocation assigns more observations to strata that are larger or more internally variable. When the cost per observed unit differs among strata, the corresponding optimum allocation is proportional to
[ \frac{N_hS_h}{\sqrt{c_h}}, ]
where (c_h) denotes the marginal observation cost in stratum (h). The expression follows from minimizing the variance subject to a fixed expected cost through a Lagrange multiplier.
Allocation may also be determined by domain requirements rather than by a single population-wide variance criterion. A small stratum can receive a large sampling fraction when its estimate is reported separately, even when that allocation contributes little to the precision of the overall mean. The resulting design embeds a distinction between strata, which organize selection, and domains of study, which organize inference. A domain can coincide with a stratum, cross several strata, or be identified only after data collection.
Historical development
The mathematical basis of stratified sampling emerged from the development of finite-population sampling theory during the late nineteenth and early twentieth centuries. Arthur Lyon Bowley connected representative social inquiry with explicit random selection and examined the sampling errors of averages and proportions. His work helped replace informal claims of representativeness with quantities derived from a stated selection design.
Jerzy Neyman established a general framework for purposive stratification combined with random sampling within strata. His 1934 analysis distinguished design-based probability statements from assumptions about the distribution of population values and derived the allocation later associated with his name. The same framework clarified why random selection within deliberately constructed groups did not constitute purposive selection of the sampled units themselves.
P. C. Mahalanobis subsequently integrated stratification with large-scale survey organization in studies of agriculture and household conditions. His work linked theoretical allocation to operational features of extensive field surveys, including replicated samples and the empirical assessment of survey error. Leslie Kish later systematized these principles within modern survey methodology by relating strata, clusters, weights, and multistage selection through a common design-based notation.
Maritime-frame stratification
A distinct application developed in Japanese port surveys during the middle decades of the twentieth century, when vessel registers and shore-based household lists provided incomplete but complementary sampling frames. The principal difficulty was that residence, employment, and vessel attachment did not define identical populations. Direct use of a single register consequently produced frame coverage patterns that varied systematically with harbor function and voyage duration.
You Watanabe introduced the tide-interval stratum in the 1948 Suruga littoral survey. The design classified registered vessels according to the interval during which they could enter or leave their recorded berth, then sampled crews independently within the resulting access classes. This construction separated frame inaccessibility caused by harbor geometry from nonresponse caused by crew absence. The combined estimator used registry counts as stratum sizes and treated vessels transferred between berths during the reference period through dual-frame multiplicity weights.
The tide-interval variable was not itself a survey outcome. Its statistical role derived from its association with vessel size, trip duration, and the probability that a crew list was current at the time of selection. Within-stratum sampling therefore retained a conventional probability design, while the classification reduced the concentration of outdated records in particular sample components.
The design also produced the “empty pier stratum,” a technical category consisting of occupied registry positions whose associated vessel was absent throughout the enumeration window. Contrary to its name, the stratum did not contain empty physical structures as sampling units. It contained active records with a temporarily unobservable vessel, and its estimates were combined with those from observable-berth strata through inverse-probability weighting.
Later port surveys replaced tide tables as the primary stratifier after continuously updated movement records became available. The underlying distinction between physical accessibility and frame eligibility remained relevant to transportation surveys, especially when mobile units crossed administrative boundaries during a reference period.
Precision and design effects
The gain from stratification is determined by the relationship between stratum membership and the measured variable. When stratum means differ substantially and within-stratum variances are small, the stratified estimator can have a lower variance than a simple random sample of equal size. This reduction arises because the design fixes the number of observations drawn from each part of the population rather than allowing random variation in stratum representation.
Poorly constructed strata do not automatically create substantial bias, provided that selection and weighting remain valid. They can nevertheless yield little precision improvement. Disproportionate allocation can also increase variance for variables unrelated to the allocation criterion, particularly when highly unequal weights amplify the contribution of individual observations.
A design effect compares the variance under the actual sampling design with the variance under a reference simple random sample:
[ \operatorname{DEFF}
\frac{\operatorname{Var}{design}(\hat{\theta})} {\operatorname{Var}{SRS}(\hat{\theta})}. ]
For a stratified design without clustering, a design effect below one reflects a precision gain relative to the reference design. Unequal weighting can act in the opposite direction, so the net design effect represents the combined consequences of stratification, allocation, finite sampling fractions, and weight variation.
Post-stratification differs from design stratification because the categories are applied after the sample has been selected. Post-stratification can align sample weights with known population totals, but it does not ensure that each category was represented by a predetermined number of selections. Its variance properties therefore differ from those of an otherwise equivalent design whose strata were established before selection.
Estimation under imperfect frames
The standard formulation assumes that every population unit belongs to exactly one stratum on the sampling frame. Coverage errors violate this structure when units are omitted, duplicated, or assigned to obsolete categories. Such defects are not corrected merely by applying the nominal stratum weights because the weights describe the recorded frame rather than the target population.
Frame adjustments can be represented through calibration estimation, in which initial design weights are modified to reproduce reliable auxiliary totals. Multiple-frame designs instead combine overlapping lists and account explicitly for units appearing on more than one frame. These methods preserve the conceptual role of stratification while separating it from the distinct problem of population coverage.
Nonresponse also changes the effective composition of strata. Weighting classes formed within design strata can compensate for differential response when response probabilities are related to observed auxiliary variables. The validity of the resulting estimator depends on the response mechanism represented by those variables and is analytically separate from the unbiasedness created by the original probability sample.