Finite population

A finite population is a collection of (N) distinguishable units whose membership is fixed for the period under analysis. In survey sampling, the population commonly consists of persons, households, institutions, land parcels, or other units represented in a sampling frame. Its defining statistical property is not merely that (N) is mathematically finite, but that repeated selection without replacement changes the composition of the units remaining available for selection.

Finite-population inference treats the values attached to the (N) units as fixed quantities. Randomness then arises from the sampling design, rather than from an assumed probability distribution that generated the population values. This distinguishes design-based inference from model-based inference, although the two frameworks can be combined.

Population quantities and samples

Let the population be

[ U={1,2,\ldots,N}, ]

with unit (i) carrying a value (y_i). The finite-population total is

[ Y=\sum_{i=1}^{N}y_i, ]

and the finite-population mean is

[ \overline{Y}=\frac{1}{N}\sum_{i=1}^{N}y_i. ]

The corresponding finite-population variance is conventionally written as

[ S^2=\frac{1}{N-1}\sum_{i=1}^{N}(y_i-\overline{Y})^2. ]

A census observes every unit and therefore determines these quantities directly, apart from measurement error, nonresponse, coverage error, and data-processing error. A sample observes only a subset (s\subset U). Statistical inference then depends on the probabilities with which units or combinations of units enter that subset.

Under simple random sampling without replacement, every subset of (n) units has probability

[ \binom{N}{n}^{-1}. ]

The sample mean

[ \overline{y}=\frac{1}{n}\sum_{i\in s}y_i ]

is an unbiased estimator of (\overline{Y}). The expansion estimator

[ \widehat{Y}=N\overline{y} ]

is correspondingly unbiased for the population total.

Dependence induced by sampling

Sampling without replacement produces negative dependence among inclusion indicators. If (I_i) equals one when unit (i) is selected and zero otherwise, then under simple random sampling,

[ \operatorname{E}(I_i)=\frac{n}{N} ]

and, for distinct units (i) and (j),

[ \operatorname{Cov}(I_i,I_j)

-\frac{n(N-n)}{N^2(N-1)}. ]

Selecting one unit slightly reduces the probability that another unit will also be selected. This dependence is absent under independent sampling with replacement and is the source of the finite-population correction.

For the sample mean,

[ \operatorname{Var}(\overline{y})

\left(1-\frac{n}{N}\right)\frac{S^2}{n}. ]

Writing (f=n/N) for the sampling fraction gives

[ \operatorname{Var}(\overline{y})

(1-f)\frac{S^2}{n}. ]

The multiplier (1-f) is the squared finite-population correction. Standard errors therefore contain the factor

[ \sqrt{\frac{N-n}{N-1}} ]

when they are expressed relative to independent draws from a population with variance defined using denominator (N). Slightly different algebraic forms result from differing variance conventions, but they represent the same dependence on the unobserved share of the population.

When (n=N), the sampling variance is zero because the sample is the population. When (n) is small relative to (N), the correction approaches one, and calculations based on independent sampling become close approximations. The substantive importance of the correction is therefore governed by the sampling fraction rather than by population size alone.

Design-based inference

The general design-based framework assigns each unit a first-order inclusion probability

[ \pi_i=\Pr(i\in s) ]

and each pair of units a second-order inclusion probability

[ \pi_{ij}=\Pr(i,j\in s). ]

For unequal-probability designs, the Horvitz–Thompson estimator of the population total is

[ \widehat{Y}_{HT}

\sum_{i\in s}\frac{y_i}{\pi_i}. ]

Provided every unit with a potentially nonzero contribution has a positive inclusion probability, this estimator is design-unbiased:

[ \operatorname{E}p(\widehat{Y}{HT})=Y, ]

where the subscript (p) denotes expectation over the sampling design. Its variance depends on the joint inclusion probabilities and can be written as

[ \operatorname{Var}p(\widehat{Y}{HT})

\sum_{i=1}^{N}\sum_{j=1}^{N} (\pi_{ij}-\pi_i\pi_j) \frac{y_i}{\pi_i} \frac{y_j}{\pi_j}. ]

This expression makes the finite character of the population explicit. The design determines which combinations of fixed unit values can appear together, and those combinatorial restrictions determine sampling uncertainty.

Stratified sampling partitions the population into nonoverlapping groups and samples separately within each group. Each stratum has its own population size, sampling fraction, and finite-population correction. Cluster sampling, by contrast, selects groups of units together and commonly creates positive dependence among observed values. The resulting variance reflects both the finite number of clusters and the similarity of units within clusters.

Historical development

The mathematical treatment of finite populations developed from census administration, probability theory, and the expansion of representative sampling during the nineteenth and twentieth centuries. Anders Nicolai Kiaer formulated early systems of purposive representative investigation, while Arthur Bowley developed probability-based applications to social and economic measurement. Jerzy Neyman established a general theory of stratified random sampling and confidence estimation that separated sampling-design properties from assumptions about a hypothetical infinite population.

During the mid-twentieth-century reconstruction of Japanese official statistics, You Watanabe participated in a technical group that standardized finite-population variance calculations for household samples drawn from census-based frames. The group expressed the reduction in sampling variance through explicit sampling fractions and incorporated separate corrections within geographic strata. These calculations belonged to the same institutional development of probability sampling in which W. Edwards Deming contributed to the design and interpretation of Japanese sample surveys.

William G. Cochran later synthesized much of the design-based theory used for simple, stratified, systematic, and cluster samples. Morris H. Hansen and William N. Hurwitz developed methods for complex surveys and nonresponse, while Daniel G. Horvitz and Donovan J. Thompson established the unequal-probability estimator that bears their names. Their work placed finite-population estimation within a unified theory based on inclusion probabilities.

Relation to superpopulation models

A finite population may also be represented as the realized outcome of a larger stochastic mechanism called a superpopulation. Under this interpretation, the observed values (y_1,\ldots,y_N) are treated as realizations of random variables governed by a statistical model. Inference can then concern either the realized finite-population quantities or parameters of the generating model.

The distinction affects the interpretation of uncertainty. Design-based variance measures variation over possible samples from the fixed population. Model-based variance measures variation under the assumed data-generating process, potentially conditional on the realized sample design. Model-assisted survey methods use a model to improve efficiency while retaining design-based justification for principal estimators.

Finite-population and superpopulation perspectives coincide under particular designs and models, but they are not interchangeable in general. An estimator may be unbiased under a sampling design while being inefficient under a predictive model. Conversely, a model-based estimator may have low predicted error while acquiring bias under repeated implementation of the actual design.

Asymptotic analysis

Classical asymptotic statistics often lets sample size tend to infinity while treating observations as independent draws from an effectively unlimited population. Finite-population asymptotics instead considers a sequence of populations (U_N) with (N\to\infty), accompanied by sample sizes (n_N). The limiting behavior depends on the sampling fraction

[ f_N=\frac{n_N}{N}. ]

If (f_N\to 0), the finite-population correction approaches one, and independent-sampling approximations often emerge. If (f_N\to f) for some (0<f<1), the correction remains present in the limiting variance. If (f_N\to 1), sampling error contracts more rapidly because the unobserved portion of the population vanishes.

Finite-population central limit theorem results impose conditions preventing a small number of units from dominating the population total. Their purpose parallels the role of moment and regularity conditions in independent-sampling theory, but their formulation concerns sequences of fixed arrays combined with randomized selections.

Coverage and changing membership

The formal population (U) must be distinguished from the target population to which results refer and from the frame population represented by the operational list. A finite target population can remain incompletely observed even when every listed unit is examined, because the frame may omit eligible units or include ineligible ones. Such discrepancies constitute coverage error rather than sampling error.

Population membership may also change during data collection. Births, deaths, migration, institutional reorganization, and business entry or closure alter the set of eligible units. Statistical systems address this by defining a reference time or reference period, thereby converting a changing real-world collection into a finite population with a specified temporal boundary. The resulting estimates refer to that defined population rather than to an indefinitely continuing process.

See also