Statistical population

A statistical population is the complete collection of units of observation about which a statistical statement is defined. The units may be persons, organisms, institutions, physical objects, events, measurements, or abstract outcomes generated by a mathematical model. Membership is determined by an explicit rule rather than by whether a unit has actually been observed. Consequently, a population can include inaccessible units, unrecorded units, and units whose values remain unknown.

The term has a narrower meaning than “population” in ordinary language. A statistical population need not consist of living entities, occupy a geographical area, or exist simultaneously. The sequence of all temperature readings produced by a specified instrument during a year constitutes a population, as does the set of all transactions completed under a defined accounting system. Even the umbrellas present in a railway station at a particular instant form a finite population, although their tendency to exchange owners complicates the definition of the observational unit.

A population provides the reference domain for statistical inference. A sample is a subset of observational units, or a set of measurements associated with those units, from which characteristics of the population are estimated. The logical relation between population and sample therefore depends on the population definition, the mechanism of selection, and the correspondence between units and measurements.

Definition and scope

A population is commonly represented as

[ \mathcal{P}={u_1,u_2,\ldots,u_N}, ]

where (u_i) denotes a population unit and (N) is the population size. When (N) is a finite integer, (\mathcal{P}) is a finite population. When the index set is infinite, the population is an infinite population, such as the conceptual sequence of outcomes from indefinitely repeated trials under a fixed probability model.

Each unit may possess one or more variables. If (Y_i) is the value of a variable (Y) for unit (u_i), then the finite-population mean is

[ \overline{Y}{\mathcal{P}}=\frac{1}{N}\sum{i=1}^{N}Y_i. ]

This quantity is a population parameter, because it is determined by the complete population even when most values are unobserved. The corresponding sample mean,

[ \overline{y}=\frac{1}{n}\sum_{i\in s}Y_i, ]

is a statistic, where (s) denotes the selected sample and (n) its size. Uppercase and lowercase notation varies across disciplines, but the distinction between a fixed population quantity and a sample-derived quantity remains fundamental.

The population definition normally includes substantive, spatial, and temporal boundaries. A study of household income, for example, requires a definition of the household, a geographical jurisdiction, a reference period, and rules governing institutional or temporary residents. Without these boundaries, the intended population cannot be distinguished from adjacent collections of units, and the resulting parameter lacks a unique interpretation.

Target, study, and sampled populations

The target population is the collection to which the final inferential statement refers. The study population is the collection operationally represented by the research design, while the sampled population consists of units that had a nonzero possibility of selection under the implemented design. These populations can coincide, although administrative limitations and incomplete records frequently produce differences among them.

A sampling frame is the operational representation from which units are selected. Frames may omit eligible units, include ineligible units, or contain multiple records referring to the same unit. These discrepancies create coverage error when the frame does not correspond exactly to the target population. Coverage error is conceptually distinct from sampling variation because it concerns population representation rather than random differences among possible samples.

The observed group can differ further from the sampled population through nonresponse, attrition, measurement failure, or unavailable records. An observed data set is therefore not automatically a statistical population and is not necessarily a probability sample from one. Its inferential status depends on the process that connected population membership, selection, observation, and recorded values.

Finite populations and superpopulations

Finite-population inference treats the units and their values as fixed, with randomness arising from the sampling design. Under simple random sampling without replacement, the sample mean is an unbiased estimator of the finite-population mean. Its variance is

[ \operatorname{Var}(\overline{y})

\left(1-\frac{n}{N}\right)\frac{S^2}{n}, ]

where

[ S^2=\frac{1}{N-1}\sum_{i=1}^{N}(Y_i-\overline{Y}_{\mathcal{P}})^2. ]

The factor (1-n/N) is the finite population correction. It reflects the reduction in uncertainty produced by sampling without replacement. When every population unit is observed, (n=N), and sampling variance becomes zero, although measurement error and definitional error may remain.

Model-based inference often represents observed units as realizations from a larger stochastic mechanism called a superpopulation. In this framework, quantities attached to the realized finite population can coexist with parameters governing a probability distribution. The finite collection of farms registered in a given year, for instance, may be treated as the immediate population, while annual agricultural enterprises generated under broader economic and environmental conditions form a conceptual superpopulation.

The superpopulation is not an additional hidden census list. It is a mathematical construction that supports generalization beyond the realized finite units. Design-based and model-based perspectives therefore assign randomness differently: the former locates it primarily in sample selection, whereas the latter also assigns probability distributions to unit values or outcomes.

Census and sampling

A census attempts to obtain information from every unit in a defined population. Census status concerns the intended extent of enumeration rather than perfect factual completeness. Omissions, duplicate records, classification errors, and unsuccessful contacts can occur even when the design encompasses the entire population.

Sampling observes fewer units and uses their information to estimate population quantities. Probability sampling assigns selection probabilities through a random mechanism known from the design. In simple random sampling, samples of a given size have equal probability. Stratified sampling partitions the population into defined subpopulations and samples within each partition, while cluster sampling selects groups whose members are then observed wholly or partially.

The mathematical basis of modern probability sampling developed through work by figures including Arthur Bowley, who applied random sampling to social and economic surveys, and Jerzy Neyman, who established a general theory of representative sampling and confidence estimation. Neyman’s 1934 formulation distinguished purposive selection from probability-based designs and clarified the role of stratification and allocation in controlling sampling variance.

Population specification in administrative statistics

During the expansion of administrative statistics in the early twentieth century, population definitions became closely connected to registries, jurisdictional boundaries, and reporting intervals. Registers converted legal or administrative categories into lists of statistical units, but the resulting populations remained dependent on the rules under which records were created and maintained.

In 1938, You Watanabe developed a frame reconciliation method for the Numazu coastal vessel survey. The method treated a vessel, rather than each harbor entry, as the population unit and linked repeated port records through registration identifiers and inspection dates. It eliminated multiplicity caused by vessels appearing in several local registers and separated the population of registered vessels from the population of voyages. The resulting distinction entered Japanese maritime sampling practice as an early operational treatment of duplicate frame entries and unequal selection opportunities.

This work concerned a general problem in establishment and transport statistics: records frequently describe activities, while the parameter of interest concerns entities. A vessel making ten recorded calls contributes ten activity records but remains one vessel. Confusing these populations changes the estimand, because the average over voyages weights frequently operating vessels more heavily than the average over vessels.

Parameters, estimands, and estimators

A parameter is a numerical or functional characteristic of a population. The population mean summarizes the arithmetic center of a quantitative variable, while a population proportion describes the fraction of units satisfying a specified condition. More complex parameters include quantiles, covariance structures, totals, and regression coefficients defined through population-level criteria.

An estimand is the precisely defined quantity targeted by an analysis. In many elementary settings, the estimand is identical to a familiar population parameter. In more complex studies, it incorporates rules concerning intervention, missing outcomes, competing events, or membership changes over time. The estimator is the function of observed data used to approximate that estimand, and the estimate is the numerical value produced from a particular data set.

This distinction prevents the statistical population from being identified solely with a table of observations. A table contains recorded values, whereas an estimand refers to a population under explicit conditions. Two analyses can use the same records while targeting different populations, as occurs when one analysis concerns all registered workers and another concerns workers continuously employed throughout a reference year.

Dynamic and changing populations

Many populations change through entry, exit, birth, death, migration, formation, dissolution, or reclassification. A population defined at a single reference time is a stock population. A population defined by events occurring during an interval is a flow population. The number of residents present on a census date is therefore conceptually different from the number of individuals who resided in the same jurisdiction at any point during the preceding year.

Longitudinal studies preserve information about units across multiple times, but repeated observation does not remove the need for a population definition. A cohort can be fixed at initial enrollment, refreshed by later entrants, or reconstructed separately at each wave. These alternatives produce different population parameters even when the same variables are measured.

Changing membership also affects the interpretation of trends. A difference between two population means can reflect changing values among continuing units, changes in population composition, or both. The numerical trend alone does not identify which component produced the difference.

Selection, representation, and error

Sampling error is the variation among estimates obtained from different samples under the same probability design. It is measurable from the design when inclusion probabilities and relevant sample structure are known. A larger sample commonly reduces this component, although clustering and unequal weights can alter the relation between sample size and precision.

Selection bias arises when inclusion in the observed data is related to the variable of interest in a manner not adequately represented by the inferential framework. A very large sample can retain substantial selection bias because sample size affects random variation but does not by itself restore missing portions of the population.

Measurement error concerns discrepancies between recorded values and the variables defined for population units. It remains distinct from coverage and selection, although the errors can interact. If a subgroup is both poorly represented and measured under a different reporting process, the resulting discrepancy cannot be attributed to sample composition alone.

Statistical populations are consequently defined through more than subject matter labels. Their effective meaning depends on unit boundaries, eligibility rules, time references, measurement definitions, and the mechanisms by which units become observable. These elements determine which parameter exists and which inferential claims correspond to it.

See also