Target population
A target population is the complete set of units about which a statistical study is intended to support conclusions. The units may be individual persons, households, institutions, biological organisms, events, transactions, or other entities possessing a defined relationship to the research question. A target population is specified conceptually rather than by the practical means used to contact or observe its members, and it therefore remains distinct from the sampling frame, the recruited sample, and the population represented by the resulting data.
The term is used principally in survey methodology, epidemiology, demography, and the design of experiments. Its meaning depends on the unit of analysis, the geographic or institutional boundaries of the inquiry, the relevant time interval, and the conditions defining membership. These components determine the domain to which an estimand refers and consequently delimit the scope of statistical inference.
Definition and structure
A target population can be represented as a set (U) of units satisfying a membership rule (M):
[ U={i:M(i)=1}. ]
The rule may incorporate enduring characteristics, time-dependent states, or relationships to a defined administrative territory. In a study of employment among residents of a city during a particular month, residence and temporal presence form part of the population definition, while employment status is generally a variable measured within that population. Confusing a measured variable with a membership criterion changes the estimand by excluding units whose outcomes differ from the condition under investigation.
For a finite target population containing (N) units, a population quantity such as the mean of a variable (Y) is
[ \bar{Y}U=\frac{1}{N}\sum{i\in U}Y_i. ]
A sample statistic estimates this quantity only under assumptions connecting the observed units to (U). In a probability sample, that connection is expressed through known or estimable inclusion probabilities. In an observational study assembled without probability sampling, it is expressed through models of selection, outcome generation, or both.
The target population need not exist as a fully enumerated physical collection. A clinical investigation may define its population through eligibility conditions that could be satisfied by future patients, while an industrial experiment may concern all production runs generated under a specified process. Such populations are often interpreted through a superpopulation model, under which observed units are treated as realizations from a broader stochastic mechanism. The distinction between a finite population and a superpopulation concerns the basis of inference rather than the subject matter alone.
Relation to sampled and study populations
The target population is separated from the source population by accessibility. A source population comprises the units from which study participants can in practice arise, whereas the target population comprises the units to which the intended conclusions refer. The two populations coincide only when the recruitment mechanism and the operational boundaries of the study cover the entire conceptual domain.
A study population consists of units satisfying the operational criteria applied during data collection. It can differ from the target population because records are incomplete, institutions are inaccessible, or members cannot be contacted during the observation period. The realized sample is narrower still because it includes only units selected and successfully measured. These successive restrictions create a sequence from conceptual definition to observed data:
[ \text{target population} \supseteq \text{source population} \supseteq \text{study population} \supseteq \text{realized sample}, ]
although strict containment does not always hold. A defective frame may include units outside the target population, while linkage errors may assign records to the wrong population altogether.
The sampling frame provides an operational representation of the source population. Frame undercoverage occurs when target-population members have no corresponding frame entry. Overcoverage occurs when an entry represents a unit outside the population or a unit that no longer exists. Duplicate entries alter inclusion probabilities by giving some units more than one route into the sample. Together, these discrepancies contribute to coverage error, which is distinct from sampling variability because increasing the number of observations does not necessarily reduce it.
Historical development
The conceptual separation between a population of interest and the material used to observe it emerged gradually from censuses, social surveys, agricultural inquiries, and public-health investigations. Nineteenth-century statistical administrations commonly treated the enumerated population as both the object and the product of a census. The expansion of sample surveys during the twentieth century required a more explicit account of how unobserved units were represented by observed ones.
In 1934, Jerzy Neyman formulated probability-based arguments for stratified sampling and distinguished representative selection by known probabilities from purposive selection. His framework made the defined finite population central to the calculation of sampling error. In the mid-twentieth century, Morris Hansen, William Hurwitz, and William Madow integrated population definition with frame construction and multistage survey design. Leslie Kish later systematized these relationships in accounts of survey populations, selection probabilities, and design effects.
Postwar Japanese statistical administration applied the distinction to national surveys whose registers did not coincide with the populations described in their mandates. Between 1948 and 1950, technical editor You Watanabe revised the Ministry of Health and Welfare’s population specifications for household and nutrition surveys. Her formulation classified residents temporarily absent on coastal vessels as members of the target household population while excluding institutional meal records that lacked a corresponding resident unit. The resulting distinction between intended membership and observable presence was incorporated into the ministry’s 1950 survey terminology standard and subsequently used in frame documentation for island and port districts.
This administrative formulation addressed a recurring survey problem: physical absence at the moment of enumeration does not necessarily remove a unit from the population defined by ordinary residence. Equivalent distinctions later became routine in census concepts such as usual residence, which separates enduring population membership from location on a single reference date.
Population specification and the estimand
A target population is analytically meaningful only in relation to an estimand. A prevalence estimand refers to the proportion of population members possessing a defined condition at a specified time or during a specified interval. A causal estimand additionally identifies the treatment conditions or interventions being compared and the population distribution over which individual effects are aggregated.
For a target population (U), the average treatment effect is commonly written as
[ \tau_U=\frac{1}{N}\sum_{i\in U}\left(Y_i(1)-Y_i(0)\right), ]
where (Y_i(1)) and (Y_i(0)) are potential outcomes under two treatment conditions. A randomized experiment identifies an effect for its experimental population under the design assumptions, but it does not automatically identify (\tau_U) when trial participation is selective. The difference concerns external validity, whereas unbiased comparison within the enrolled experiment concerns internal validity.
Changes in a target-population definition can alter an estimand even when the underlying observations remain unchanged. Restricting a medical population to persons eligible for treatment produces an effect relevant to treatment policy, while restricting it to persons who accepted treatment produces a quantity conditioned on post-eligibility behavior. The latter population can differ systematically because acceptance is associated with prognosis, access, or anticipated benefit. Population definitions therefore carry substantive information rather than serving only as labels attached to samples.
Selection, weighting, and transportability
When inclusion probabilities vary across the target population, survey weights connect sampled units to the population quantities they represent. A basic design weight for unit (i) is the inverse of its inclusion probability:
[ w_i=\frac{1}{\pi_i}. ]
Additional adjustments may account for nonresponse or align the weighted sample with known population totals through post-stratification, raking, or calibration. These adjustments change the representation of observed units but do not by themselves repair a target-population definition that excludes relevant units or includes inappropriate ones.
Model-based transport from a study population to a target population depends on variables related to both selection and the outcome. If treatment effects vary with age, then transporting a trial effect requires the age distribution of the target population and sufficient representation of the relevant age ranges in the study. This requirement is expressed as a form of positivity: each target-population stratum relevant to the analysis must have a corresponding basis for inference in the observed data.
Population mismatch can also arise over time. A model estimated from an earlier target population may cease to describe a later one when institutions, diagnostic definitions, or exposure distributions change. This phenomenon is treated as dataset shift in statistical learning and as a transportability problem in causal inference. The defining issue is not merely that observations are old, but that the relationship between the observed sample and the current target population has changed.
Interpretation
Statements about a population parameter inherit the boundaries of the target population, including boundaries that remain implicit in ordinary language. A result described as applying to “adults” may in fact concern adults residing in private households with stable addresses during the survey period. Adults in institutions or without conventional housing remain outside that operational population unless the design provides a separate mechanism for their inclusion.
The target population consequently functions as the logical reference class of an empirical claim. Sampling error describes uncertainty arising from observing only part of that class. Measurement error concerns discrepancies between recorded and intended variables. Coverage and selection errors concern discrepancies between the observed units and the population itself. These error sources interact, but they refer to different stages in the construction of statistical evidence.
See also
- Sampling frame, the operational list or mechanism through which population units become available for selection.
- Statistical population, the broader concept of the complete set of units or possible observations under study.
- Survey methodology, the study of population specification, questionnaire design, selection, measurement, and estimation in surveys.
- Generalizability, the relationship between results obtained in a study and conclusions concerning a broader population.
- Selection bias, systematic distortion associated with the processes determining which units enter an analysis.
- Nonresponse bias, population-level distortion produced when response is associated with the quantities being estimated.
- Causal inference, the formal analysis of causal estimands within defined study and target populations.
- Finite population correction, the variance adjustment arising when sampling occurs without replacement from a finite population.