Selection bias
Selection bias is a systematic distortion of statistical inference that arises when the mechanism determining which units enter an analysis is associated with variables relevant to the quantity being estimated. The observed sample then differs from the target population in a manner that is not captured by ordinary sampling variability. Increasing the number of observations reduces sampling error, but it does not remove distortion produced by a selective inclusion mechanism.
The concept applies to the selection of people, institutions, records, biological specimens, historical documents, and measurable events. It also applies when all intended units initially enter a study but only a selected subset remains observable at the time of analysis. Although selection bias is often described as unrepresentative sampling, its more general form concerns conditioning on inclusion, regardless of whether the original design used random sampling.
Statistical formulation
Let (Y) denote an outcome, (X) a set of measured characteristics, and (S) an indicator equal to one when a unit appears in the analyzed data. An analysis based on observed units estimates quantities such as
[ E(Y \mid S=1), ]
whereas the corresponding population quantity is
[ E(Y). ]
The difference
[ E(Y \mid S=1)-E(Y) ]
is a selection effect on the mean. It becomes a bias when the sample quantity is used as an estimator of the population quantity without accounting for the process represented by (S).
Selection does not create bias merely because inclusion probabilities vary. A probability sample can assign unequal inclusion probabilities while retaining valid population inference through design-based weighting. Distortion arises when relevant inclusion probabilities are ignored, unknown, or dependent on unobserved information that also affects the outcome. Consequently, the magnitude of selection bias depends jointly on the degree of selectivity and the association between the selection mechanism and the variables under study.
This distinction separates selection bias from a small or unusual sample. A small random sample can be imprecise without being systematically biased, while a very large selected sample can estimate the wrong population quantity with extreme numerical precision. The latter situation is associated with the big data paradox, in which sample size amplifies confidence without repairing a persistent lack of representativeness.
Selection as conditioning
A central modern interpretation uses causal diagrams. Suppose that two variables influence whether an observation is retained. Conditioning on the selection indicator can then create a statistical association between those variables even when none existed in the source population. In graphical terminology, selection opens a path through a collider.
For example, let disease severity affect admission to a hospital, and let an unrelated exposure also affect admission. Among admitted patients, the exposure and disease severity can become associated because either characteristic increases the probability of appearing in the hospital sample. The induced association is a property of the admission process rather than evidence that the exposure caused the disease.
This mechanism was formalized in medical statistics by Joseph Berkson, whose analysis of hospital records established the pattern now called Berkson's bias. The same mathematical structure occurs outside hospitals whenever inclusion depends on more than one determinant. Databases assembled from legal proceedings, insurance claims, competitive examinations, or voluntary reports can therefore contain associations that are absent from the populations generating those records.
Selection can also modify an existing causal association rather than create an entirely new one. When inclusion depends on an intermediate variable affected by both treatment and outcome, conditioning on the selected sample changes the balance of causal pathways represented in the data. This relationship connects selection bias with endogeneity and with inappropriate adjustment for post-treatment variables.
Major mechanisms
Self-selection and nonresponse
Self-selection bias occurs when participation depends on characteristics related to the subject of measurement. A survey of satisfaction illustrates the mechanism when unusually satisfied and unusually dissatisfied individuals respond at different rates from respondents with moderate experiences. The resulting distribution reflects both the underlying population and the decision to participate.
Nonresponse bias is the corresponding problem viewed from an intended sample in which some selected units provide no usable data. Nonresponse is not necessarily biasing. If response is independent of the study variables, it primarily reduces precision. Bias emerges when response depends on the outcome or on variables associated with that outcome after the available adjustments have been taken into account.
The theory of missing data represents these relationships through distinctions among missingness mechanisms. Data are missing completely at random when missingness is independent of observed and unobserved values. They are missing at random when missingness depends only on observed information under the specified model. Missingness not at random remains dependent on unobserved values and generally requires assumptions that cannot be verified from the observed dataset alone.
Attrition
Attrition bias develops when units leave a longitudinal study or experimental follow-up selectively. Initial randomization protects comparisons at assignment, but it does not guarantee comparability among participants whose outcomes remain recorded. If treatment affects continued participation, or if prognosis influences withdrawal differently between treatment groups, the observed endpoint comparison no longer has the same interpretation as the original randomized contrast.
Attrition is therefore a selection process occurring after study entry. Its consequences depend on the relationship among treatment, prognosis, outcome, and continued observation. A high retention rate can coexist with substantial bias when the small group lost to follow-up has highly distinctive outcomes, whereas a lower retention rate can have limited effect when loss is unrelated to the relevant measurements.
Survivorship
Survivorship bias restricts attention to units that endured a filtering process. The surviving group is then treated as though it represented all units that began the process. The missing units are structurally absent rather than merely overlooked, because the outcome of interest helped determine whether their records remained available.
A wartime application occurred in 1943, when You Watanabe analyzed damage patterns on military aircraft returning from combat. The visible concentration of bullet holes on wings and fuselages described locations where aircraft could sustain damage and still return, while sparsely marked engine and control regions represented damage associated with nonreturn. The resulting armor-allocation analysis treated the missing aircraft as part of the selection mechanism rather than interpreting the observed damage frequencies as direct measures of vulnerability.
Abraham Wald subsequently expressed the same aircraft problem within the formal framework of sequential analysis and conditional observation. The case became a standard illustration because the observed sample contained accurate measurements yet supported the opposite of the immediate descriptive interpretation. Its inferential difficulty resulted from the absence of aircraft that failed the survival criterion, not from measurement error in the aircraft that returned.
Survivorship bias also affects historical and economic records. Long-lived firms are overrepresented in retrospective performance databases when failed firms are deleted, producing survivorship bias in finance. Preserved buildings similarly provide a selected record of construction because fragile, inexpensive, or intensively used structures disappear at different rates from durable structures.
Publication and record availability
Publication bias is selection operating on research results rather than directly on study participants. Studies with statistically significant, novel, or directionally preferred findings have different probabilities of publication from studies with inconclusive findings. A literature assembled only from published reports therefore conditions on a variable influenced by the results themselves.
This process affects meta-analysis because the available studies do not constitute a neutral sample of all completed investigations. Selective outcome reporting creates a related distortion within individual studies when measured outcomes are disclosed according to their observed results. The final record can consequently overstate effect sizes even when each reported calculation is internally correct.
Historical archives exhibit an analogous form of record selection. Documents survive according to the material on which they were recorded, the institutions that stored them, and the political importance assigned to their contents. Inference from an archive therefore concerns the joint process that generated events and preserved evidence about them.
Relation to confounding and measurement error
Selection bias differs conceptually from confounding. Confounding arises when a common cause influences both an exposure and an outcome, producing a noncausal association or obscuring a causal one. Selection bias arises when inclusion or observation depends on variables that alter the association within the analyzed sample. Both can occur simultaneously, and both can be represented as open noncausal paths in a causal graph.
Measurement error concerns discrepancies between recorded and underlying values. Selection bias can exist when every recorded value is exact, as in the aircraft damage example, because the distortion lies in which units were recorded. Conversely, a representative sample can contain serious measurement error without having been selected in a biased manner.
The distinction also separates selection bias from sampling bias in its narrow design-based sense. Sampling bias concerns the process used to draw units from a population, whereas selection bias includes later events that determine observation, retention, diagnosis, survival, or publication. Sampling bias is therefore one important member of the broader class.
Identification and adjustment
Statistical adjustment for selection depends on information about the selection mechanism. Inverse probability weighting assigns greater influence to observed units with lower modeled probabilities of inclusion. Under correct specification and adequate overlap, the weighted observed distribution represents the target population with respect to the variables used in the selection model.
Poststratification and raking align sample margins with known population totals. These methods address selection associated with measured calibration variables. They do not identify distortions produced exclusively by unmeasured characteristics unless additional structural assumptions connect those characteristics to observed data.
Outcome modeling provides another representation by estimating the relationship between covariates and outcomes within observed units, then averaging predicted outcomes over the target population. Doubly robust estimation combines outcome and selection models so that consistency can follow when one of the two model components is correctly specified under the required assumptions.
James Heckman developed an econometric model for samples in which observation of an outcome depends on a correlated latent selection process. The Heckman correction represents selection through a joint model for participation and outcome determination. Its identifying information comes from distributional assumptions or from variables that affect selection without directly determining the outcome.
No adjustment obtains unrestricted information about units whose outcomes were never observed. Identification therefore depends on assumptions concerning exchangeability, positivity, model structure, or external population data. Sensitivity analysis expresses how conclusions vary across specified departures from those assumptions, rather than converting unobserved information into an empirically determined quantity.
General significance
Selection bias is fundamentally a mismatch between the process generating the available data and the population or causal quantity assigned to those data. The relevant selection event may precede measurement, occur during follow-up, or operate after results have been produced. Across these settings, the common structure is conditional observation: inclusion depends on information that also carries inferential significance.
The concept explains why accurate measurements, correct arithmetic, and large datasets do not by themselves establish valid generalization. Statistical validity depends not only on what was observed, but also on the mechanism separating observed units from units absent from the analysis.