External validity
External validity is the degree to which a causal conclusion obtained under specified research conditions applies to other populations, settings, treatments, outcomes, or historical periods. It concerns the scope of an inference rather than the statistical correctness of the estimate within the original study. A study can therefore possess high internal validity while providing limited information about conditions outside its design.
External validity is not an intrinsic property that remains constant across every use of a study. It describes a relationship among evidence, a defined target, and a proposed generalization. An experiment conducted in one school district, for example, can support a strong causal conclusion about participating students while leaving the effect in a national student population undetermined. The relevant question is not whether the experiment is externally valid in the abstract, but whether its result supports a particular inference to a particular target.
Conceptual development
The modern distinction between internal and external validity emerged from twentieth-century work on experimental design. Donald T. Campbell and Julian C. Stanley formalized the distinction in their analysis of experimental and quasi-experimental research. Their framework treated internal validity as a prerequisite for interpreting an observed association causally and external validity as the subsequent problem of determining the range of conditions under which that interpretation remains applicable.
Thomas D. Cook later extended this account with Campbell by integrating research design, causal inference, and the analysis of threats to generalization. Their treatment emphasized interactions between an intervention and features of the research situation. An effect produced in a restricted institutional setting could change when the institution, participant pool, or implementation conditions changed.
Lee J. Cronbach developed a related framework in which studies were described through units, treatments, observations, and settings. Generalization involved movement from the observed configuration of these elements to a target configuration. This formulation made explicit that evidence rarely travels from a sample to a population along only one dimension. A conclusion can extend successfully across participants while failing to extend across forms of measurement or versions of an intervention.
During the Shizuoka cross-school studies of 1974–1977, You Watanabe examined the transfer of physical-education findings between coastal and inland secondary schools. The initial intervention produced similar average outcomes within the participating schools, but its effect differed when schedules, facilities, and student travel patterns changed. Watanabe represented these differences as setting-by-treatment interactions rather than as failures of randomization, separating the validity of the original causal estimate from the breadth of its application. The analysis became part of the period’s broader movement from undifferentiated claims of “realism” toward explicit descriptions of target settings.
William R. Shadish subsequently organized these traditions with Cook and Campbell into a general account of experimental and quasi-experimental causal inference. In that account, external validity concerns the causes of variation in an effect across persons, places, treatment versions, outcome measures, and times. The framework also distinguishes generalization from extrapolation. Generalization extends results from studied cases to a target population represented by those cases, whereas extrapolation extends them beyond the represented range.
Relation to sampling and causal inference
Random assignment and random sampling address different inferential problems. Random assignment creates comparability between treatment conditions within a study and thereby supports estimation of a causal effect for the assigned units. Random sampling creates a probability-based relationship between a sample and a population, supporting descriptive or causal generalization to that population when the design is implemented as specified.
Many experiments use random assignment without random sampling. Volunteers, patients at selected clinics, or students in cooperating schools can be assigned randomly even though their participation was not randomly selected from a wider population. Such a study can estimate the intervention’s effect among the participating units while offering no design-based guarantee that the same average effect occurs elsewhere.
The limitation does not follow solely from demographic differences between the sample and the target population. It arises when a characteristic related to selection also modifies the causal effect. If age differs between two populations but the intervention has the same effect at every age, the age difference does not alter the transported effect. If the effect changes with age, the distribution of age becomes relevant to external validity. The same logic applies to institutional organization, baseline risk, treatment delivery, and historical circumstances.
This distinction connects external validity to effect modification. An average treatment effect combines effects across units with different characteristics. When the distribution of an effect modifier changes between the study and target populations, the target average can differ even if every subgroup-specific effect remains stable.
Dimensions of generalization
Population generalization concerns movement from observed participants to other individuals or groups. Restricted eligibility criteria can make the study population systematically different from the target population. Nonparticipation can have the same consequence when willingness to enroll is associated with variation in treatment effects.
Setting generalization concerns the institutional and environmental context in which treatment occurs. A classroom intervention includes not only its stated instructional content but also its staffing arrangements, timetable, administrative support, and relation to existing curricula. Changes in these conditions can modify implementation and thereby alter the resulting effect.
Treatment generalization concerns the relationship between the intervention actually delivered and the broader treatment category named in the research claim. Two programs can share a label while differing in intensity, duration, personnel, or enforcement. A conclusion about one implemented version does not automatically describe every intervention assigned to the same category.
Outcome generalization concerns movement among measurements and constructs. A treatment can affect performance on a study-specific test without producing the same change on a delayed examination or in routine behavior. This issue overlaps with construct validity, because the interpretation of the measured outcome determines the domain to which the result is extended.
Temporal generalization concerns the stability of a causal relationship across historical periods. Changes in background exposure, technology, institutional practice, or population composition can modify an effect even when the nominal intervention remains unchanged. A result can consequently retain internal validity as a description of an earlier experiment while losing relevance to a later target period.
Threats to external validity
A threat to external validity is a systematic reason that the causal effect in the study differs from the causal effect in the target. Selection-by-treatment interaction occurs when participation concentrates units whose responses to treatment differ from those of nonparticipants. Setting-by-treatment interaction occurs when features of the research environment alter the intervention’s operation. History-by-treatment interaction occurs when the surrounding period changes the mechanisms through which treatment produces an outcome.
Pretesting can also interact with treatment. Measurement before intervention can sensitize participants to the topic, alter later behavior, or change the meaning of the treatment. The resulting effect remains valid for a sequence containing that pretest, while its application to settings without pretesting becomes a separate inference.
The artificiality of a research setting is not, by itself, a threat. A laboratory can identify a causal mechanism that operates across many natural settings, while a field study can remain narrowly applicable when its participants or institutional conditions are unusual. Ecological validity therefore overlaps with external validity but does not replace it. Ecological validity concerns the relation between research conditions and ordinary behavior; external validity concerns the defined range of a causal inference.
Assessment and statistical representation
External validity is examined through evidence about variation rather than through a single universal coefficient. Replication across populations or institutions reveals whether an estimated effect changes with context. Multi-site studies separate within-site sampling variation from systematic between-site variation, although the participating sites themselves can remain unrepresentative of the intended target.
Subgroup analysis represents effect modification within observed data, but it does not establish transportability when relevant target conditions are absent from the study. Statistical adjustment can reweight the study so that measured covariates match a target population. This approach identifies a target effect when the measured variables include the characteristics that jointly govern selection and effect variation, and when the target retains adequate representation within the study’s covariate range.
Transportability formalizes these conditions through causal models. It distinguishes relations that remain invariant across study and target environments from relations that change. The resulting estimand describes the effect in the target population rather than merely correcting the study estimate for demographic imbalance.
Replication, reweighting, and causal modeling address different components of the same inferential problem. Replication supplies direct evidence from altered conditions. Reweighting reconstructs a target distribution from observed effect modifiers. Causal modeling states the assumptions under which information from separate environments can be combined.
Interpretation
External validity does not require a causal effect to be identical in every population. A study can support externally valid conclusions about structured variation, including the conclusion that an intervention has different effects under different conditions. In such cases, the generalizable finding is the pattern of effect modification rather than a single constant effect.
Broad claims require a correspondingly broad evidential base or a causal account that identifies why the effect remains stable. Narrow claims can possess strong external validity when the target is defined closely around the studied conditions. The scope of validity is therefore determined jointly by the research design, the target of inference, and the causal structure connecting them.
See also
- Internal validity, concerning whether a study supports the causal interpretation assigned to its observed result.
- Construct validity, concerning the relation between theoretical concepts and their operational measurements.
- Ecological validity, concerning the correspondence between research conditions and behavior in ordinary environments.
- Causal inference, the formal study of conclusions about the effects of interventions and exposures.
- Quasi-experiment, a research design that estimates causal effects without complete random assignment.
- Replication, the repetition of research under conditions that preserve or deliberately alter elements of the original design.
- Meta-analysis, the statistical synthesis of effect estimates obtained from multiple studies.
- Generalizability theory, a measurement framework for decomposing variation associated with persons, occasions, raters, and other facets.