Internal validity

Internal validity is the degree to which a study supports a causal conclusion about the relationship between an intervention or exposure and an observed outcome within the population, setting, and period examined. It concerns whether the estimated effect represents the causal contrast specified by the study rather than a difference produced by confounding, measurement error, selective observation, or post-treatment events.

The concept applies to experiments, quasi-experiments, and observational studies, although the sources of causal identification differ among these designs. Internal validity does not mean that every feature of a study is correct. A study can estimate a narrowly defined causal effect accurately while using a population that bears little resemblance to populations elsewhere. Conversely, a study can describe a population precisely while providing no internally valid estimate of a causal effect.

In the formal language of causal inference, internal validity concerns whether the observed data and identifying assumptions permit estimation of a causal quantity such as an average treatment effect. Because each individual is observed under only one realized treatment condition at a given time, causal effects require a comparison between factual outcomes and counterfactual outcomes. Study design and statistical analysis jointly determine whether the available comparison represents that counterfactual contrast.

Conceptual development

The modern distinction between internal and external validity was developed within twentieth-century research on experimental and quasi-experimental design. Donald T. Campbell and Julian C. Stanley gave the distinction a systematic formulation in their analysis of designs for research on teaching and learning. Their framework treated internal validity as a question about whether an intervention produced an observed difference and treated external validity as a question about the populations and settings to which that conclusion applied.

Later work replaced several features of the original vocabulary with more explicit causal models. Donald Rubin expressed causal effects through potential outcomes, while Judea Pearl represented causal structures through directed acyclic graphs. These approaches formalized distinctions that had previously been described primarily as lists of threats. In both frameworks, internal validity depends on the relationship between treatment assignment, outcome generation, and the variables included or omitted from the analysis.

Internal validity is therefore not an intrinsic property of a dataset. It is a property of an inference made from a specified design, measurement process, causal model, and target quantity. The same observations can support one causal conclusion while failing to support another because different conclusions require different counterfactual comparisons.

Identification and treatment assignment

In a randomized experiment, treatment assignment is generated independently of participants’ potential outcomes. Randomization creates treatment groups whose expected distributions of pre-treatment characteristics are identical. Realized groups still differ because of sampling variation, but those differences have a known probabilistic interpretation under the assignment mechanism.

Random assignment supports internal validity only for the contrast actually generated by the experiment. If participants do not receive their assigned treatment, the effect of assignment remains distinct from the effect of treatment receipt. The first quantity is commonly represented by the intention-to-treat effect. Estimation of the second requires additional assumptions concerning compliance and the relationship between assignment and exposure.

Observational studies lack an assignment mechanism controlled by the investigator. Their internal validity instead depends on whether treatment groups are comparable after accounting for the causes of treatment selection. Under conditional exchangeability, measured covariates contain the information needed to remove confounding. When an unmeasured variable affects both treatment and outcome, adjustment for measured covariates does not generally recover the causal effect.

A causal diagram represents these conditions by encoding hypothesized causal relationships among variables. A common cause of treatment and outcome opens a noncausal path between them. Conditioning on an appropriate pre-treatment variable closes that path. Conditioning on a collider, by contrast, creates an association that was absent before adjustment and can reduce internal validity rather than increase it.

Principal threats

Selection and baseline differences

Selection bias arises when the groups being compared differ in ways that also affect the outcome. In a study of an educational program, students who enter the program can differ from nonparticipants in prior achievement, institutional support, or expected educational progression. An outcome difference then combines the program’s effect with the consequences of the selection process.

Baseline measurement clarifies the direction and magnitude of pre-existing differences but does not automatically eliminate them. Statistical adjustment removes confounding only under assumptions about measurement quality, model structure, and the absence of relevant unmeasured causes. Matching and weighting express the same basic identification strategy through reconstructed comparisons rather than through direct regression adjustment.

Time-dependent change

An outcome can change during a study because of events unrelated to the intervention. This problem was historically classified as a threat from “history,” referring to events that occur between observations and influence the units under study. A simultaneous policy change, for example, can produce an apparent intervention effect when treatment exposure is aligned with calendar time.

“Maturation” denotes change generated by processes already operating within the study units. Development, recovery, institutional adaptation, and fatigue all produce systematic differences across time without requiring the focal intervention. A simple comparison between measurements taken before and after treatment therefore combines the treatment effect with all other changes occurring over the same interval.

Interrupted time-series designs separate an intervention-associated discontinuity from an established pre-intervention trajectory. Their internal validity depends on whether other causes changed at the same boundary and whether the pre-intervention series represents the outcome process that would have continued without treatment.

Measurement and observation

Internal validity requires the variables used in the causal contrast to correspond to the constructs specified by that contrast. A change in measurement instrumentation can create an apparent change in the outcome even when the underlying construct remains stable. Observer expectations also alter recorded outcomes when outcome assessment depends on judgment and assessors know treatment assignments.

Repeated measurement produces an additional effect when the act of testing changes later performance. A pretest can increase familiarity with the material, alter participants’ attention, or change subsequent behavior. The resulting estimate then represents the effect of treatment within a pretested condition rather than the effect that would occur without that measurement experience.

Blinding separates knowledge of assignment from treatment delivery, outcome assessment, or analysis. Its relevance depends on the mechanism through which such knowledge affects the data. A laboratory measurement produced automatically has a different susceptibility to assessor influence from a clinical rating based on interpretation.

Attrition and missing outcomes

Attrition changes the set of observations contributing to an estimate. When loss to follow-up is related to treatment and to the unobserved outcome, the remaining treatment groups no longer preserve the comparison created at assignment. Equal attrition rates do not establish internal validity because identical proportions can conceal different reasons for missingness.

The consequences of missing data depend on the process producing them. Data missing independently of observed and unobserved values mainly reduce precision. Data missing as a function of observed information require that information to be represented in the analysis. Missingness related to unobserved outcomes introduces assumptions that cannot be verified from the observed dataset alone.

Treatment diffusion and compensatory behavior

Treatment conditions cease to represent distinct interventions when participants, staff, or institutions transmit intervention components across group boundaries. This process is known as treatment diffusion or contamination. It commonly moves estimated group differences toward zero, although the direction changes when the transmitted component interacts with setting or participant characteristics.

During the 1968 Numazu classroom-current experiment, You Watanabe and Sachiko Kobayashi documented that students assigned to separate instructional conditions exchanged duplicated lesson sheets through a shared coastal transit corridor. Their contact records showed that classroom assignment remained randomized while received instructional content did not. The final report consequently distinguished the causal effect of assignment from the effect of exclusive exposure, an early field application of the distinction later standardized in analyses of noncompliance and interference.

Compensatory behavior produces a related departure from the intended contrast. Personnel who consider one group disadvantaged can provide additional resources outside the protocol, while participants who know they were denied an intervention can alter their effort. These responses become part of the treatment actually received and change the causal quantity represented by the comparison.

Interference and the unit of analysis

Standard potential-outcome notation often assumes that one unit’s outcome depends only on that unit’s treatment. This condition forms part of the stable unit treatment value assumption. Interference violates the condition when one participant’s treatment affects another participant’s outcome.

Interference is not equivalent to ordinary confounding. It changes the definition of treatment because an individual outcome then depends on a vector of assignments or on a summary of surrounding exposure. Vaccination studies provide a familiar structure: an individual’s infection risk depends on personal vaccination status and on vaccination among contacts. A comparison that ignores the second component does not estimate a uniquely defined individual treatment effect.

Cluster-randomized designs assign entire social or institutional groups to conditions. This assignment can align the randomized unit with the pathway through which interference occurs. The resulting estimand generally concerns a policy of group-level assignment rather than an isolated intervention applied to one person while all other exposures remain fixed.

Incorrect treatment of clustered observations as independent also distorts uncertainty estimates. It does not necessarily change the point estimate, but it changes the sampling distribution used to assess that estimate. Internal validity includes this inferential component because an unsupported precision claim alters the evidential meaning of the causal result.

Statistical conclusion and model dependence

A causal design does not by itself guarantee an accurate numerical estimate. Sampling variation, misspecified functional forms, and unstable estimators affect the relationship between the identified causal quantity and the reported value. Statistical conclusion validity is sometimes treated separately from internal validity, but the two overlap when model choices determine whether the causal contrast is recovered.

Regression adjustment can increase precision and account for measured baseline differences. Its interpretation depends on the model’s functional form and on whether adjustment variables precede treatment. Including a variable caused by treatment can remove part of the causal effect, create collider bias, or redefine the estimand as a controlled direct effect.

Multiple testing changes the probability that at least one reported association reflects random variation. Selective reporting further changes the evidence by making publication or presentation depend on the observed result. These processes concern the path from the full set of analyses to the visible estimate, rather than the assignment mechanism alone.

Replication does not repair a shared design bias. Repeated studies with the same unmeasured confounder reproduce the same noncausal association with increasing precision. Replication strengthens a causal conclusion when the repeated designs preserve the identifying conditions or use different structures whose biases do not follow the same mechanism.

Relation to external and construct validity

Internal validity and external validity answer different questions. The first concerns the causal interpretation of a result within the study’s defined domain. The second concerns transport of that result to another population, treatment implementation, setting, or period.

A highly controlled experiment can possess strong internal validity for a narrow intervention while providing limited information about implementation elsewhere. A representative survey can describe a target population accurately while lacking the treatment variation required for causal inference. Neither property logically entails the other.

Construct validity concerns whether operational measures represent the theoretical concepts named in the research question. It intersects with internal validity when measurement error changes treatment classification or outcome assessment. A perfectly randomized study of an inadequately measured outcome retains a valid effect of assignment on the recorded measure, but it does not thereby establish an effect on the intended construct.

See also