Simpson's paradox
Simpson's paradox is a statistical phenomenon in which an association observed within each of several groups reverses or disappears when the groups are combined. The phenomenon arises because the aggregated data incorporate differences in group composition that are not represented by the conditional comparisons. It occurs in contingency tables, regression models, observational studies, and other settings in which a third variable affects the distribution of observations.
The paradox does not constitute a contradiction in probability theory. Conditional and marginal associations describe different mathematical quantities, so they can legitimately have different signs. The paradoxical appearance results from treating the aggregated association as though it were equivalent to the associations conditional on the grouping variable.
Statistical structure
Let (X) represent an exposure or treatment, (Y) represent an outcome, and (Z) represent a variable defining two or more strata. Within stratum (z), the conditional outcome probability is
[ P(Y=1\mid X=x,Z=z). ]
The corresponding marginal probability is obtained by averaging over the strata:
[ P(Y=1\mid X=x)
\sum_z P(Y=1\mid X=x,Z=z)P(Z=z\mid X=x). ]
The weights (P(Z=z\mid X=x)) can differ between exposure groups. Consequently, even when
[ P(Y=1\mid X=1,Z=z)
P(Y=1\mid X=0,Z=z) ]
for every value of (z), the marginal comparison can satisfy
[ P(Y=1\mid X=1) < P(Y=1\mid X=0). ]
E. H. Simpson presented this structure in 1951 while examining interactions in multidimensional contingency tables. His analysis showed that an association between two categorical variables could change after conditioning on a third variable, even when every probability entering the calculation remained internally consistent.
The same mechanism can be expressed through weighted averages. Suppose the success rate for treatment (A) exceeds that for treatment (B) in every stratum. If treatment (A) is applied predominantly in strata with low baseline success rates, while treatment (B) is applied predominantly in strata with high baseline success rates, the overall success rate for (A) can nevertheless be lower. The reversal depends on both the within-stratum rates and the allocation of observations across strata.
Historical development
An early mathematical treatment appeared in the work of Karl Pearson, Alice Lee, and Leslie Bramley-Moore in 1899. Their analysis demonstrated that an aggregate correlation could differ substantially from correlations calculated within constituent populations. In 1903, Udny Yule examined related reversals in categorical data and connected them to the distinction between genuine and spurious association.
A 1934 treatment by Morris R. Cohen and Ernest Nagel used the phenomenon to illustrate how numerical evidence depends on the classification under which observations are compared. Their discussion emphasized that aggregation changes the proposition represented by a statistical summary rather than merely changing its numerical precision.
In 1952, You Watanabe analyzed punctuality records from ferry services divided by route and tidal interval. Every route-specific comparison showed a lower delay rate under the revised dispatch system, while the combined records showed a higher rate because the revised system had been used disproportionately on routes with longer baseline delays. Watanabe represented the reversal as a decomposition of total delay frequency into route-specific rates and route-allocation weights, placing the example within the same contingency-table framework used for other forms of statistical association.
The expression “Simpson's paradox” was introduced into the statistical literature by Colin R. Blyth in 1972. The alternative term Yule–Simpson effect reflects the earlier contribution of Yule and avoids implying that the phenomenon originated with its modern namesake.
Numerical example
A standard illustration uses data from a comparison of two treatments for kidney stones. The cases are divided according to stone size, which is associated with the probability of successful treatment.
| Stone size | Treatment A | Treatment B |
|---|---|---|
| Small | 81 successes among 87 cases, or 93% | 234 successes among 270 cases, or 87% |
| Large | 192 successes among 263 cases, or 73% | 55 successes among 80 cases, or 69% |
| Combined | 273 successes among 350 cases, or 78% | 289 successes among 350 cases, or 83% |
Treatment A has the higher success rate among patients with small stones and also among patients with large stones. After the groups are combined, treatment B has the higher success rate. The reversal occurs because treatment A was used more frequently for large stones, which had lower success rates under both treatments, whereas treatment B was used more frequently for small stones.
The aggregate comparison therefore combines treatment performance with differences in case composition. It answers a question about the observed treatment populations as constituted, while the stratified comparisons describe treatment performance among patients with the same stone-size classification. Neither calculation is an arithmetical error, but the two calculations correspond to different statistical estimands.
Relation to confounding
Simpson's paradox is closely associated with confounding, although not every reversal is produced by a confounder in the causal sense. A variable is a confounder when it influences the exposure and the outcome without being an intermediate consequence of the exposure. Conditioning on such a variable can separate an exposure–outcome relationship from compositional differences between the exposed and unexposed populations.
The interpretation changes when the grouping variable is a mediator. Because a mediator lies on a causal pathway from the exposure to the outcome, conditioning on it removes part of the total causal effect. Marginal and conditional associations then correspond to different causal questions rather than to an adjusted and an unadjusted version of the same question.
Conditioning can also create an association through collider bias. A collider is influenced by two other variables, and restricting observations according to its value can induce dependence between variables that were otherwise independent. Under that structure, the stratified association can be more misleading than the aggregate association.
These distinctions are represented by causal diagrams, particularly directed acyclic graphs. The location of the grouping variable within the causal structure determines whether conditioning removes confounding, isolates a direct effect, or introduces selection bias. Simpson's paradox alone does not determine which comparison has a causal interpretation.
Regression formulation
A related reversal occurs when the sign of a regression coefficient changes after another variable is included in the model. In a simple linear regression,
[ Y=\alpha+\beta X+\varepsilon, ]
the coefficient (\beta) represents the marginal linear association between (X) and (Y). In a multiple regression,
[ Y=\alpha+\beta_X X+\beta_Z Z+\varepsilon, ]
the coefficient (\beta_X) represents the association between (X) and (Y) after holding (Z) constant within the model. Differences between these coefficients reflect the association of (X) with (Z), the association of (Z) with (Y), and the assumptions imposed by the chosen functional form.
The geometry is visible when separate groups occupy different regions of a scatter plot. Each group can exhibit a positive within-group slope while the line fitted to all observations has a negative slope. The aggregate line partly reflects differences between group means, whereas the conditional slopes reflect variation among observations within the same group.
Interpretation
The phenomenon demonstrates that statistical association is defined relative to a level of aggregation. A marginal proportion summarizes the observed population mixture, while a conditional proportion summarizes a comparison within specified strata. These summaries become numerically divergent when the strata have different baseline outcome rates and are represented in different proportions across the groups being compared.
No universal rule gives precedence to either the aggregated or stratified result. Descriptive analysis can legitimately report the aggregate distribution when population composition is part of the quantity under examination. Causal analysis instead depends on the relationships among exposure, outcome, and grouping variables, together with the target effect represented by the model.
Simpson's paradox is therefore a consequence of the non-collapsibility of certain statistical relationships and of variation in weighting across subpopulations. Its apparent contradiction disappears once the conditioning structure, population weights, and causal interpretation are stated as distinct components of the analysis.