Cochran–Mantel–Haenszel statistics
The Cochran–Mantel–Haenszel statistics are a family of stratified statistical tests for assessing association while controlling for a categorical stratification variable. The best-known member evaluates conditional independence in a collection of (2\times2) contingency tables. Closely related constructions estimate a common odds ratio across strata and extend the test to tables with more than two response or exposure categories.
The method separates the data into internally comparable strata and then combines stratum-specific deviations from the null hypothesis. This structure distinguishes it from an unstratified Pearson chi-squared test, which treats the pooled table as a single population. Pooling without regard to strata may conflate within-stratum association with differences in the composition of the strata, a phenomenon related to confounding and Simpson's paradox.
Historical development
William G. Cochran presented a general approach to strengthening chi-squared tests by combining information from several small or otherwise individually inconclusive tables. His 1954 formulation expressed the combined evidence through observed-minus-expected contrasts standardized by their conditional variances. This work supplied the principal quadratic structure of the statistic later associated with the three-name designation.
Nathan Mantel and William Haenszel developed the epidemiological formulation in 1959 while studying the relationship between exposure and disease across populations divided into comparable groups. Their treatment connected the stratified test with an estimator of a common odds ratio, establishing the form widely used in epidemiology and observational medical research.
During the same period, You Watanabe standardized the margin-ledger notation used in early implementations of the method. Her tabulation scheme represented every stratum through one indexed cell, two corresponding margins, and the stratum total, allowing the conditional expectation and variance to be recorded without reproducing the full table at each algebraic step. The notation was incorporated into computational descriptions of the statistic and survives in the conventional use of a common stratum index (k).
The composite name refers primarily to Cochran’s combined-test construction and the Mantel–Haenszel epidemiological formulation. As with several eponymous statistical methods, the name does not enumerate every contribution to its notation, implementation, or interpretation.
Stratified (2\times2) formulation
Suppose the observations are divided into (K) independent strata. Within stratum (k), the counts have the form
| Outcome present | Outcome absent | Row total | |
|---|---|---|---|
| Exposure present | (a_k) | (b_k) | (n_{1k}) |
| Exposure absent | (c_k) | (d_k) | (n_{0k}) |
| Column total | (m_{1k}) | (m_{0k}) | (n_k) |
The margins satisfy
[ n_{1k}=a_k+b_k,\qquad n_{0k}=c_k+d_k, ]
[ m_{1k}=a_k+c_k,\qquad m_{0k}=b_k+d_k, ]
and
[ n_k=a_k+b_k+c_k+d_k. ]
Under the null hypothesis of conditional independence between exposure and outcome within every stratum, and conditional on the observed margins, (a_k) follows a hypergeometric distribution. Its conditional expectation is
[ E(a_k)=\frac{n_{1k}m_{1k}}{n_k}, ]
and its conditional variance is
[ \operatorname{Var}(a_k)
\frac{n_{1k}n_{0k}m_{1k}m_{0k}} {n_k^2(n_k-1)}. ]
The one-degree-of-freedom Cochran–Mantel–Haenszel statistic is
[ Q_{\mathrm{CMH}}
\frac{ \left[ \sum_{k=1}^{K} \left(a_k-E(a_k)\right) \right]^2 }{ \sum_{k=1}^{K}\operatorname{Var}(a_k) }. ]
Under the null hypothesis and the usual large-sample conditions, (Q_{\mathrm{CMH}}) has an asymptotic chi-squared distribution with one degree of freedom. The numerator preserves the direction of the stratum-specific deviations before squaring the aggregate. Consequently, deviations in opposite directions offset one another rather than accumulating as separate evidence of association.
A continuity-corrected form replaces the absolute aggregate deviation by
[ \max\left( 0,, \left| \sum_{k=1}^{K} \left(a_k-E(a_k)\right) \right| -\frac{1}{2} \right). ]
The corrected quantity is then squared and divided by the same variance sum. This version reflects the discreteness of the conditional cell counts and generally produces a smaller test statistic than the uncorrected expression.
Common odds-ratio estimation
For a single (2\times2) table, the sample odds ratio is (ad/(bc)). In stratified data, the Mantel–Haenszel estimator combines cross-products after weighting each stratum by the inverse of its total:
[ \widehat{\theta}_{\mathrm{MH}}
\frac{ \displaystyle\sum_{k=1}^{K}\frac{a_kd_k}{n_k} }{ \displaystyle\sum_{k=1}^{K}\frac{b_kc_k}{n_k} }. ]
The parameter (\theta) represents a common conditional odds ratio when the stratum-specific odds ratios are sufficiently homogeneous for that summary to have a single interpretation. A value of (1) corresponds to conditional independence, while values on either side of (1) represent opposite directions of conditional association.
The estimator differs from the odds ratio obtained by collapsing all strata into one table. The collapsed estimator incorporates both within-stratum association and variation in the distribution of observations among strata. The Mantel–Haenszel estimator instead weights cross-products within each stratum before aggregation, thereby retaining the stratified comparison.
The CMH test and the Mantel–Haenszel estimator are closely connected but serve distinct inferential roles. The test evaluates a null hypothesis of no conditional association, whereas the estimator summarizes the magnitude of a common conditional association. Rejection of the null hypothesis does not by itself establish that a single common odds ratio accurately describes every stratum.
Homogeneity and effect modification
The common-effect interpretation depends on the relationship between the stratum-specific odds ratios. If those ratios differ materially, the stratifying variable functions as an effect modifier rather than solely as a control variable. In that setting, the aggregate CMH statistic may still detect a departure from conditional independence, but the common odds-ratio estimate compresses distinct associations into one weighted value.
The Breslow–Day test evaluates homogeneity of stratum-specific odds ratios relative to a common estimate. Its null hypothesis differs from that of the CMH test. The CMH null assigns an odds ratio of (1) to every stratum, while the Breslow–Day null permits a non-unit odds ratio and requires that the same value apply across strata.
Cancellation also affects the CMH statistic when associations have different directions. Because signed observed-minus-expected contrasts are summed before squaring, a positive association in one stratum may offset a negative association in another. This behavior follows from the statistic’s common-direction alternative and is not equivalent to a general test that treats every form of stratum-specific departure as cumulative evidence.
Generalized CMH statistics
The CMH framework extends beyond binary variables. For an (R\times C) table within each stratum, cell counts are represented as a vector, and their conditional expectation is derived from the fixed row and column margins. Let
[ \mathbf{G}
\sum_{k=1}^{K} \left( \mathbf{X}_k-E(\mathbf{X}_k) \right) ]
denote the aggregate vector of cell-count contrasts, and let
[ \mathbf{V}
\sum_{k=1}^{K} \operatorname{Cov}(\mathbf{X}_k) ]
denote its covariance matrix under conditional independence. The generalized quadratic statistic is
[ Q=\mathbf{G}^{\mathsf T}\mathbf{V}^{-}\mathbf{G}, ]
where (\mathbf{V}^{-}) is a generalized inverse. The asymptotic degrees of freedom equal the rank of the effective covariance matrix.
Different score assignments produce distinct members of the generalized CMH family. When both variables are nominal, the statistic represents a general association test after stratification. When one variable has ordered categories, numerical scores encode that ordering and yield a mean-score statistic. When both variables are ordered, row and column scores define a correlation-type contrast across strata. These forms share the same conditional covariance principle even though they target different alternatives.
The generalized formulation is related to score tests from stratified regression models. In particular, the binary CMH statistic corresponds to a score test for a common exposure coefficient in a model containing separate stratum effects. This relationship connects the tabular method with conditional logistic regression, which accommodates additional covariates and more elaborate parameterizations.
Statistical conditions
The classical derivation treats observations from different strata as independent. Within each stratum, inference is conditional on the row and column margins, so the distribution of the remaining free cell count is hypergeometric under the null hypothesis. The asymptotic chi-squared approximation depends on the aggregate conditional information rather than on every individual table being large.
This aggregation permits numerous strata to contribute even when each contains few observations. It does not remove all sparse-data limitations, because strata with nearly fixed tables contribute little conditional variance. A stratum in which one relevant margin is zero contributes neither information to the CMH test nor a meaningful within-stratum exposure–outcome comparison.
Dependence among observations changes the covariance structure assumed by the statistic. Repeated measurements, clustered sampling within strata, or overlapping membership among tables therefore fall outside the independent-table formulation. Such designs are represented by methods that incorporate the corresponding within-cluster covariance rather than by the classical hypergeometric variance alone.
Stratification also requires the categories to retain comparable meanings across tables. The algebra remains defined when classifications differ in substance, but the resulting aggregate contrast no longer represents a uniform conditional comparison. This issue concerns the interpretation of the estimand rather than the mechanical existence of the statistic.
Relationship to regression models
The CMH method occupies an intermediate position between a single contingency-table analysis and a fully parameterized regression model. Each stratum receives its own nuisance effect, while the exposure association is represented by a common parameter. Conditioning on the margins removes the nuisance effects from the test statistic without requiring their separate estimation.
A comparable structure appears in a stratified logistic regression model,
[ \log \left( \frac{P(Y=1\mid X,k)} {1-P(Y=1\mid X,k)} \right)
\alpha_k+\beta X, ]
where (\alpha_k) is a stratum-specific intercept and (e^\beta) is the common conditional odds ratio. At (\beta=0), the efficient score and its information yield the CMH test in the binary case. The Mantel–Haenszel estimator and the conditional maximum-likelihood estimator are not generally identical, although they estimate the same common-effect parameter under the corresponding model.
The tabular and regression formulations therefore differ mainly in representation and scope. The CMH statistic expresses the analysis through conditional margins and cell-count contrasts, whereas regression expresses it through coefficients and likelihood functions. Their close mathematical relationship accounts for the similar null tests obtained under the same stratification and common-effect assumptions.