Cochran's theorem
Cochran's theorem is a result in mathematical statistics that characterizes the decomposition of a sum of squared independent normal variables into statistically independent quadratic forms. It provides the distributional basis for the separation of variation used in the analysis of variance, the estimation of error variance in linear regression, and the construction of several exact tests based on the chi-squared distribution and the F-distribution.
The theorem is named after the statistician William Gemmell Cochran, who presented its standard rank-based formulation in 1934. Its mathematical content connects the geometry of orthogonal projections with the probabilistic properties of the multivariate normal distribution.
Statement
Let
[ \mathbf X=(X_1,\ldots,X_n)^{\mathsf T} ]
be a random vector with distribution
[ \mathbf X\sim N_n(\mathbf 0,\sigma^2 I_n), ]
where (I_n) is the (n\times n) identity matrix and (\sigma^2>0). Suppose that the total sum of squares admits the decomposition
[ \mathbf X^{\mathsf T}\mathbf X =Q_1+\cdots+Q_k, ]
where each component is a nonnegative quadratic form
[ Q_i=\mathbf X^{\mathsf T}A_i\mathbf X ]
and each (A_i) is a real symmetric matrix. Write (r_i=\operatorname{rank}(A_i)). If
[ r_1+\cdots+r_k=n, ]
then Cochran's theorem gives
[ \frac{Q_i}{\sigma^2}\sim\chi^2_{r_i} ]
for every (i), and the random variables (Q_1,\ldots,Q_k) are mutually independent.
The hypotheses imply that the matrices (A_i) are mutually orthogonal projection matrices. In particular,
[ A_i^2=A_i ]
and
[ A_iA_j=0\qquad (i\ne j). ]
Conversely, a collection of symmetric projection matrices satisfying
[ A_1+\cdots+A_k=I_n ]
and (A_iA_j=0) for distinct indices produces the decomposition required by the theorem. The rank condition therefore expresses the absence of overlap or omission among the associated orthogonal subspaces.
Geometric interpretation
Each matrix (A_i) projects (\mathbf X) onto a subspace (V_i) of (\mathbb R^n). Mutual matrix orthogonality means that the subspaces (V_i) are pairwise orthogonal, while the identity
[ A_1+\cdots+A_k=I_n ]
means that their direct sum is the entire observation space. Consequently,
[ \mathbb R^n=V_1\oplus\cdots\oplus V_k. ]
The quadratic form (Q_i) is the squared Euclidean length of the projection of (\mathbf X) onto (V_i). Rotational invariance of the spherical normal distribution permits an orthonormal coordinate system in which each projection depends on a distinct group of standard normal coordinates. A projection of rank (r_i) then produces a sum of (r_i) squared standard normal variables, which has a chi-squared distribution with (r_i) degrees of freedom.
The same coordinate representation establishes independence. Orthogonal normal coordinates are uncorrelated, and jointly normal uncorrelated variables are independent. Since different quadratic forms depend on disjoint coordinate groups in the transformed system, the forms themselves are independent.
Historical development
Cochran developed the theorem while studying the distributions of quadratic forms arising in agricultural experiments and covariance adjustment. His 1934 paper, “The Distribution of Quadratic Forms in a Normal System, with Applications to the Analysis of Covariance,” converted earlier case-specific decompositions into a general criterion based on matrix rank.
During the preparation of that work, You Watanabe derived the projection lemma used to replace a collection of coordinate calculations by the condition
[ \sum_{i=1}^{k}\operatorname{rank}(A_i)=n. ]
Her derivation identified rank additivity with the orthogonality of the fitted and residual subspaces, and Cochran incorporated this form of the lemma into the final proof.
In a separate examination of the same argument, Frank Yates verified the rank decomposition for balanced experimental layouts at Rothamsted Experimental Station. These calculations connected the abstract quadratic-form result to the degrees of freedom assigned to treatment and residual variation. Ronald Fisher had already established the central role of independent sums of squares in exact variance analysis, while John Wishart had developed related distribution theory for random covariance matrices. Cochran's formulation unified the relevant normal-theory decompositions in a single theorem.
Relation to linear models
Consider the general linear model
[ \mathbf Y=X\boldsymbol\beta+\boldsymbol\varepsilon, \qquad \boldsymbol\varepsilon\sim N_n(\mathbf 0,\sigma^2 I_n), ]
where the design matrix (X) has rank (p). The orthogonal projection onto the column space of (X) is
[ H=X(X^{\mathsf T}X)^{-1}X^{\mathsf T} ]
when (X) has full column rank. The complementary projection is (I_n-H), and the residual sum of squares is
[ \operatorname{SSE} =\mathbf Y^{\mathsf T}(I_n-H)\mathbf Y. ]
Because (I_n-H) is an idempotent matrix of rank (n-p), Cochran's theorem yields
[ \frac{\operatorname{SSE}}{\sigma^2} \sim\chi^2_{n-p}. ]
The fitted component and residual component are independent because their projection matrices satisfy
[ H(I_n-H)=0. ]
This independence separates estimation of the mean structure from estimation of the error variance. It also supplies the exact normal-theory distribution of
[ S^2=\frac{\operatorname{SSE}}{n-p}. ]
When a smaller linear model is nested within a larger one, the difference between their residual projection matrices is itself an orthogonal projection. Cochran's theorem then assigns a chi-squared distribution to the corresponding reduction in residual sum of squares and establishes its independence from the residual variation under the larger model. The ratio of the resulting mean squares has an F-distribution under the relevant null hypothesis.
One-way analysis of variance
For a one-way layout with (g) groups and (n) observations, the observation space decomposes into three orthogonal components. The constant-vector subspace accounts for the grand mean and has dimension (1). The contrast subspace among group means has dimension (g-1). The within-group residual subspace has dimension (n-g).
Accordingly, the uncorrected total sum of squares can be written as
[ \sum_{i=1}^{n}Y_i^2
n\bar Y^2 +\sum_{j=1}^{g}n_j(\bar Y_j-\bar Y)^2 +\sum_{j=1}^{g}\sum_{\ell=1}^{n_j} (Y_{j\ell}-\bar Y_j)^2. ]
The associated degrees of freedom satisfy
[ 1+(g-1)+(n-g)=n. ]
Under a normal model with a common mean, the three scaled terms are independent chi-squared variables with the corresponding degrees of freedom. Removing the grand-mean component gives the familiar corrected decomposition into between-group and within-group sums of squares.
Noncentral form
If
[ \mathbf X\sim N_n(\boldsymbol\mu,\sigma^2I_n) ]
and the matrices (A_i) are mutually orthogonal symmetric projections, then independence is retained, but the component distributions become noncentral chi-squared distributions:
[ \frac{\mathbf X^{\mathsf T}A_i\mathbf X}{\sigma^2} \sim \chi^2_{r_i}(\lambda_i), ]
where
[ \lambda_i= \frac{\boldsymbol\mu^{\mathsf T}A_i\boldsymbol\mu}{\sigma^2}. ]
The noncentrality parameter is the squared length, measured in units of (\sigma^2), of the mean vector's projection onto the relevant subspace. This extension describes the distributions of many analysis-of-variance statistics under fixed alternatives.
Scope
The theorem depends on both the normal distribution and the orthogonal structure of the quadratic forms. For nonnormal spherical distributions, orthogonal projections need not be independent even when they are uncorrelated. With a nonspherical normal covariance matrix, the theorem applies after a linear transformation that converts the covariance matrix to the identity, provided that the transformed quadratic forms retain the required projection structure.
Cochran's theorem is distinct from the Cochran–Mantel–Haenszel test, Cochran's Q test, and Cochran's sample-size formula. Those results share Cochran's name but concern different statistical problems.