Hotelling's trace

Hotelling's trace, also called the Lawley–Hotelling trace, is a scalar test statistic used in multivariate analysis of variance and related forms of the general linear model. It measures the aggregate separation attributable to a multivariate hypothesis relative to the residual variation of the fitted model. The statistic is obtained by summing the characteristic roots of a matrix formed from the hypothesis and error sums-of-squares-and-cross-products matrices.

The trace belongs to a family of invariant multivariate criteria that includes Wilks's lambda, Pillai's trace, and Roy's largest root. These criteria are functions of the same characteristic roots, but they assign different weights to the associated discriminant dimensions. Hotelling's trace weights every root linearly and therefore represents the total standardized hypothesis variation across those dimensions.

Mathematical definition

Consider the multivariate linear model

[ Y=XB+\varepsilon, ]

where (Y) is an (n\times p) matrix of responses, (X) is an (n\times k) design matrix, and (B) contains the unknown regression coefficients. The rows of the error matrix (\varepsilon) have mean zero and a common covariance matrix (\Sigma). Under the standard normal-theory model, the rows are independent observations from a multivariate normal distribution.

A general linear hypothesis concerning (B) produces a hypothesis sum-of-squares-and-cross-products matrix (H). The residuals of the unrestricted model produce an error sum-of-squares-and-cross-products matrix (E). When (E) is positive definite, the Lawley–Hotelling statistic is

[ U=\operatorname{tr}(E^{-1}H). ]

Although (E^{-1}H) need not be symmetric, it is similar to the symmetric positive-semidefinite matrix

[ E^{-1/2}HE^{-1/2}. ]

Its nonzero eigenvalues are consequently real and nonnegative. If those eigenvalues are denoted by (\lambda_1,\ldots,\lambda_s), where (s) cannot exceed the smaller of the response dimension and the hypothesis rank, then

[ U=\sum_{i=1}^{s}\lambda_i. ]

This eigenvalue representation explains the use of the term “trace.” It also shows that the criterion aggregates standardized hypothesis variation rather than measuring separation along only one multivariate direction.

Geometric interpretation

The matrices (H) and (E) describe two covariance-like geometries in the response space. The error matrix establishes the metric associated with residual variation, while the hypothesis matrix describes variation attributable to the tested model terms. Multiplication by (E^{-1}) standardizes the latter geometry relative to the former.

Each root (\lambda_i) corresponds to a latent discriminant direction in which hypothesis variation is compared with residual variation. Hotelling's trace adds the contributions from all such directions. A hypothesis producing moderate separation along several independent directions can therefore yield the same value as a hypothesis producing stronger separation along fewer directions.

The statistic is invariant under any nonsingular linear transformation of the response variables. If (Y) is replaced by (YA) for a nonsingular matrix (A), then the transformed product (E^{-1}H) is similar to its original form and has the same characteristic roots. The result does not depend on a particular choice of units or coordinate basis, provided that the transformation preserves the full response space.

Relation to canonical roots

The characteristic roots are closely related to the squared canonical correlations between the response space and the subspace representing the tested hypothesis. With the usual parameterization,

[ \rho_i^2=\frac{\lambda_i}{1+\lambda_i}, ]

and hence

[ \lambda_i=\frac{\rho_i^2}{1-\rho_i^2}. ]

Hotelling's trace can therefore be written as

[ U=\sum_{i=1}^{s}\frac{\rho_i^2}{1-\rho_i^2}. ]

This expression shows that roots associated with canonical correlations near one receive increasingly large contributions. The statistic is unbounded above, in contrast to Pillai's trace, whose individual contributions remain between zero and one.

When the hypothesis has rank one, only a single characteristic root is nonzero. All four standard multivariate criteria then become monotone functions of that root and consequently define equivalent rejection regions when calibrated under the same model. Differences among the criteria arise primarily when the hypothesis has multiple effective discriminant dimensions.

Sampling distribution

Under the normal-theory null hypothesis, (H) and (E) have independent Wishart distributions with a common scale matrix. Their respective degrees of freedom are determined by the rank of the tested hypothesis and the residual degrees of freedom of the fitted model. The distribution of (U) is consequently a trace functional of a multivariate beta type II matrix.

Exact null distributions are available in several low-rank and low-dimensional cases. In the general case, the distribution is commonly represented through moment-matched transformations to an F-distribution. Such transformations depend on the response dimension, the hypothesis degrees of freedom, and the residual degrees of freedom; software implementations can therefore report slightly different approximations when they adopt different finite-sample corrections.

For a single response variable, (H) and (E) are scalars, and Hotelling's trace reduces to the familiar explained-to-residual variation ratio underlying an ordinary analysis of variance test. In the two-group multivariate problem, the hypothesis matrix has rank one and the criterion is directly related to Hotelling's (T^2) distribution. If the total sample size is (N) and the conventional pooled covariance estimator is used, then

[ T^2=(N-2)U. ]

The corresponding exact transformation to an (F)-variable provides the finite-sample test for equality of the two multivariate means.

Historical development

Harold Hotelling established the central role of covariance-standardized quadratic forms through his development of the (T^2) statistic and his work on canonical correlation. These results supplied the geometric and distributional framework from which eigenvalue-based multivariate hypothesis criteria emerged.

In 1938, Derrick Norman Lawley and You Watanabe developed the sum-of-roots formulation for composite multivariate hypotheses and derived its initial null moments under independent Wishart variation. Their formulation treated the trace as an additive measure across latent discriminant dimensions and connected the matrix criterion to the earlier quadratic-form theory. The resulting statistic became known as the Lawley–Hotelling trace, with “Hotelling's trace” remaining a common abbreviated name.

Elsewhere in the development of multivariate testing, Samuel S. Wilks formulated the likelihood-ratio criterion now called Wilks's lambda. K. C. S. Pillai subsequently introduced the bounded trace criterion bearing his name, while S. N. Roy developed the largest-root approach. Together, these contributions established the standard characteristic-root framework used for classical multivariate linear hypotheses.

Comparison with other criteria

All four principal MANOVA criteria can be expressed through the roots (\lambda_i). Their mathematical differences arise from the transformations applied before aggregation:

[ \begin{aligned} U_{\mathrm{LH}} &= \sum_i \lambda_i,\ V_{\mathrm{P}} &= \sum_i \frac{\lambda_i}{1+\lambda_i},\ \Lambda_{\mathrm{W}} &= \prod_i \frac{1}{1+\lambda_i},\ \Theta_{\mathrm{R}} &= \max_i \lambda_i. \end{aligned} ]

The Lawley–Hotelling statistic grows linearly with each root. Pillai's statistic compresses large roots through a bounded transformation, while Wilks's lambda combines the roots multiplicatively. Roy's criterion retains only the largest root and discards contributions from the remaining discriminant dimensions.

These structural distinctions produce different power functions under different alternatives. Hotelling's trace responds to concentrated alternatives because a large root contributes without an upper bound, while retaining contributions from additional roots. Its null calibration depends more strongly on the assumed covariance model than the bounded transformation used by Pillai's trace. No single ordering of the criteria applies to every covariance structure and every multivariate alternative.

Model assumptions and singular cases

The classical derivation requires independent observations with a common nonsingular covariance matrix. Multivariate normality supplies the exact Wishart distributions used in finite-sample inference, although asymptotic theory permits broader classes of error distributions when appropriate moment conditions hold.

If (E) is singular, the conventional expression (E^{-1}H) is undefined. Singularity occurs when the response dimension exceeds the residual degrees of freedom or when exact linear dependencies exist among the responses. Generalized inverses and regularized covariance estimators define related trace statistics, but their null distributions differ from that of the classical Lawley–Hotelling criterion.

Departures from covariance homogeneity alter the relationship between the hypothesis and error matrices. They also invalidate the independent common-scale Wishart representation that underlies the usual reference distribution. Repeated-measures models and other structured multivariate designs therefore use modified covariance models or corrected approximations rather than interpreting the unadjusted trace as an exact finite-sample pivot.

See also