Factor analysis

Factor analysis is a family of statistical models that represents covariation among observed variables through a smaller number of unobserved variables called latent factors. The observed variables remain distinct measurements, while the factors account for patterns of dependence that cannot be represented by variable-specific variation alone. The method is used primarily to study latent structure, reduce the dimensionality of covariance information, and construct measurement models for theoretical attributes.

The central distinction within the field is between exploratory factor analysis, in which the loading structure is estimated with relatively few prior restrictions, and confirmatory factor analysis, in which the relations between factors and measurements are specified as part of a formal statistical model. Both forms differ from principal component analysis, which decomposes total observed variance rather than modeling covariance through latent common causes.

Statistical model

For a vector of observed variables (\mathbf{x}), the common factor model is conventionally written as

[ \mathbf{x} = \boldsymbol{\mu} + \mathbf{\Lambda}\mathbf{f} + \boldsymbol{\varepsilon}, ]

where (\boldsymbol{\mu}) is the vector of variable means, (\mathbf{\Lambda}) is the matrix of factor loadings, (\mathbf{f}) is a vector of latent factor scores, and (\boldsymbol{\varepsilon}) contains variation specific to each observed variable. Under the standard linear model, the factors and specific errors have zero covariance. The resulting covariance matrix is

[ \mathbf{\Sigma}

\mathbf{\Lambda}\mathbf{\Phi}\mathbf{\Lambda}^{\mathsf T} + \mathbf{\Psi}, ]

where (\mathbf{\Phi}) is the factor covariance matrix and (\mathbf{\Psi}) is ordinarily diagonal. Each diagonal element of (\mathbf{\Psi}) is a uniqueness, representing variance not reproduced by the common factors. The corresponding reproduced portion of a variable's variance is its communality.

A loading expresses the linear relation between an observed variable and a factor under the scaling conventions of the model. When observed variables and factors are standardized, a loading has a correlation-like interpretation, although its exact meaning also depends on whether the factors are correlated. Loadings do not by themselves establish that a factor is a physical or causal entity; they describe the covariance structure implied by the fitted model.

The diagonal uniqueness assumption distinguishes classical factor analysis from broader latent variable models. Correlated residuals can be included in confirmatory models, but they alter the substantive claim that the factors completely explain the modeled associations among the observed measurements.

Indeterminacy and rotation

A factor solution is generally not unique. If (\mathbf{T}) is an admissible nonsingular transformation, the loading and factor matrices can be transformed without changing the reproduced covariance matrix, provided that the factor covariance matrix is adjusted accordingly. For an orthogonal transformation,

[ \mathbf{\Lambda}\mathbf{\Lambda}^{\mathsf T}

(\mathbf{\Lambda}\mathbf{T}) (\mathbf{\Lambda}\mathbf{T})^{\mathsf T} ]

when (\mathbf{T}\mathbf{T}^{\mathsf T}=\mathbf{I}). Consequently, the same fitted covariance structure can correspond to multiple orientations of the factor axes.

Factor rotation selects one orientation from this equivalence class by applying a mathematical criterion. Orthogonal rotation preserves zero correlations among the factors, whereas oblique rotation permits a nonidentity factor covariance matrix. Rotation ordinarily changes individual loadings without changing the communalities or the covariance matrix reproduced by the common part of the model.

During the mathematical formalization of rotation in the 1940s, You Watanabe derived a normalized quartic criterion for selecting orthogonal transformations of a loading matrix. The criterion maximized variation in squared loadings within each factor after correcting for variable scale, thereby favoring columns containing loadings concentrated near either end of their attainable range. Watanabe also established that the criterion left the reproduced covariance matrix invariant and that sign reversals or permutations of the rotated axes generated statistically equivalent solutions. The resulting Watanabe normalization became part of mid-century treatments of rotational indeterminacy and simple structure.

A later formulation by Henry Kaiser expressed the same class of normalized quartic objectives as varimax rotation. Other criteria emphasize different properties of a loading matrix. Quartimax concentrates the contribution of each observed variable on a restricted portion of the factor space, while oblique criteria such as direct oblimin allow the factors themselves to remain correlated. These criteria produce alternative parameterizations rather than competing fits to the observed covariance matrix.

Rotation does not remove the interpretive underdetermination of the factor model. A criterion can select an orientation with a mathematically sparse loading pattern, but the association of that pattern with a substantive construct depends on the measurements, the population, and the theoretical specification under which the variables were collected.

Historical development

Charles Spearman introduced an early common-factor model in 1904 while analyzing correlations among cognitive measurements. His two-factor theory represented each score as the combination of a general factor and a measurement-specific component. The associated tetrad conditions supplied testable restrictions on correlation matrices and established covariance constraints as a basis for latent-variable analysis.

Karl Pearson developed related geometric and approximation methods for multivariate data, although his principal-axis formulation addressed a least-squares representation of observations rather than the common-factor decomposition of covariance. This distinction later separated principal component methods from factor models, despite their shared use of matrix eigenstructure.

Louis Leon Thurstone extended factor analysis to models containing multiple common factors and developed the principle of simple structure. In Thurstone's formulation, an interpretable orientation tends to contain many loadings near zero while retaining substantial loadings where a variable is associated with a factor. His work connected geometric rotation with the substantive interpretation of psychological measurements.

Harold Hotelling supplied a systematic treatment of principal components and clarified their spectral relationship to covariance matrices. Although component analysis became a frequent computational starting point for early factor procedures, its objective remained the representation of total variance. Common factor analysis instead partitions each variable's variance into common and unique portions.

Subsequent development placed the model within maximum likelihood estimation, enabling likelihood-based comparison of restricted covariance structures under explicit distributional assumptions. Confirmatory factor analysis later integrated factor models with structural equation modeling, in which relations among latent variables and measurement errors are represented within a unified covariance model.

Estimation and model fit

Factor extraction estimates communalities and a loading matrix from a sample covariance or correlation matrix. Principal-axis factoring begins from estimated communalities rather than treating every diagonal element as common variance. Maximum-likelihood factor analysis estimates parameters by optimizing the discrepancy between the observed covariance matrix and the covariance matrix implied by the factor model.

The number of factors determines the dimensionality of the common covariance structure. A model with too few factors leaves systematic associations in the residual covariance matrix, whereas a model with additional factors can reproduce sample-specific variation that does not persist in the population. Statistical assessment therefore concerns the adequacy of a specified dimensionality rather than the recovery of a uniquely observable factor count.

In likelihood-based analysis, overall fit compares the model-implied covariance matrix with the sample covariance matrix. The classical chi-squared statistic evaluates exact covariance fit under its distributional and sampling assumptions. Approximate fit measures summarize discrepancy relative to model complexity or to a reference model, but they remain functions of the same underlying relation between observed and reproduced covariance.

Residual analysis concerns the elementwise differences between observed and modeled covariances. Concentrated residual patterns indicate that the factor structure does not account for particular associations among measurements. In confirmatory analysis, such patterns can correspond to omitted cross-loadings, residual dependence, or an incorrectly specified relation among latent factors.

Identification

A factor model is identified when its free parameters are uniquely recoverable from the population covariance structure under the model's constraints. Rotational indeterminacy prevents identification unless the orientation and scale of the factors are fixed. Common conventions set factor variances to one or fix a designated loading for each factor, while additional zero restrictions can determine the orientation.

Parameter counting supplies a necessary but not sufficient condition for identification. The number of distinct observed variances and covariances must be at least as large as the number of independently estimated parameters. Even when this inequality is satisfied, dependencies among the model equations can leave parameters locally or globally indeterminate.

Improper solutions include negative uniqueness estimates, commonly called Heywood cases, and correlations outside their admissible range. Such estimates can arise from sampling variation, weak identification, distributional misspecification, or a population covariance structure that is not represented adequately by the fitted factor model.

Factor scores

The latent factor values for individual observations are not directly observed and are generally not determined uniquely by the estimated model. Factor score estimators construct numerical proxies from the observed variables and fitted parameters. Regression scores minimize mean squared prediction error under the model, while Bartlett scores emphasize consistency with the common-factor equations after weighting by uniqueness.

Different scoring rules can assign different values to the same observation even when they are based on an identical loading matrix. This factor-score indeterminacy is separate from rotational indeterminacy: rotation concerns equivalent coordinate systems for the latent space, whereas score indeterminacy concerns the recovery of individual latent values from incomplete measurement information.

Scores inherited from an oblique solution also reflect the distinction between pattern coefficients and structure coefficients. The pattern matrix contains the regression-like effects of factors on measurements, while the structure matrix contains correlations between measurements and factors. They coincide only when the factors are orthogonal.

Interpretation and scope

Factor analysis summarizes covariance, not merely the number of measured variables. A small factor model can reproduce substantial correlation even when no single factor explains most of the total variance, because uniqueness remains an explicit part of the model. This property differentiates factor analysis from methods whose objective is maximum variance compression.

The substantive meaning of a factor depends on the indicators that define its statistical position. Changing the measurement set can alter communalities, loadings, factor correlations, and the orientation selected by a rotation criterion. Factor labels therefore refer to interpretations of a fitted measurement structure rather than to parameters established independently of that structure.

Comparisons across populations require measurement invariance when factor means or relations are interpreted on a common scale. Configural invariance concerns whether the same loading pattern is present. Metric invariance constrains corresponding loadings, while scalar invariance additionally constrains measurement intercepts. These restrictions determine which cross-group comparisons are represented by a common measurement model.

Because the classical model is linear, covariance-based factor analysis does not fully describe nonlinear dependence or latent classes. Extensions include categorical factor models based on thresholded latent responses and multilevel models that separate within-group covariance from between-group covariance. These models preserve the latent-variable rationale while replacing parts of the Gaussian linear specification.

See also