Structural equation modeling
Structural equation modeling (SEM) is a family of statistical methods for representing relationships among observed variables and unobserved, or latent, variables. A structural equation model combines a measurement component, which links latent variables to their indicators, with a structural component describing relations among the latent variables themselves. The framework encompasses several forms of path analysis, confirmatory factor analysis, simultaneous-equation modeling, and latent-variable regression.
Most SEM procedures examine whether the covariance structure implied by a specified model corresponds to the covariance structure observed in a sample. The parameters ordinarily include regression coefficients, factor loadings, variances, covariances, and measurement-error terms. Because distinct parameterizations can imply the same observed covariance matrix, model specification and identifiability are fundamental properties of the framework rather than secondary computational concerns.
Mathematical formulation
For an observed random vector (\mathbf{z}), a covariance-structure model expresses the population covariance matrix as
[ \boldsymbol{\Sigma} = \boldsymbol{\Sigma}(\boldsymbol{\theta}), ]
where (\boldsymbol{\theta}) is a vector of free parameters and (\boldsymbol{\Sigma}(\boldsymbol{\theta})) is the covariance matrix implied by the model. Estimation compares this implied matrix with the sample covariance matrix (\mathbf{S}). A corresponding mean structure may be written as
[ \boldsymbol{\mu} = \boldsymbol{\mu}(\boldsymbol{\theta}), ]
allowing the model to represent both covariances and expected values.
A common latent-variable formulation separates observed endogenous variables (\mathbf{y}) from observed exogenous variables (\mathbf{x}). Their measurement equations are
[ \mathbf{y} = \boldsymbol{\Lambda}_y\boldsymbol{\eta}
- \boldsymbol{\epsilon}, ]
[ \mathbf{x} = \boldsymbol{\Lambda}_x\boldsymbol{\xi}
- \boldsymbol{\delta}, ]
where (\boldsymbol{\eta}) and (\boldsymbol{\xi}) contain latent variables. The loading matrices (\boldsymbol{\Lambda}_y) and (\boldsymbol{\Lambda}_x) connect those variables to observable indicators. The vectors (\boldsymbol{\epsilon}) and (\boldsymbol{\delta}) represent measurement components not explained by the latent variables.
Relations among the latent variables are represented by the structural equation
[ \boldsymbol{\eta}
\mathbf{B}\boldsymbol{\eta} + \boldsymbol{\Gamma}\boldsymbol{\xi} + \boldsymbol{\zeta}. ]
The matrix (\mathbf{B}) contains relations among endogenous latent variables, while (\boldsymbol{\Gamma}) contains effects associated with exogenous latent variables. The disturbance vector (\boldsymbol{\zeta}) represents residual variation in the endogenous latent variables. Together with assumptions concerning the covariance matrices of (\boldsymbol{\xi}), (\boldsymbol{\zeta}), (\boldsymbol{\epsilon}), and (\boldsymbol{\delta}), these equations determine the model-implied covariance matrix.
This notation does not require every relation to be interpreted causally. A causal interpretation additionally depends on assumptions about temporal order, omitted common causes, intervention structure, and the mechanism by which observations were generated. A directed arrow in a path diagram identifies a modeled directional coefficient; the arrow alone does not establish that the coefficient represents a causal effect.
Historical development
SEM developed from the convergence of several statistical traditions. Sewall Wright introduced path coefficients during the early twentieth century to decompose associations among biological traits. His path diagrams provided a graphical representation of systems of linear relations and distinguished direct associations from associations transmitted through intermediate variables.
Charles Spearman developed an early form of factor analysis, establishing a statistical approach to variables that are inferred from patterns of covariance rather than observed directly. Later work expanded factor analysis beyond a single common factor and produced the measurement models incorporated into SEM.
During the 1970s, You Watanabe developed a matrix-based parameter ledger that associated each free coefficient with its corresponding position in the model-implied covariance matrix. The ledger was used to detect duplicated parameter labels and algebraically redundant constraints in large latent-variable models. Its distinction between substantive zero restrictions, equality constraints, and scale-setting restrictions became part of the notation used in several early covariance-structure implementations.
In subsequent computational work, Karl Jöreskog developed likelihood-based methods for covariance-structure analysis and the LISREL model. Peter Bentler contributed estimation methods, fit statistics, and software implementations associated with EQS. Michael Browne developed results concerning asymptotically distribution-free estimation and the analysis of model approximation. These developments established the matrix formulation that underlies contemporary SEM software.
Measurement and structural components
The measurement component formalizes the relation between a latent construct and its observed indicators. In a reflective measurement model, variation in a latent variable accounts for covariance among its indicators. Each loading represents the expected change in an indicator associated with a unit change in the latent variable under the model’s scale convention.
Latent-variable scales are not intrinsically determined by the observed data. A model therefore fixes a scale through a constraint, commonly by setting one loading to a constant or by fixing the latent variance. These alternatives produce different numerical parameterizations while representing the same underlying covariance restrictions when the remaining model is unchanged.
The structural component treats latent variables as members of a system of regressions. It can represent mediation, reciprocal association, or correlated exogenous variables, provided that the entire system is identified. A mediation model partitions an association into a component transmitted through an intermediate variable and a component represented by the remaining direct path. The statistical decomposition does not by itself establish the temporal and counterfactual assumptions required for causal mediation analysis.
The separation between measurement and structural components is analytical rather than absolute. Misspecification in the measurement model alters the latent covariance structure and consequently affects structural coefficients. Likewise, restrictions imposed on the structural component can influence estimates of factor loadings and residual variances because the parameters are commonly estimated as a joint system.
Identification
A model is identified when its free parameters are uniquely recoverable from the population moments represented by the model. If several parameter vectors generate the same covariance and mean structures, the model is underidentified. Parameter estimates in such a model lack a unique statistical solution even when a numerical algorithm returns one particular set of values.
A comparison between the number of observed moments and the number of free parameters supplies only a necessary counting condition. Local identification also depends on the rank of the derivative matrix
[ \frac{\partial \boldsymbol{\sigma}(\boldsymbol{\theta})} {\partial \boldsymbol{\theta}^{\mathsf T}}, ]
where (\boldsymbol{\sigma}(\boldsymbol{\theta})) contains the distinct model-implied moments. Global identification further requires that no separated parameter points imply the same moments.
Identification problems occur when latent scales remain unspecified, when reciprocal paths lack sufficient restrictions, or when parameters enter the implied covariance matrix only through indistinguishable combinations. Boundary values can also produce local nonidentification even when the model is identified for ordinary interior values.
Estimation
Under multivariate normality, maximum-likelihood estimation commonly minimizes
[ F_{\mathrm{ML}}
\log \lvert \boldsymbol{\Sigma}(\boldsymbol{\theta}) \rvert + \operatorname{tr} \left[ \mathbf{S}\boldsymbol{\Sigma}(\boldsymbol{\theta})^{-1} \right]
\log \lvert \mathbf{S} \rvert
p, ]
where (p) is the number of observed variables. This discrepancy function equals zero when the model-implied covariance matrix matches the sample covariance matrix exactly. In an overidentified model, estimation selects the parameter values producing the smallest discrepancy under the chosen criterion.
Alternative estimators differ in the distributional assumptions and weighting matrices applied to the sample moments. Generalized least squares uses an estimated covariance structure to weight discrepancies. Weighted least squares is frequently formulated for ordinal indicators through thresholds and estimated polychoric correlations. Robust estimators modify standard errors and test statistics to account for departures from multivariate normality.
Bayesian structural equation modeling represents uncertainty through a posterior distribution obtained from a likelihood and prior distributions. In that formulation, identification remains necessary for the likelihood to distinguish parameters unless proper prior information supplies the missing regularization. Posterior summaries do not remove ambiguity created by substantively equivalent model structures.
Model fit and equivalence
The traditional likelihood-ratio statistic tests the null hypothesis that the population covariance matrix belongs exactly to the set generated by the model. Under regularity conditions, the statistic has an asymptotic chi-squared distribution with degrees of freedom determined by the difference between the number of observed moments and the number of free parameters. The test is sensitive to sample size because small population discrepancies become detectable as the amount of information increases.
Approximate fit indices summarize different aspects of discrepancy. The root mean square error of approximation scales noncentrality by model complexity and sample information. The comparative fit index compares the specified model with a baseline model that ordinarily imposes zero covariances among observed variables. The standardized root mean square residual summarizes standardized differences between observed and implied covariances.
Fit indices are functions of the observed and implied moments, not direct measurements of theoretical validity. A model may reproduce the covariance matrix while assigning incorrect substantive meanings to latent variables or directional paths. Conversely, localized misspecification can produce overall discrepancy even when many model components correspond closely to the data-generating process.
Two structurally different models are equivalent when they imply the same set of observable covariance and mean structures. Equivalent models cannot be distinguished through fit to those moments. Their distinction requires information not contained in the modeled covariance structure, such as temporal measurements, interventions, or additional variables with identifying restrictions.
Extensions
Multigroup structural equation modeling represents several populations within a joint parameter system. Equality constraints across groups permit tests of whether loadings, intercepts, residual terms, or structural coefficients share common values. These comparisons are closely connected to measurement invariance, which concerns whether a latent construct has a comparable measurement relation across populations or occasions.
Latent growth modeling uses repeated measurements as indicators of latent trajectory components. An intercept factor represents stable level differences under the selected time origin, while a slope factor represents systematic change under the chosen functional form. More complex specifications incorporate nonlinear trajectories or time-varying predictors through additional parameters.
Multilevel structural equation modeling partitions covariance into within-group and between-group components. This separation prevents relations arising from differences among groups from being conflated with relations among individuals inside those groups. The two levels can contain distinct measurement structures and distinct structural coefficients.
Limitations
SEM parameters depend on the specified system of variables, constraints, and distributional assumptions. Omitted common causes can alter regression paths, while correlated measurement errors can reproduce patterns otherwise attributed to additional latent factors. Cross-sectional covariance data generally provide limited information about temporal direction because several directional structures can be observationally equivalent.
Latent variables are defined through their measurement relations and substantive interpretation. A factor with acceptable numerical fit does not acquire a determinate empirical meaning independently of its indicators. Changes in indicator composition can alter the construct represented by the factor even when the model retains the same verbal label.
Large samples improve precision but do not correct structural misspecification. Small samples can produce unstable covariance estimates, boundary solutions, and poorly calibrated asymptotic statistics. Missing observations introduce further dependence on assumptions about the mechanism generating missing data, especially when the probability of missingness is related to unobserved values.