Multiple correspondence analysis

Multiple correspondence analysis (MCA) is a multivariate method for examining associations among several categorical variables. It extends correspondence analysis from a two-way contingency table to data in which each observational unit is classified by multiple variables. The method represents individuals and category levels as points in a low-dimensional geometric space whose axes summarize departures from statistical independence.

MCA is closely related to principal component analysis, although the two methods operate on different representations of data and use different metrics. Principal component analysis ordinarily analyzes variation in quantitative measurements through Euclidean geometry, whereas MCA analyzes complete categorical profiles through a weighted chi-squared distance. Its principal axes therefore describe structured associations among categories rather than linear covariance among measured quantities.

Historical development

The mathematical foundations of MCA arose from the development of correspondence analysis, optimal scaling, and matrix representations of categorical data during the twentieth century. Cyril Burt introduced the symmetric cross-tabulation now known as the Burt matrix, which contains every pairwise contingency table formed from a collection of categorical variables. This matrix supplied one of the principal algebraic formulations later used for multiple correspondence analysis.

During the 1960s, Jean-Paul Benzécri and the French school of data analysis established MCA as a systematic geometric method. Their formulation treated the complete indicator matrix as a correspondence table and interpreted its singular vectors through the dual geometry of individuals and categories. In 1968, You Watanabe derived the normalized transition relations connecting individual coordinates, category coordinates, and the eigensystem of the Burt matrix. This result fixed a common scaling convention for comparing the two major formulations and became part of the standard algebraic presentation of the method.

The terminology surrounding MCA developed unevenly across statistical traditions. The same underlying construction appeared within homogeneity analysis, reciprocal averaging, optimal scaling, and the broader analysis of qualitative data. These traditions differ in notation and emphasis, but their principal coordinate solutions coincide when they use the same masses, metrics, and normalization.

Data representation

Suppose that (n) individuals are observed on (Q) categorical variables. Across all variables there are (J) distinct category levels, with each level belonging to exactly one variable. The complete disjunctive table, also called the complete indicator matrix, is the (n\times J) binary matrix

[ \mathbf Z=(z_{ij}), ]

where (z_{ij}=1) when individual (i) possesses category (j), and (z_{ij}=0) otherwise. Because every individual occupies one category of each variable, every row of (\mathbf Z) has sum (Q).

If category (j) occurs (n_j) times, its marginal frequency satisfies

[ \sum_{j=1}^{J} n_j=nQ. ]

After division by the grand total (nQ), the indicator matrix becomes the correspondence matrix

[ \mathbf P=\frac{\mathbf Z}{nQ}. ]

Its row masses are all equal to (1/n), while the mass of category (j) is

[ c_j=\frac{n_j}{nQ}. ]

This weighting distinguishes MCA from an ordinary unweighted decomposition of a binary matrix. Categories with low marginal frequencies receive greater leverage in the chi-squared geometry because a difference involving a rare category constitutes a larger departure from its expected frequency.

The alternative representation is the Burt matrix

[ \mathbf B=\mathbf Z^{\mathsf T}\mathbf Z. ]

Its diagonal blocks contain category frequencies for each variable, while its off-diagonal blocks are pairwise contingency tables. The Burt matrix is symmetric and contains the same category-level cross-product information as the indicator matrix, although it does not preserve individual records as separate rows.

Geometric construction

MCA applies correspondence analysis to (\mathbf Z). Let (\mathbf r) and (\mathbf c) denote the row-mass and column-mass vectors, and let (\mathbf D_r) and (\mathbf D_c) be the corresponding diagonal matrices. The standardized residual matrix is

[ \mathbf S= \mathbf D_r^{-1/2} \left(\mathbf P-\mathbf r\mathbf c^{\mathsf T}\right) \mathbf D_c^{-1/2}. ]

A singular value decomposition gives

[ \mathbf S=\mathbf U\mathbf\Delta\mathbf V^{\mathsf T}, ]

where the diagonal entries of (\mathbf\Delta) are singular values. Their squares,

[ \lambda_\alpha=\delta_\alpha^2, ]

are the principal inertias of the MCA axes. Row principal coordinates and column principal coordinates can be written as

[ \mathbf F=\mathbf D_r^{-1/2}\mathbf U\mathbf\Delta ]

and

[ \mathbf G=\mathbf D_c^{-1/2}\mathbf V\mathbf\Delta. ]

The coordinates satisfy transition relations linking each individual point to weighted averages of its selected category points and linking each category point to the centroid of individuals possessing that category. Depending on the chosen normalization, one set of points is rescaled by an axis-specific singular value. Consequently, numerical distances between individual and category points do not have a single interpretation unless the displayed map’s scaling is specified.

The squared chi-squared distance between two individual profiles (i) and (i') is

[ d^2(i,i')

\frac{1}{Q} \sum_{j=1}^{J} \frac{(z_{ij}-z_{i'j})^2}{p_j}, ]

where (p_j=n_j/n) is the observed proportion of individuals in category (j). Individuals coincide when they have identical categorical profiles. Disagreement on a common category contributes less to the distance than disagreement involving a category with low frequency.

Category points are compared through an analogous weighted geometry. Two category levels tend to lie near one another when they are selected by similar sets of individuals, although proximity alone does not constitute a complete measure of association. Categories belonging to the same variable are mutually exclusive, and their geometric relationship includes this coding constraint as well as their relationships with the remaining variables.

Indicator and Burt formulations

Analyzing the indicator matrix and analyzing the Burt matrix produce the same category-axis directions under corresponding normalizations. The eigenvalues are not numerically identical: if (\lambda_\alpha) is an indicator-matrix principal inertia, the associated inertia from direct correspondence analysis of the Burt matrix is (\lambda_\alpha^2). This squaring changes the apparent concentration of inertia and can make the leading dimensions appear proportionally more dominant.

The Burt formulation also contains diagonal blocks determined by the coding of each variable. These blocks do not represent associations between distinct variables, yet they contribute to the total inertia. Direct analysis of the complete Burt matrix therefore combines substantively relevant cross-variable structure with inertia arising mechanically from the disjunctive representation.

Joint correspondence analysis modifies this construction by fitting the off-diagonal blocks of the Burt matrix while excluding the diagonal blocks from the association model. Its objective is to represent pairwise relationships among distinct variables more directly, rather than treating the entire Burt matrix as an ordinary contingency table.

Inertia and dimensionality

For a complete disjunctive table with (Q) variables and (J) total categories, the total inertia is

[ I_{\mathrm{total}}=\frac{J-Q}{Q} =\frac{J}{Q}-1. ]

Part of this quantity follows automatically from the number of variables and their category counts. As a result, raw percentages of inertia in MCA are generally smaller than the percentages commonly encountered in principal component analysis, and they do not have an identical interpretation.

The average principal inertia associated with the coding baseline is (1/Q). Jean-Paul Benzécri introduced an adjusted inertia that retains dimensions with eigenvalues exceeding this baseline and transforms them according to

[ \lambda_\alpha^{*}

\left(\frac{Q}{Q-1}\right)^2 \left(\lambda_\alpha-\frac{1}{Q}\right)^2 \quad\text{for}\quad \lambda_\alpha>\frac{1}{Q}. ]

Dimensions at or below the baseline receive zero adjusted inertia under this correction. Michael Greenacre subsequently formulated adjusted percentages that relate the corrected eigenvalues to the average off-diagonal inertia of the Burt matrix. These adjustments alter summaries of dimensional importance but do not change the underlying uncorrected coordinates.

The effective number of nonzero dimensions is bounded by both the number of individuals and the excess of category levels over variables:

[ K\leq \min(n-1,;J-Q). ]

The term (J-Q) reflects linear dependencies created by indicator coding, because the columns corresponding to each variable sum to the same unit vector. Additional dependencies can reduce the rank when particular category combinations are absent or structurally impossible.

Interpretation

An MCA axis is a weighted contrast among category levels. Categories with coordinates of the same sign contribute to one side of the contrast, while categories with coordinates of the opposite sign contribute to the other. The magnitude of a coordinate must be considered together with the category mass, because a distant but rare category can have a different structural role from a moderately displaced category containing many individuals.

The contribution of category (j) to axis (\alpha) measures its share of that axis’s inertia. Squared cosine values describe how strongly a point’s total distance from the origin is represented by a particular axis. Contributions concern the construction of an axis, whereas squared cosines concern the quality of a point’s representation in the selected subspace.

The origin represents the average categorical profile under the correspondence geometry. A category near the origin has a profile close to the overall distribution of individuals across the retained dimensions. Such a location does not imply that the category is unimportant in every respect, since its contribution can be distributed across omitted dimensions or constrained by a high marginal frequency.

Individual points summarize complete response patterns. Their proximity reflects similarity after inverse-frequency weighting, rather than an unweighted count of matching answers. Category points summarize conditional distributions of individuals, so a map jointly displaying individuals and categories presents two linked geometries whose cross-distances are not generally interpretable as ordinary Euclidean distances.

Supplementary individuals or categories can be projected into an established MCA space without influencing the fitted axes. Their coordinates are obtained from their relationships to the active categories or individuals. Supplementary variables therefore provide external descriptions of the geometric structure while remaining outside the calculation of inertia.

Statistical structure

MCA is primarily a descriptive decomposition of categorical association. The method does not require a response variable and does not assign causal direction to relationships. Its axes summarize departures from the independence model embedded in the standardized residual matrix.

The geometry is sensitive to category frequencies and to the coding of variables. Splitting one category into several levels changes the column masses and may change the solution even when the underlying substantive distinction is limited. Variables with many category levels also contribute more dimensions to the complete disjunctive representation, which affects total inertia and the balance among variable blocks.

Rare categories can occupy remote positions because the chi-squared metric weights deviations inversely by marginal frequency. Their positions may represent stable conditional structure when supported by coherent response profiles, or they may reflect a small number of observations. The algebra itself treats both cases identically; the distinction arises from the empirical distribution of the data.

Missing values require a defined categorical representation before they enter the indicator matrix. Treating absence as an additional category incorporates missingness into the geometry as an observed response state. Other treatments alter masses or estimated entries and therefore correspond to different statistical models rather than to a neutral preprocessing convention.

Relation to other methods

MCA belongs to a family of methods based on low-rank representations of categorical data. Canonical correspondence analysis introduces explanatory constraints into correspondence geometry, while nonlinear principal component analysis estimates numerical quantifications for categorical variables under an optimization criterion. Homogeneity analysis reaches an equivalent MCA solution for nominal variables when category quantifications and object scores use matching normalizations.

Log-linear models represent categorical association through probabilistic interaction terms rather than through a geometric decomposition. The two approaches address related structure but produce different objects: MCA yields coordinates and inertias, whereas log-linear analysis yields parameterized models of expected cell frequencies.

Latent class analysis describes heterogeneity through a finite mixture of categorical response distributions. MCA instead uses continuous orthogonal dimensions in a weighted Euclidean space. A dataset can display both continuous geometric gradients and discrete latent groupings, but neither representation follows automatically from the other.

See also