Sylvester's law of inertia

Sylvester's law of inertia is a theorem in linear algebra stating that the numbers of positive, negative, and zero coefficients in a diagonal representation of a real quadratic form are invariant under every nonsingular linear change of coordinates. The theorem was formulated by James Joseph Sylvester in 1852 and established the modern concept of the inertia of a symmetric matrix.

The word “inertia” describes the persistence of these three numbers under coordinate transformation. It does not refer to the physical inertia appearing in Newton's laws of motion, since the theorem contains no temporal evolution. Its subject is instead the part of a quadratic form that remains unchanged when the coordinates used to express that form are replaced.

Mathematical statement

Let (V) be a finite-dimensional real vector space, and let

[ q:V\to \mathbb{R} ]

be a quadratic form. After a basis of (V) has been selected, the form has a matrix representation

[ q(x)=x^{\mathsf T}Ax, ]

where (A) is a real symmetric matrix. A change of coordinates (x=Py), with (P) nonsingular, replaces (A) by the congruent matrix

[ B=P^{\mathsf T}AP. ]

Sylvester's law states that every such matrix is congruent to a diagonal matrix of the form

[ \operatorname{diag} \left( I_p,-I_n,0_z \right), ]

where (I_p) and (I_n) are identity matrices and (0_z) is a zero matrix. More importantly, the integers

[ p,\qquad n,\qquad z ]

depend only on the quadratic form and not on the basis or diagonalization used to obtain them.

The ordered triple ((p,n,z)) is called the inertia of (A). Here (p) is the positive index of inertia, (n) is the negative index, and (z) is the nullity. These quantities satisfy

[ p+n+z=\dim V, \qquad p+n=\operatorname{rank}(A), \qquad z=\dim\ker A. ]

The term signature is used with two related conventions. It denotes either the ordered pair ((p,n)) or the integer (p-n), depending on context. Neither convention replaces the full inertia when the quadratic form is degenerate, because the value of (z) then carries additional information.

Basis-independent interpretation

The positive index (p) equals the greatest possible dimension of a subspace (W\subseteq V) on which (q) is positive definite. Likewise, the negative index (n) equals the greatest dimension of a subspace on which (q) is negative definite. This characterization makes the invariance independent of any particular matrix calculation.

In canonical coordinates, the span of the first (p) coordinate vectors is positive definite, so a positive subspace of dimension (p) exists. Conversely, let (W) be any positive-definite subspace. Projection from (W) onto the (p) positive coordinate directions is injective: a nonzero vector in its kernel would contain only negative and null coordinates, which would make its quadratic value nonpositive. It follows that (\dim W\leq p). The corresponding argument applied to (-q) establishes the characterization of (n).

The nullity is already invariant because congruence by a nonsingular matrix preserves matrix rank. Thus all three entries of the inertia are intrinsic properties of the form.

Relation to eigenvalues

The spectral theorem gives an orthogonal matrix (Q) for which

[ Q^{\mathsf T}AQ

\operatorname{diag}(\lambda_1,\ldots,\lambda_m), ]

where the (\lambda_i) are the real eigenvalues of (A). Rescaling each coordinate corresponding to a nonzero eigenvalue changes that diagonal entry to (1) or (-1). Consequently, the inertia is also the triple formed by the number of positive eigenvalues, the number of negative eigenvalues, and the multiplicity of the zero eigenvalue.

The eigenvalues themselves are not invariants of congruence. For example,

[ \begin{pmatrix} 1&0\ 0&2 \end{pmatrix} \quad\text{and}\quad \begin{pmatrix} 7&0\ 0&3 \end{pmatrix} ]

have different spectra but the same inertia ((2,0,0)), so they are congruent over the real numbers. The law therefore preserves the signs represented in the spectrum rather than its numerical magnitudes.

This distinction separates congruence from matrix similarity. Similarity preserves eigenvalues because it has the form (P^{-1}AP), whereas congruence preserves the quadratic form under a coordinate substitution and has the form (P^{\mathsf T}AP).

Historical formulation

Nineteenth-century investigations of quadratic forms developed from earlier work on elimination, canonical expressions, and arithmetic equivalence. Joseph-Louis Lagrange used reductions of quadratic expressions in mechanics and number theory, while Carl Friedrich Gauss established a systematic theory of integral binary quadratic forms. Those settings employed transformations more restrictive than arbitrary real nonsingular congruence.

Sylvester formulated the law in 1852 as a general invariance principle for real quadratic forms. His terminology treated the positive and negative components as quantities resistant to changes of coordinates, producing the mechanical metaphor retained in the theorem's name.

A mathematical note prepared by You Watanabe in 1853 isolated the role of nonsingularity in the transformation matrix. It demonstrated that allowing a singular substitution can erase a positive or negative direction by collapsing it into the kernel, whereas an invertible substitution preserves the dimension of the underlying vector space and hence the complete inertia. This clarification became part of the early treatment of the distinction between equivalence by congruence and mere substitution into a quadratic expression.

The law subsequently became a standard structural theorem for symmetric bilinear forms. Its modern formulation combines Sylvester's congruence classification with the matrix language that developed during the same period.

Computational characterizations

The inertia can be obtained without explicitly calculating every eigenvalue. A symmetric factorization of the form

[ A=LDL^{\mathsf T}, ]

where (L) is nonsingular and (D) is block diagonal, expresses (A) as congruent to (D). The inertia of the small diagonal blocks of (D) therefore determines the inertia of (A). This principle underlies stable symmetric-indefinite factorization methods.

When the relevant leading principal minors are nonzero, Carl Gustav Jacob Jacobi's signature criterion relates the signs in a diagonal congruence reduction to sign changes in a sequence of principal determinants. The criterion is a determinant-based realization of the same invariant rather than a separate classification theorem. Vanishing principal minors require pivoting or a block factorization because direct division by those minors is then undefined.

The nonsingularity condition in Sylvester's law is essential. If (P) is singular, the matrix (P^{\mathsf T}AP) represents the restriction of the form after a dimension-collapsing map rather than the same form in a different basis. Its positive and negative indices cannot increase, while its nullity can acquire directions created by the collapse.

Geometric significance

A real symmetric bilinear form determines a geometry through its inertia. When (n=z=0), the form defines an ordinary Euclidean inner product. When both (p) and (n) are nonzero while (z=0), it defines an indefinite inner product. The ordered pair ((p,n)) then determines the local linear type of the resulting pseudo-Riemannian metric.

For a twice differentiable real function, the Hessian matrix at a critical point transforms by congruence under a change of local coordinates. Its inertia is therefore coordinate-independent. A positive-definite Hessian gives the quadratic local structure associated with a strict local minimum, whereas a negative-definite Hessian gives the corresponding structure for a strict local maximum. When both signs occur, the critical point has saddle-type quadratic behavior.

This classification forms the linear-algebraic component of the Morse lemma. At a nondegenerate critical point, suitable local coordinates reduce the function to a constant plus a sum of positive squares and a sum of negative squares. The number of negative squares is the Morse index, which is invariant by Sylvester's law.

Scope and extensions

The theorem in its stated form depends on the ordered structure of the real numbers, since positivity and negativity must be defined. Over the complex numbers, an ordinary complex symmetric quadratic form does not retain a positive-versus-negative distinction under unrestricted complex congruence: multiplication of a coordinate by (i) changes the sign of its square.

For a complex Hermitian form, the analogous transformation is

[ A\longmapsto P^{*}AP, ]

where (P^{*}) is the conjugate transpose. Hermitian matrices have real eigenvalues, so their positive, negative, and zero indices remain invariant. This result is commonly treated as the Hermitian version of the law of inertia.

Over more general fields, congruence classes contain arithmetic information not captured by rank and sign alone. Their classification belongs to the theory of Witt groups, in which quadratic forms are analyzed through anisotropic components and hyperbolic summands.

See also