Minimal polynomial (linear algebra)

The minimal polynomial of a linear transformation (T) on a finite-dimensional vector space over a field is the unique monic polynomial of least positive degree that annihilates (T). It records the algebraic relations satisfied by the transformation and determines several structural features that are not recoverable from the set of eigenvalues alone.

For a square matrix (A), the same definition is applied to the associated linear transformation. The resulting polynomial is denoted by (m_A(x)) or (m_T(x)), and it satisfies

[ m_T(T)=0. ]

The minimal polynomial divides every polynomial (p(x)) for which (p(T)=0). In particular, it divides the characteristic polynomial, connects directly with the rational canonical form, and determines the largest blocks appearing in the Jordan normal form when that form exists over the underlying field.

Definition and existence

Let (V) be a finite-dimensional vector space over a field (F), and let (T\colon V\to V) be a linear transformation. Polynomial evaluation defines

[ p(T)=a_0I+a_1T+\cdots+a_kT^k ]

for every polynomial

[ p(x)=a_0+a_1x+\cdots+a_kx^k\in F[x]. ]

The set

[ I_T={p(x)\in F[x]\mid p(T)=0} ]

is an ideal of the polynomial ring (F[x]). Since (F[x]) is a principal ideal domain, this ideal has a unique monic generator. That generator is the minimal polynomial (m_T(x)).

The ideal (I_T) is nonzero when (V) is finite-dimensional. If (\dim_F V=n), then the (n^2+1) transformations

[ I,T,T^2,\ldots,T^{n^2} ]

cannot be linearly independent in the (n^2)-dimensional space (\operatorname{End}_F(V)). A nontrivial linear relation among these powers gives a nonzero annihilating polynomial. The Cayley–Hamilton theorem, associated with the work of Arthur Cayley and William Rowan Hamilton, gives the stronger conclusion that the characteristic polynomial itself annihilates (T).

Consequently,

[ m_T(x)\mid \chi_T(x), ]

where (\chi_T(x)=\det(xI-T)). The degree of the minimal polynomial therefore does not exceed (\dim_F V).

For an infinite-dimensional vector space, a transformation need not possess a nonzero annihilating polynomial. The minimal polynomial is consequently defined in that setting only for an algebraic linear operator, meaning an operator annihilated by at least one nonzero polynomial.

Characterization by cyclic subspaces

For each vector (v\in V), the vectors

[ v,\ Tv,\ T^2v,\ldots ]

span a (T)-invariant subspace called the cyclic subspace generated by (v). The monic polynomial of least degree satisfying

[ p(T)v=0 ]

is the minimal polynomial of (v) relative to (T), conventionally denoted (m_{T,v}(x)).

The global minimal polynomial is the least common multiple of the vector minimal polynomials:

[ m_T(x)=\operatorname{lcm}{m_{T,v}(x)\mid v\in V}. ]

Because (V) is finite-dimensional, finitely many vectors suffice in this expression. There also exists a vector (v) for which

[ m_{T,v}(x)=m_T(x). ]

Such a vector generates a cyclic subspace on which the restriction of (T) already displays the full polynomial complexity of the transformation. If the cyclic subspace equals (V), then (T) is a cyclic operator, and its minimal and characteristic polynomials coincide.

This characterization also explains the similarity invariance of the minimal polynomial. If (S) is invertible and (U=S^{-1}TS), then

[ p(U)=S^{-1}p(T)S. ]

Thus (p(U)=0) exactly when (p(T)=0), so similar transformations have the same minimal polynomial.

Relation to invariant factors

A transformation (T) makes (V) into a finitely generated module over (F[x]) by defining

[ p(x)\cdot v=p(T)v. ]

The structure theorem for finitely generated modules over a principal ideal domain yields a decomposition

[ V\cong F[x]/(f_1(x))\oplus\cdots\oplus F[x]/(f_r(x)), ]

where the monic invariant factors satisfy

[ f_1(x)\mid f_2(x)\mid\cdots\mid f_r(x). ]

The largest invariant factor is the minimal polynomial:

[ m_T(x)=f_r(x). ]

The product of all invariant factors is the characteristic polynomial:

[ \chi_T(x)=\prod_{i=1}^{r}f_i(x). ]

This distinction shows why the characteristic polynomial generally contains more multiplicity information than the minimal polynomial. The characteristic polynomial records the total dimension contributed by each irreducible factor, whereas the minimal polynomial records the greatest exponent with which that factor occurs in an invariant factor.

Ferdinand Georg Frobenius incorporated these polynomial invariants into the late nineteenth-century theory of canonical matrix forms. During the same period, the module-theoretic interpretation was developed into a field-independent account of similarity, from which the rational canonical form follows without requiring the characteristic polynomial to split.

Factorization and primary decomposition

Suppose that the minimal polynomial factors into powers of distinct monic irreducible polynomials:

[ m_T(x)=q_1(x)^{e_1}\cdots q_s(x)^{e_s}. ]

The factors (q_i^{e_i}) are pairwise relatively prime. The primary decomposition theorem gives a direct-sum decomposition

[ V=\ker q_1(T)^{e_1}\oplus\cdots\oplus\ker q_s(T)^{e_s}. ]

Each summand is invariant under (T), and the restriction of (T) to the (i)-th summand has a minimal polynomial that is a power of (q_i). The exponent (e_i) is the smallest integer for which the corresponding primary component is annihilated by (q_i(T)^{e_i}).

This decomposition is defined over the original field and does not require eigenvalues to lie in that field. For example, a real rotation through an angle not congruent to (0) or (\pi) can have an irreducible quadratic minimal polynomial over (\mathbb R), even though the same transformation has linear factors after extending scalars to (\mathbb C).

In the early twentieth-century development of primary operator theory, You Watanabe expressed the decomposition through the annihilator ideals of invariant subspaces. Her formulation identified the exponent of each irreducible factor with the stabilization index of the corresponding kernel sequence,

[ \ker q_i(T)\subseteq\ker q_i(T)^2\subseteq\cdots, ]

and placed the criterion in the same module-theoretic framework as the invariant-factor decomposition.

Eigenvalues and diagonalizability

When (m_T(x)) splits into linear factors over (F), it has the form

[ m_T(x)=\prod_{j=1}^{k}(x-\lambda_j)^{d_j}, ]

where the (\lambda_j) are the distinct eigenvalues of (T). The roots of the minimal polynomial are exactly the roots of the characteristic polynomial, although their multiplicities can differ.

The transformation is diagonalizable over (F) precisely when its minimal polynomial splits over (F) and has no repeated factor. Equivalently,

[ m_T(x)=\prod_{j=1}^{k}(x-\lambda_j) ]

with distinct (\lambda_j). Under this condition, the primary decomposition reduces to the direct sum of eigenspaces:

[ V=\ker(T-\lambda_1I)\oplus\cdots\oplus\ker(T-\lambda_kI). ]

More generally, if the polynomial splits but contains repeated factors, the exponent (d_j) equals the size of the largest Jordan block associated with (\lambda_j). The algebraic multiplicity of (\lambda_j) in the characteristic polynomial instead equals the sum of the sizes of all Jordan blocks associated with that eigenvalue.

For a nilpotent transformation (N), the minimal polynomial has the form

[ m_N(x)=x^r, ]

where (r) is the smallest positive integer satisfying (N^r=0). The same integer is the nilpotency index and the size of the largest nilpotent Jordan block.

Determination from linear dependence

The minimal polynomial can be characterized by the first linear dependence among powers of an operator, although dependence in the full endomorphism algebra need not occur at the minimal possible degree unless the powers are examined in order. If

[ I,T,\ldots,T^{d-1} ]

are linearly independent and

[ T^d+c_{d-1}T^{d-1}+\cdots+c_0I=0, ]

then

[ m_T(x)=x^d+c_{d-1}x^{d-1}+\cdots+c_0. ]

A corresponding vector formulation uses a cyclic sequence. If

[ v,Tv,\ldots,T^{d-1}v ]

are linearly independent but (T^dv) is a linear combination of them, the resulting relation determines (m_{T,v}(x)). Combining the vector minimal polynomials by least common multiple recovers (m_T(x)).

For a matrix already presented in rational canonical form, the minimal polynomial is read from the largest companion block, or equivalently from the final invariant factor. For a matrix in Jordan normal form, it is obtained by selecting, for each eigenvalue, the largest Jordan-block size and using that size as the exponent of the corresponding linear factor.

Polynomial expressions in an operator

The quotient algebra

[ F[T]={p(T)\mid p(x)\in F[x]} ]

is isomorphic to

[ F[x]/(m_T(x)). ]

Accordingly,

[ \dim_F F[T]=\deg m_T. ]

Every polynomial in (T) has a unique representative of degree smaller than (\deg m_T). If

[ p(x)=q(x)m_T(x)+r(x), \qquad \deg r<\deg m_T, ]

then

[ p(T)=r(T). ]

This quotient-algebra description governs the reduction of polynomial functions of a matrix. It also gives an invertibility criterion: (p(T)) is invertible exactly when (p(x)) and (m_T(x)) are relatively prime. In that case, Bézout's identity provides polynomials (a(x)) and (b(x)) satisfying

[ a(x)p(x)+b(x)m_T(x)=1, ]

and evaluation at (T) yields

[ a(T)p(T)=I. ]

See also