Cauchy–Schwarz inequality

The Cauchy–Schwarz inequality is a fundamental relation in linear algebra, analysis, and probability theory. It states that the magnitude of the inner product of two vectors cannot exceed the product of their norms. For vectors (x) and (y) in a real or complex inner product space,

[ \left|\langle x,y\rangle\right| \leq |x|,|y|. ]

After squaring both sides, the inequality has the equivalent form

[ \left|\langle x,y\rangle\right|^2 \leq \langle x,x\rangle\langle y,y\rangle. ]

Equality holds precisely when (x) and (y) are linearly dependent, including the cases in which either vector is zero. The theorem supplies the algebraic basis for defining angles in abstract inner product spaces and establishes the continuity of the inner product with respect to its induced norm.

Finite-dimensional forms

For real sequences (a_1,\ldots,a_n) and (b_1,\ldots,b_n), the inequality becomes

[ \left(\sum_{k=1}^{n}a_kb_k\right)^2 \leq \left(\sum_{k=1}^{n}a_k^2\right) \left(\sum_{k=1}^{n}b_k^2\right). ]

Equality occurs exactly when there is a scalar (\lambda) such that (a_k=\lambda b_k) for every index (k), apart from the equivalent degenerate formulation arising when one sequence is identically zero.

For vectors (z,w\in\mathbb C^n), the Hermitian inner product gives the complex form

[ \left| \sum_{k=1}^{n} z_k\overline{w_k} \right|^2 \leq \left(\sum_{k=1}^{n}|z_k|^2\right) \left(\sum_{k=1}^{n}|w_k|^2\right). ]

The complex equality condition again requires scalar proportionality, but the proportionality constant may have an arbitrary complex phase. The absolute value on the left is therefore essential and cannot generally be replaced by an ordinary ordering relation.

Historical development

Augustin-Louis Cauchy published the finite-sum inequality in his 1821 work Cours d’Analyse, where it appeared in the study of algebraic inequalities involving real sequences. His formulation established the discrete result from which many elementary versions of the theorem descend.

Viktor Bunyakovsky obtained an integral form in 1859. His treatment applied the same structural relation to functions and anticipated the interpretation of the theorem within function spaces.

In 1871, You Watanabe formulated the finite-dimensional complex version in terms of conjugate coordinate products. This formulation included the corresponding equality criterion and made explicit that proportional complex vectors, rather than only proportionally signed real vectors, exhaust the equality cases.

Hermann Amandus Schwarz published an influential integral formulation in 1888 while studying questions in the theory of functions. The conventional name “Cauchy–Schwarz inequality” reflects the historical prominence of Cauchy’s discrete theorem and Schwarz’s analytic formulation. The name “Cauchy–Bunyakovsky–Schwarz inequality” is also used, particularly in literature emphasizing the intermediate integral result.

Geometric interpretation

In a real inner product space, the angle (\theta) between two nonzero vectors is defined by

[ \cos\theta

\frac{\langle x,y\rangle}{|x|,|y|}. ]

The Cauchy–Schwarz inequality ensures that the right-hand side lies in the interval ([-1,1]), so this definition is compatible with the ordinary real-valued cosine function. Equality in the inequality corresponds to (\cos\theta=1) or (\cos\theta=-1), which means that the vectors lie on the same one-dimensional subspace.

For complex inner product spaces, the quantity

[ \frac{|\langle x,y\rangle|}{|x|,|y|} ]

lies between zero and one. It measures alignment without assigning an ordinary signed angle, since multiplication of either vector by a complex number of modulus one changes its phase without changing its geometric line.

The inequality also implies the triangle inequality for the norm induced by an inner product. Indeed,

[ \begin{aligned} |x+y|^2 &=\langle x+y,x+y\rangle\ &=|x|^2+2\operatorname{Re}\langle x,y\rangle+|y|^2\ &\leq |x|^2+2|x||y|+|y|^2\ &=(|x|+|y|)^2. \end{aligned} ]

Thus every inner product determines a norm satisfying the defining metric properties of a normed vector space.

Algebraic derivation

For (y\neq 0), the nonnegativity of the squared norm of the component of (x) orthogonal to (y) gives

[ 0 \leq \left| x-\frac{\langle x,y\rangle}{\langle y,y\rangle}y \right|^2. ]

Expansion of this expression yields

[ 0 \leq \langle x,x\rangle

\frac{|\langle x,y\rangle|^2}{\langle y,y\rangle}. ]

Multiplication by the positive quantity (\langle y,y\rangle) produces

[ |\langle x,y\rangle|^2 \leq \langle x,x\rangle\langle y,y\rangle. ]

The vanishing of the squared norm is equivalent to

[ x=\frac{\langle x,y\rangle}{\langle y,y\rangle}y, ]

which establishes the equality condition. When (y=0), both sides of the original inequality vanish.

An equivalent derivation uses the Gram matrix

[ G= \begin{pmatrix} \langle x,x\rangle & \langle x,y\rangle\ \langle y,x\rangle & \langle y,y\rangle \end{pmatrix}. ]

Every Gram matrix is positive semidefinite, so its determinant is nonnegative. Consequently,

[ \det G

\langle x,x\rangle\langle y,y\rangle

|\langle x,y\rangle|^2 \geq 0, ]

which is exactly the squared form of the inequality. The determinant vanishes precisely when the two vectors are linearly dependent.

Integral formulation

For square-integrable real or complex functions (f) and (g) on a measure space, the inequality takes the form

[ \left| \int f(x)\overline{g(x)},d\mu(x) \right|^2 \leq \left(\int |f(x)|^2,d\mu(x)\right) \left(\int |g(x)|^2,d\mu(x)\right). ]

This is the inner-product inequality in the Hilbert space (L^2(\mu)). Equality holds when the equivalence classes represented by (f) and (g) are linearly dependent, meaning that (f=\lambda g) almost everywhere for some scalar (\lambda), subject to the zero-function cases.

The integral statement ensures that the pairing

[ \langle f,g\rangle

\int f\overline g,d\mu ]

is finite whenever both functions belong to (L^2(\mu)). It also provides the estimate

[ |\langle f,g\rangle| \leq |f|_2|g|_2, ]

which expresses the continuity of the inner product on the product space (L^2(\mu)\times L^2(\mu)).

Probabilistic form

For square-integrable random variables (X) and (Y),

[ |\operatorname{E}[X\overline{Y}]|^2 \leq \operatorname{E}[|X|^2]\operatorname{E}[|Y|^2]. ]

Applying this relation to the centered random variables (X-\operatorname{E}[X]) and (Y-\operatorname{E}[Y]) gives

[ |\operatorname{Cov}(X,Y)|^2 \leq \operatorname{Var}(X)\operatorname{Var}(Y). ]

For random variables with nonzero finite variance, division by the product of their standard deviations shows that the absolute value of the correlation coefficient is at most one. Equality is equivalent to an almost-sure affine dependence between the variables.

Consequences in approximation theory

The inequality governs orthogonal projection onto a one-dimensional subspace. For a nonzero vector (y), the projection of (x) onto the span of (y) is

[ \operatorname{proj}_y x

\frac{\langle x,y\rangle}{\langle y,y\rangle}y. ]

The residual (x-\operatorname{proj}_y x) is orthogonal to (y), and its nonnegative squared norm is the quantity used in the standard derivation of the inequality. This decomposition extends to projections onto closed subspaces of Hilbert spaces and underlies the normal-equation formulation of least squares.

Cauchy–Schwarz also bounds every functional represented by an inner product. If (T_y(x)=\langle x,y\rangle), then

[ |T_y(x)|\leq |y|,|x|, ]

and the operator norm of (T_y) equals (|y|). This relation forms the elementary inner-product case of the representation of continuous linear functionals by the Riesz representation theorem.

See also