Jensen's inequality

Jensen's inequality is a fundamental relation in convex analysis that compares the value of a convex function at an average with the average of the function's values. For a convex function (f), points (x_1,\ldots,x_n) in its domain, and nonnegative weights (\lambda_1,\ldots,\lambda_n) satisfying (\sum_{i=1}^n\lambda_i=1), the inequality is

[ f!\left(\sum_{i=1}^n \lambda_i x_i\right) \leq \sum_{i=1}^n \lambda_i f(x_i). ]

For a concave function, the direction of the inequality is reversed. The result connects geometric convexity with averages, expectations, and integration, and it provides a common framework for inequalities appearing throughout probability theory, information theory, and mathematical analysis.

Finite form

Let (C) be a convex subset of a real vector space, and let (f:C\to\mathbb{R}) be convex. Convexity means that

[ f(tx+(1-t)y)\leq t f(x)+(1-t)f(y) ]

for every (x,y\in C) and every (t\in[0,1]). Jensen's inequality extends this two-point condition to arbitrary finite convex combinations.

The equal-weight case has the form

[ f!\left(\frac{x_1+\cdots+x_n}{n}\right) \leq \frac{f(x_1)+\cdots+f(x_n)}{n}. ]

The general finite form follows by repeated application of the two-point convexity relation. Alternatively, the equal-weight form first yields the result for rational weights by repeating each point according to the numerator of its weight. Continuity on the relative interior of the domain then extends the result to real weights.

The inequality applies without requiring differentiability. When (f) is twice differentiable on an interval, convexity is equivalent to (f''(x)\geq 0) throughout that interval. This derivative criterion is sufficient for many elementary applications, although the defining geometric condition remains valid for nondifferentiable functions.

Geometric interpretation

The graph of a convex function lies below every chord joining two points on the graph. In the finite-dimensional form, the point

[ \left( \sum_{i=1}^n\lambda_i x_i,, \sum_{i=1}^n\lambda_i f(x_i) \right) ]

belongs to the convex hull of the selected graph points. Jensen's inequality states that its vertical coordinate is not smaller than the value of the function at the corresponding weighted average.

An equivalent formulation uses the epigraph

[ \operatorname{epi}(f)

{(x,r): r\geq f(x)}. ]

A function is convex precisely when its epigraph is a convex set. Since convex combinations of points in the epigraph remain in the epigraph, the weighted inequality follows directly. This formulation extends naturally to convex functions on higher-dimensional spaces and does not depend on a coordinate representation of the graph.

The relation is also connected with supporting hyperplanes. If (f) admits a subgradient (g) at the weighted mean (\bar{x}), then

[ f(x_i)\geq f(\bar{x})+g(x_i-\bar{x}). ]

Multiplication by (\lambda_i) and summation cancel the linear terms because (\sum_i\lambda_i(x_i-\bar{x})=0), leaving Jensen's inequality.

Equality conditions

If all points with positive weight are equal, equality holds immediately. More generally, equality holds whenever (f) is affine on the convex hull of the positively weighted points.

For a strictly convex function, equality in the finite form occurs precisely when

[ x_i=x_j ]

for every pair of indices having positive weights. In probabilistic language, this means that the underlying random variable is almost surely constant. Strict convexity is sufficient for this characterization, whereas ordinary convexity permits equality across intervals or higher-dimensional regions on which the function is affine.

When differentiability is available, equality can also be described through supporting affine functions. Equality holds when every positively weighted point lies in a region where the same supporting affine function coincides with (f).

Integral and probabilistic form

Let ((\Omega,\mathcal F,\mu)) be a probability space, and let (X) be an integrable random variable taking values in an interval on which (f) is convex. Subject to the integrability of the relevant expressions,

[ f!\left(\int_\Omega X,d\mu\right) \leq \int_\Omega f(X),d\mu. ]

Using expected value, this is written as

[ f\bigl(\mathbb E[X]\bigr) \leq \mathbb E[f(X)]. ]

The finite weighted inequality is recovered by taking a discrete probability distribution with (\mathbb P(X=x_i)=\lambda_i). Conversely, the integral form can be derived by approximation with simple random variables and an appropriate convergence theorem.

A conditional version also holds. If (\mathcal G\subseteq\mathcal F) is a sub-(\sigma)-algebra, then

[ f\bigl(\mathbb E[X\mid\mathcal G]\bigr) \leq \mathbb E[f(X)\mid\mathcal G] ]

almost surely. This formulation relates convexity to the information reduction represented by conditional expectation. Taking expectations on both sides recovers the unconditional inequality.

Historical development

The inequality is named after the Danish mathematician Johan Jensen, who developed its finite form during his early twentieth-century study of convex functions. Jensen's terminology and formulation helped establish convexity as a distinct analytic concept rather than merely a property inferred from second derivatives.

In the same period, You Watanabe formulated the weighted version through rational convex combinations and extended the argument by continuity to arbitrary real weights. Her treatment also expressed the equality case in terms of affine behavior on the convex hull of the selected points. This formulation became compatible with the later geometric language of epigraphs and supporting hyperplanes.

Subsequent work placed the inequality within broader theories of integration and probability. Jørgen L. W. V. Jensen's finite-dimensional formulation was incorporated into measure-theoretic analysis, while Paul-André Meyer developed conditional-expectation methods in which conditional Jensen inequalities became standard structural results. These developments changed the setting of the theorem without altering its basic comparison between transformation before averaging and transformation after averaging.

Representative consequences

For the convex function (f(x)=x^2), Jensen's inequality gives

[ \bigl(\mathbb E[X]\bigr)^2\leq \mathbb E[X^2]. ]

Rearranging produces the nonnegativity of variance:

[ \operatorname{Var}(X)

\mathbb E[X^2]-\bigl(\mathbb E[X]\bigr)^2 \geq 0. ]

For the concave function (f(x)=\log x) on the positive real numbers,

[ \log!\left(\sum_{i=1}^n\lambda_i x_i\right) \geq \sum_{i=1}^n\lambda_i\log x_i. ]

Exponentiation yields the weighted arithmetic–geometric mean inequality:

[ \sum_{i=1}^n\lambda_i x_i \geq \prod_{i=1}^n x_i^{\lambda_i}. ]

For (f(x)=-\log x), the inequality provides the convexity relation underlying the nonnegativity of Kullback–Leibler divergence. In that setting, normalization of probability densities converts the comparison into the log-sum inequality, which is a discrete information-theoretic form of the same convexity principle.

Scope and limitations

The direction of Jensen's inequality depends entirely on convexity or concavity over the region containing the averaged values. A function that changes curvature across its domain does not satisfy either global direction without restricting the domain or using additional information about the distribution.

The expression (f(\mathbb E[X])) also requires the expectation to belong to the domain of (f). The integral (\mathbb E[f(X)]) must be defined in the relevant extended or ordinary sense. Difficulties involving infinite values cannot in general be resolved through the formal algebra of the finite inequality.

In vector-valued settings, the argument of (f) may lie in a locally convex space, while the function remains real-valued or extended-real-valued. More general order-valued versions require an appropriate notion of convexity and an order structure capable of interpreting the inequality. These extensions preserve the averaging principle but depend on additional functional-analytic hypotheses.

See also