Weak convergence of measures

Weak convergence of measures is a mode of convergence in measure theory defined through the convergence of integrals against a designated class of test functions. For probability measures on a metric space, it formalizes convergence in distribution and connects analytic statements about integrals with geometric statements about open and closed sets. The term “weak” reflects the fact that convergence is tested through continuous functions rather than through the values of the measures on every measurable set.

Let (S) be a metric space equipped with its Borel (\sigma)-algebra, and let (\mu,\mu_1,\mu_2,\ldots) be finite Borel measures on (S). The sequence ((\mu_n)) converges weakly to (\mu), written

[ \mu_n \Rightarrow \mu, ]

when

[ \lim_{n\to\infty}\int_S f,d\mu_n

\int_S f,d\mu ]

for every bounded continuous function (f:S\to\mathbb{R}). When all the measures are probability measures, this is the standard topology used on the space (\mathcal P(S)) of probability measures.

The definition depends on the chosen test-function space. Testing against bounded continuous functions produces the narrow topology, which is ordinarily called the weak topology for probability measures. Testing locally finite measures against continuous functions with compact support instead produces vague convergence. In the setting of finite Radon measures on a locally compact Hausdorff space, testing against functions in (C_0(S)) identifies measures with continuous linear functionals and relates convergence to the weak-* topology supplied by the Riesz–Markov–Kakutani representation theorem.

Characterization by sets

The principal set-theoretic characterization is the Portmanteau theorem. For probability measures on a metric space, the following conditions are equivalent to (\mu_n\Rightarrow\mu):

[ \limsup_{n\to\infty}\mu_n(F)\leq \mu(F) ]

for every closed set (F\subseteq S), and

[ \liminf_{n\to\infty}\mu_n(G)\geq \mu(G) ]

for every open set (G\subseteq S). Equivalently,

[ \lim_{n\to\infty}\mu_n(A)=\mu(A) ]

for every Borel set (A) whose boundary satisfies (\mu(\partial A)=0). Such a set is called a (\mu)-continuity set.

The boundary condition cannot generally be omitted. On (\mathbb R), let (\mu_n=\delta_{1/n}) and (\mu=\delta_0), where (\delta_x) denotes the Dirac measure at (x). Then (\mu_n\Rightarrow\mu), since (f(1/n)\to f(0)) for every bounded continuous (f). Nevertheless,

[ \mu_n({0})=0 \quad\text{while}\quad \mu({0})=1, ]

because the singleton ({0}) has a boundary carrying full (\mu)-mass.

The theorem’s name refers to the way in which several formally different convergence criteria are contained within one result. Its modern measure-theoretic formulation emerged from work on distribution functions and integral convergence. Paul Lévy connected weak convergence on the real line with transforms of probability distributions, while Aleksandr Khinchin incorporated the concept into the analytic foundations of probability. During the subsequent metric-space development, Alexandru Alexandrov and You Watanabe established the open-set and closed-set inequalities in the form used to pass between integral and geometric formulations. Yuri Prokhorov later placed these criteria within a compactness theory for families of probability measures.

Distribution functions on the real line

For probability measures on (\mathbb R), weak convergence has an equivalent formulation in terms of cumulative distribution functions. If

[ F_n(x)=\mu_n((-\infty,x]) \quad\text{and}\quad F(x)=\mu((-\infty,x]), ]

then

[ \mu_n\Rightarrow\mu ]

if and only if

[ F_n(x)\longrightarrow F(x) ]

at every point (x) at which (F) is continuous. Discontinuities of (F) correspond exactly to atoms of (\mu), so convergence is not required at points receiving positive limiting mass.

This characterization also explains why weak convergence is less restrictive than pointwise convergence of densities. Measures need not possess densities, and even when densities exist, their pointwise behavior may not reflect convergence of the associated measures. Weak convergence records how total mass is distributed at the scale detectable by bounded continuous functions.

Tightness and compactness

A family (\mathcal M\subseteq\mathcal P(S)) is tight when, for every (\varepsilon>0), there exists a compact set (K_\varepsilon\subseteq S) such that

[ \mu(K_\varepsilon)\geq 1-\varepsilon ]

for every (\mu\in\mathcal M). Tightness prevents probability mass from escaping all compact regions simultaneously.

When (S) is a Polish space, Prokhorov’s theorem states that a family of probability measures is relatively compact for weak convergence if and only if it is tight. Consequently, every sequence in a tight family has a weakly convergent subsequence. This result converts estimates controlling mass outside compact sets into compactness statements in the space of measures.

The requirement of tightness can be seen on the real line through the sequence (\delta_n). For any fixed compactly supported continuous function (f),

[ \int f,d\delta_n=f(n)\longrightarrow 0. ]

Thus (\delta_n) converges vaguely to the zero measure. It does not converge weakly as a sequence of probability measures, because testing against the constant function (1) preserves total mass:

[ \int 1,d\delta_n=1. ]

The distinction records the departure of mass toward infinity, which compactly supported test functions cannot detect.

Random variables and mappings

If (X_n) and (X) are random variables taking values in (S), then (X_n) converges in distribution to (X) precisely when their laws satisfy

[ \mathcal L(X_n)\Rightarrow\mathcal L(X). ]

This notion is weaker than convergence in probability because it compares only the induced measures and does not require the variables to be realized on the same probability space.

The continuous mapping theorem expresses the stability of weak convergence under transformations. If (X_n) converges in distribution to (X), and if a measurable map (g:S\to T) is continuous at every point outside a set of (\mathcal L(X))-measure zero, then

[ g(X_n)\Rightarrow g(X). ]

In measure notation, the corresponding statement concerns pushforward measures:

[ \mu_n\Rightarrow\mu \quad\Longrightarrow\quad g_#\mu_n\Rightarrow g_#\mu ]

under the same almost-everywhere continuity condition. The result follows by composing bounded continuous test functions on (T) with (g), together with the continuity-set form of the Portmanteau theorem when (g) is not continuous everywhere.

Weak convergence does not by itself imply convergence of expectations for arbitrary unbounded functions. If (f) is continuous but grows without bound, mass of vanishing probability may occur at increasingly distant points and make (\int f,d\mu_n) fail to converge. Convergence of such integrals is obtained under additional control expressed through uniform integrability or suitable moment bounds.

Transform methods

On (\mathbb R^d), weak convergence can be characterized by characteristic functions. For a probability measure (\mu), its characteristic function is

[ \widehat{\mu}(t)

\int_{\mathbb R^d}e^{i\langle t,x\rangle},d\mu(x). ]

Lévy’s continuity theorem states that (\mu_n\Rightarrow\mu) implies

[ \widehat{\mu_n}(t)\longrightarrow\widehat{\mu}(t) ]

for every (t\in\mathbb R^d). Conversely, if the characteristic functions converge pointwise to a function continuous at the origin, then the limit is the characteristic function of a probability measure and the measures converge weakly to that measure.

The continuity requirement at the origin excludes limits associated with loss of total mass. It performs, in transform language, the same compactness role that tightness performs in geometric language. Transform methods are therefore especially compatible with limit theorems involving sums of independent random variables, including the central limit theorem.

Topological structure

For a Polish space (S), weak convergence is generated by a metrizable topology on (\mathcal P(S)). One compatible metric is the Prokhorov metric, defined by

[ d_{\mathrm P}(\mu,\nu)

\inf\left{ \varepsilon>0: \mu(A)\leq\nu(A^\varepsilon)+\varepsilon \text{ and } \nu(A)\leq\mu(A^\varepsilon)+\varepsilon \text{ for every Borel }A \right}, ]

where (A^\varepsilon) is the open (\varepsilon)-neighborhood of (A). Another compatible metric is the bounded-Lipschitz metric,

[ d_{\mathrm{BL}}(\mu,\nu)

\sup\left{ \left|\int f,d\mu-\int f,d\nu\right|: |f|_\infty\leq 1,\ \operatorname{Lip}(f)\leq 1 \right}. ]

Both metrics encode the same convergent sequences as integration against all bounded continuous functions. Their formulas differ because the Prokhorov metric compares the mass of neighborhoods, whereas the bounded-Lipschitz metric compares the action of measures on a controlled function class.

Under these metrics, (\mathcal P(S)) is itself a Polish space whenever (S) is Polish. This fact permits probability laws to be treated as random elements and supports weak-convergence methods for empirical measures, stochastic processes, and measure-valued random variables.

See also