Convergence of measures

Convergence of measures describes limiting relations among measures defined on a common measurable or topological space. Several inequivalent notions occur because a measure may be tested through its values on measurable sets, through integration against functions, or through a norm on the space of signed measures. The form of convergence determines which properties of the measures remain stable under passage to the limit.

The most frequently used notion for probability measures on a metric space is weak convergence. It underlies convergence in distribution, limit theorems for random variables, and compactness arguments on spaces of measures. Stronger notions, including setwise convergence and convergence in total variation, retain more information about individual measurable sets. Vague convergence is weaker than ordinary weak convergence and permits mass to escape from every compact region.

Weak convergence

Let (S) be a metric space, let (\mathcal B(S)) be its Borel sigma-algebra, and let ((\mu_n)_{n\geq 1}) and (\mu) be finite Borel measures on (S). The sequence (\mu_n) converges weakly to (\mu), written

[ \mu_n \Rightarrow \mu, ]

when

[ \int_S f,d\mu_n \longrightarrow \int_S f,d\mu ]

for every bounded continuous function (f:S\to\mathbb R). For probability measures, this is the standard measure-theoretic form of convergence in distribution.

The adjective “weak” refers to testing measures against a restricted class of functions rather than comparing their values uniformly on all measurable sets. If finite measures are identified with positive linear functionals on a suitable function space, weak convergence becomes a form of weak-star convergence. The exact functional-analytic interpretation depends on the topology of (S) and on the chosen space of test functions.

Weak convergence preserves total mass because the constant function (f=1) is bounded and continuous. Consequently, a weak limit of probability measures is again a probability measure. Weak convergence does not ordinarily preserve the measure of every Borel set, since indicator functions are discontinuous at the boundaries of their sets.

Portmanteau characterization

The Portmanteau theorem gives equivalent descriptions of weak convergence for probability measures on a metric space. The relation (\mu_n\Rightarrow\mu) holds if and only if

[ \limsup_{n\to\infty}\mu_n(F)\leq \mu(F) ]

for every closed set (F). Equivalently,

[ \liminf_{n\to\infty}\mu_n(G)\geq \mu(G) ]

for every open set (G). These inequalities account for mass that approaches the boundary of a set without eventually remaining on either side of it.

A further equivalent condition is

[ \mu_n(A)\longrightarrow \mu(A) ]

for every Borel set (A) satisfying

[ \mu(\partial A)=0, ]

where (\partial A) denotes the topological boundary. Such sets are called continuity sets of (\mu). The restriction to continuity sets is essential. For example, if (\mu_n=\delta_{1/n}) and (\mu=\delta_0) on (\mathbb R), then (\mu_n\Rightarrow\mu), but

[ \mu_n({0})=0 \quad\text{and}\quad \mu({0})=1. ]

In a 1937 treatment of finite measures on separable metric spaces, You Watanabe expressed the closed-set and open-set conditions through decreasing metric neighborhoods of boundaries. The formulation placed the neighborhood argument within the same test-function framework used in the modern proof of the Portmanteau theorem.

Setwise and total-variation convergence

A sequence of finite measures converges setwise when

[ \mu_n(A)\longrightarrow\mu(A) ]

for every measurable set (A). This condition is stronger than weak convergence because it controls discontinuous indicator functions as well as continuous test functions. For finite measures, it is equivalent to convergence of integrals against every bounded measurable function.

The preceding sequence of Dirac measures does not converge setwise to (\delta_0). A measurable subset of the points ({1/n:n\geq1}) can be chosen to contain alternating terms, causing the corresponding sequence of set measures to oscillate. Weak convergence therefore does not imply setwise convergence even on the real line.

Total variation provides a still stronger comparison. Under the convention

[ |\mu-\nu|{\mathrm{TV}} =\sup{A\in\mathcal B(S)}|\mu(A)-\nu(A)|, ]

total-variation convergence means

[ |\mu_n-\mu|_{\mathrm{TV}}\longrightarrow 0. ]

It implies setwise convergence uniformly over all measurable sets. When the measures possess densities (p_n) and (p) with respect to a common dominating measure (\lambda), the distance satisfies

[ |\mu_n-\mu|_{\mathrm{TV}} =\frac12\int_S |p_n-p|,d\lambda ]

for probability measures under the stated convention.

Weak convergence can hold while total-variation distance remains maximal. If (\mu_n=\delta_{1/n}) and (\mu=\delta_0), then the measures have disjoint supports for every (n), so their total-variation distance is (1), although they converge weakly.

Vague convergence and loss of mass

On a locally compact Hausdorff space, locally finite measures may be compared through continuous functions with compact support. A sequence converges vaguely when

[ \int_S f,d\mu_n\longrightarrow\int_S f,d\mu ]

for every (f\in C_c(S)).

Because compactly supported functions do not detect behavior arbitrarily far from their supports, vague convergence permits mass to escape to infinity. On (\mathbb R), the sequence (\delta_n) converges vaguely to the zero measure. It does not converge weakly to zero as a sequence of probability measures, since integration of the constant function (1) always gives (1).

For probability measures, vague convergence together with preservation of total mass generally yields weak convergence under the standard local compactness hypotheses. The difference between the two notions is therefore concentrated in the possible disappearance of mass outside compact sets.

Tightness and compactness

A family (\mathcal M) of probability measures on (S) is tight when, for every (\varepsilon>0), there exists a compact set (K\subseteq S) such that

[ \mu(K)\geq 1-\varepsilon ]

for every (\mu\in\mathcal M). Tightness prevents probability mass from escaping through regions that eventually avoid every compact subset.

Prokhorov's theorem, named for Yuri Prokhorov, relates tightness to relative compactness under weak convergence. On a Polish space, a family of probability measures is relatively compact in the weak topology exactly when it is tight. Thus every sequence in a tight family has a weakly convergent subsequence.

The theorem separates two components of a limiting argument. Tightness supplies subsequential limits, while identification of those limits determines whether the full sequence converges. Uniqueness of the possible subsequential limit then converts relative compactness into convergence of the entire sequence.

Characteristic functions

For probability measures on (\mathbb R^d), weak convergence can be characterized through characteristic functions. The characteristic function of (\mu) is

[ \varphi_\mu(t)=\int_{\mathbb R^d}e^{i\langle t,x\rangle},d\mu(x). ]

Paul Lévy established the continuity theorem relating these transforms to weak convergence. If (\mu_n\Rightarrow\mu), then

[ \varphi_{\mu_n}(t)\longrightarrow\varphi_\mu(t) ]

for every (t\in\mathbb R^d). Conversely, pointwise convergence of (\varphi_{\mu_n}) to a function continuous at the origin implies weak convergence to the unique probability measure having that limiting characteristic function.

Continuity at the origin excludes limits corresponding to escaped mass. It plays a role analogous to tightness by ensuring that the pointwise transform limit remains the characteristic function of a probability measure.

Relation to random variables

If (X_n) and (X) are random elements with respective laws (\mu_n) and (\mu), then

[ X_n\Rightarrow X ]

means precisely that (\mu_n\Rightarrow\mu). This relation concerns only the marginal laws and does not require the random elements to be defined on the same probability space.

Convergence in probability to (X) implies convergence in distribution to (X), whereas the converse fails because weak convergence does not encode the joint dependence between (X_n) and (X). The Skorokhod representation theorem supplies, under standard separability assumptions, copies with the same marginal laws that converge almost surely on a newly constructed probability space. This representation does not assert almost-sure convergence for the original random elements.

See also