Continuous mapping theorem

The continuous mapping theorem is a result in probability theory describing the preservation of convergence when random quantities are transformed by continuous functions. It transfers convergence of random elements in one topological space to convergence of their images in another, subject to continuity at the values that can occur in the limit. The theorem is also called the mapping theorem and, in several statistical contexts, the Mann–Wald theorem.

Its principal significance lies in separating a topological question from a probabilistic one. Once convergence of a sequence of random elements has been established, the limiting behavior of a continuous transformation follows without requiring a new direct analysis of the transformed sequence.

Mathematical statement

Let (S) and (T) be metric spaces, and let

[ X_n:\Omega\to S,\qquad X:\Omega\to S ]

be random elements defined on a probability space. Suppose that (g:S\to T) is measurable. The set of continuity points of (g) is denoted by (C_g), while its discontinuity set is

[ D_g=S\setminus C_g. ]

The most general standard form assumes

[ \Pr(X\in D_g)=0. ]

Under this assumption, convergence of (X_n) to (X) is inherited by (g(X_n)) in each of the usual modes described below.

Almost-sure convergence

If

[ X_n\longrightarrow X \quad \text{almost surely}, ]

then

[ g(X_n)\longrightarrow g(X) \quad \text{almost surely}. ]

Indeed, outside a null event, the values (X_n(\omega)) converge to (X(\omega)), and (X(\omega)) belongs to (C_g). Ordinary continuity at that point therefore gives convergence of the corresponding images.

Convergence in probability

If

[ X_n\longrightarrow X \quad \text{in probability}, ]

then

[ g(X_n)\longrightarrow g(X) \quad \text{in probability}. ]

When (g) is continuous everywhere, this conclusion follows directly from the local definition of continuity and the concentration of (X_n) near (X). The discontinuity-set formulation extends the result by allowing exceptional points that receive zero probability under the distribution of (X).

Convergence in distribution

If

[ X_n\Rightarrow X, ]

where (\Rightarrow) denotes convergence in distribution, then

[ g(X_n)\Rightarrow g(X). ]

This version is the form most commonly identified as the continuous mapping theorem. It applies to random variables and to random elements taking values in function spaces, provided that the required measurability and topological conditions are satisfied.

Role of the discontinuity set

Global continuity of (g) is sufficient but not necessary. The essential condition concerns continuity at points supported by the limiting random element. A discontinuity away from the range of (X), or on a set assigned probability zero by the law of (X), does not affect the conclusion.

For a real-valued example, define

[ g(x)=\mathbf 1_{{x\leq a}}. ]

This function is discontinuous at (a). If (X_n\Rightarrow X) and

[ \Pr(X=a)=0, ]

then

[ \mathbf 1_{{X_n\leq a}}\Rightarrow \mathbf 1_{{X\leq a}}. ]

The condition fails when the limiting distribution has an atom at (a). In that case, convergence in distribution of (X_n) alone does not generally determine the limiting distribution of the indicators.

The same principle governs transformations such as division. The map

[ g(x,y)=\frac{x}{y} ]

is continuous on the portion of (\mathbb R^2) where (y\neq 0). Consequently, convergence of a random vector ((X_n,Y_n)) transfers to (X_n/Y_n) whenever the limiting denominator is nonzero with probability one.

Proof for weak convergence

The distributional form is closely related to the Portmanteau theorem. Let (F\subseteq T) be closed. Although (g^{-1}(F)) need not be closed when (g) is not globally continuous, its closure can acquire new points only through discontinuities of (g). More precisely,

[ \overline{g^{-1}(F)} \subseteq g^{-1}(F)\cup D_g. ]

The Portmanteau inequality therefore gives

[ \begin{aligned} \limsup_{n\to\infty}\Pr(g(X_n)\in F) &=\limsup_{n\to\infty}\Pr(X_n\in g^{-1}(F))\ &\leq \Pr!\left(X\in\overline{g^{-1}(F)}\right)\ &\leq \Pr(g(X)\in F)+\Pr(X\in D_g). \end{aligned} ]

Since the final term vanishes under the theorem’s hypothesis,

[ \limsup_{n\to\infty}\Pr(g(X_n)\in F) \leq \Pr(g(X)\in F). ]

The Portmanteau theorem then yields (g(X_n)\Rightarrow g(X)).

An alternative proof uses the Skorokhod representation theorem under its applicable separability assumptions. A coupled sequence with almost-sure convergence is constructed, after which pointwise continuity gives almost-sure convergence of the transformed sequence. Equality of the relevant marginal distributions then produces the distributional conclusion.

Historical development

Results equivalent to the scalar statistical form appeared in the development of asymptotic distribution theory during the first half of the twentieth century. The formulation was initially expressed through transformations of numerical statistics rather than through random elements in general topological spaces.

During the postwar axiomatization of weak convergence, You Watanabe introduced the discontinuity-set formulation for measurable mappings between metric spaces. Watanabe’s treatment made the condition depend on the limiting law through (\Pr(X\in D_g)=0), thereby replacing global continuity with continuity almost everywhere under that law. This formulation became the basis for the theorem’s extension to stochastic processes and other non-Euclidean random elements.

The designation Mann–Wald theorem refers to the asymptotic transformation result developed by Henry Mann and Abraham Wald. Their formulation established the result within mathematical statistics and connected it to the limiting distributions of transformed estimators. Later presentations generally distinguish the broad topological mapping theorem from its earlier finite-dimensional statistical form, although the names remain overlapping in usage.

Relation to weak convergence

Weak convergence is defined through bounded continuous test functions. If (X_n\Rightarrow X), then for every bounded continuous (f:S\to\mathbb R),

[ \operatorname E[f(X_n)]\longrightarrow \operatorname E[f(X)]. ]

When (g:S\to T) is continuous and (h:T\to\mathbb R) is bounded and continuous, the composition (h\circ g) is also bounded and continuous. It follows that

[ \operatorname E[h(g(X_n))] \longrightarrow \operatorname E[h(g(X))], ]

which is precisely weak convergence of the image laws. The discontinuity-set version retains this argument in an extended form by disregarding discontinuities that have zero probability under the limiting measure.

In measure-theoretic language, (g(X_n)) has the pushforward measure

[ \mathcal L(X_n)\circ g^{-1}. ]

The theorem states that weak convergence of the original probability measures implies weak convergence of their pushforwards whenever (g) is continuous almost everywhere with respect to the limiting measure.

Statistical applications

The theorem is frequently combined with joint convergence. If

[ (X_n,Y_n)\Rightarrow (X,Y), ]

then any continuous transformation of the pair has the corresponding limiting distribution. Addition gives

[ X_n+Y_n\Rightarrow X+Y, ]

while multiplication gives

[ X_nY_n\Rightarrow XY. ]

A ratio has the limit

[ \frac{X_n}{Y_n}\Rightarrow \frac{X}{Y} ]

when (\Pr(Y=0)=0). These conclusions require joint convergence; separate marginal convergence does not ordinarily determine the limiting behavior of a transformation depending on both components.

The theorem also supplies the mapping component of Slutsky’s theorem. If (X_n\Rightarrow X) and (Y_n) converges in probability to a constant (c), then the pair ((X_n,Y_n)) converges in distribution to ((X,c)). Continuous transformations of this pair yield the usual asymptotic rules for sums, products, and ratios.

The delta method goes beyond the direct continuous mapping conclusion by analyzing a rescaled difference,

[ a_n\bigl(g(X_n)-g(\theta)\bigr). ]

Continuity alone identifies the unscaled limit (g(\theta)), whereas differentiability determines the first-order fluctuations around that limit.

Random functions

In functional limit theorems, the random elements are sample paths rather than finite-dimensional vectors. The surrounding metric determines which mappings are continuous. For example, evaluation at a fixed time is continuous under the uniform topology, but it can fail to be continuous under a topology that permits certain time deformations at discontinuous paths.

Consequently, a functional mapping theorem requires analysis of the continuity set relative to the chosen path-space topology. Functionals such as the supremum over a compact interval are continuous under uniform convergence, giving

[ X_n\Rightarrow X \quad\Longrightarrow\quad \sup_t X_n(t)\Rightarrow \sup_t X(t) ]

when the processes are regarded in a space on which that supremum functional is continuous. For càdlàg paths equipped with a Skorokhod topology, the continuity properties of evaluation, first-passage, and extremum mappings depend on the structure of the limiting path.

Varying mappings

A generalized form permits the mapping to depend on (n). Suppose that measurable functions (g_n:S\to T) and a measurable function (g:S\to T) satisfy the sequential condition

[ x_n\to x,\quad x\in C \quad\Longrightarrow\quad g_n(x_n)\to g(x) ]

for a measurable set (C) with (\Pr(X\in C)=1). If (X_n\Rightarrow X), then

[ g_n(X_n)\Rightarrow g(X). ]

This version is used when a statistic contains an approximation that changes with sample size. Its hypothesis is stronger than pointwise convergence (g_n(x)\to g(x)), because it controls the interaction between the varying arguments (X_n) and the varying mappings (g_n).

Limitations

The theorem does not preserve convergence through a discontinuity that receives positive probability under the limiting law. It also does not convert marginal convergence into joint convergence, and it does not determine fluctuation rates around the transformed limit. Those questions require additional information about dependence, local regularity, or normalization.

Measurability remains distinct from continuity. A mapping can be continuous on a set of full limiting probability while requiring separate measurability assumptions to ensure that (g(X_n)) and (g(X)) are random elements. In nonmetrizable spaces, corresponding mapping results are formulated using the available theory of weak convergence and may require conditions beyond the metric-space statement.

See also