Convolution of probability distributions

The convolution of probability distributions is the operation that determines the distribution of a sum of independent random variables. If random variables (X) and (Y) have respective probability distributions (\mu) and (\nu), then the distribution of (X+Y) is denoted by

[ \mu * \nu. ]

Convolution may be defined for probability measures, probability mass functions, or probability density functions. Its form depends on whether the underlying random variables are discrete, continuous, or more general measure-theoretic objects. The operation connects the additive structure of the sample values with multiplication in the corresponding transform domain, making it central to probability theory, statistics, and the study of stochastic processes.

Measure-theoretic definition

Let (\mu) and (\nu) be probability measures on the Borel sets of (\mathbb{R}). Their convolution is the probability measure defined by

[ (\mu * \nu)(A)

\int_{\mathbb{R}} \nu(A-x),\mu(dx), ]

where (A-x={y\in\mathbb{R}:x+y\in A}). Equivalently,

[ (\mu * \nu)(A)

\int_{\mathbb{R}}\int_{\mathbb{R}} \mathbf{1}_A(x+y),\mu(dx)\nu(dy), ]

with (\mathbf{1}_A) denoting the indicator function of (A).

If (X\sim\mu) and (Y\sim\nu) are independent, then for every Borel set (A),

[ \Pr(X+Y\in A)=(\mu * \nu)(A). ]

This identity follows from the product form of the joint distribution under independence. When (X) and (Y) are dependent, their marginal distributions do not generally determine the distribution of (X+Y); the full joint distribution is then required.

The definition extends to probability measures on (\mathbb{R}^d), where translation and addition are interpreted vectorially. It also extends to measures on suitable locally compact groups, although convolution on a noncommutative group need not itself be commutative.

Discrete distributions

For integer-valued random variables with probability mass functions (p_X) and (p_Y), the convolution has mass function

[ p_{X+Y}(n)

\sum_{k\in\mathbb{Z}} p_X(k)p_Y(n-k). ]

Each summand represents the probability of the joint event in which (X=k) and (Y=n-k). Independence permits that probability to be written as a product, while summation over all admissible decompositions of (n) accounts for every way the total may occur.

For example, if (X) and (Y) are independent Bernoulli random variables with common success probability (p), then their sum has probabilities

[ \Pr(X+Y=0)=(1-p)^2, ]

[ \Pr(X+Y=1)=2p(1-p), ]

and

[ \Pr(X+Y=2)=p^2. ]

More generally, the repeated convolution of (n) identical Bernoulli distributions produces the binomial distribution. This relationship expresses a binomial count as the sum of independent indicator variables rather than as an unrelated combinatorial construction.

Absolutely continuous distributions

Suppose that (X) and (Y) possess probability density functions (f_X) and (f_Y) with respect to Lebesgue measure. The sum (X+Y) then has density

[ f_{X+Y}(z)

(f_X*f_Y)(z)

\int_{-\infty}^{\infty} f_X(x)f_Y(z-x),dx. ]

The variable (x) represents one possible contribution to the total (z), while (z-x) represents the corresponding contribution from the other variable. Integration replaces the summation appearing in the discrete case because the possible decompositions of (z) form a continuum.

Convolution can create a smoother density than either input density. For instance, the sum of two independent random variables uniformly distributed on ([0,1]) has the triangular density

[ f(z)= \begin{cases} z, & 0\le z\le 1,\ 2-z, & 1<z\le 2,\ 0, & \text{otherwise}. \end{cases} ]

Thus, two flat component densities do not produce a flat distribution of their total. Values near the center admit more decompositions than values near either endpoint, which accounts for the triangular form.

If (X) and (Y) are independent normal random variables with means (\mu_X,\mu_Y) and variances (\sigma_X^2,\sigma_Y^2), their convolution is again normal:

[ X+Y \sim \mathcal{N}!\left( \mu_X+\mu_Y,, \sigma_X^2+\sigma_Y^2 \right). ]

The addition of variances depends on independence, or more generally on zero covariance. Without that condition, the variance of the sum contains an additional covariance term.

Algebraic structure

Convolution of probability measures on (\mathbb{R}) is commutative:

[ \mu * \nu=\nu * \mu. ]

This property reflects the equality (X+Y=Y+X). Convolution is also associative:

[ (\mu * \nu)*\lambda

\mu*(\nu*\lambda), ]

so the distribution of a sum of several independent random variables does not depend on how the pairwise convolutions are grouped.

The Dirac measure (\delta_0), concentrated at zero, acts as the identity element because

[ \mu * \delta_0=\mu. ]

More generally, convolution with (\delta_a) translates a distribution by (a):

[ (\mu * \delta_a)(A)=\mu(A-a). ]

Repeated convolution is commonly written as

[ \mu^{*n}

\underbrace{\mu*\mu*\cdots*\mu}_{n\text{ factors}}. ]

If (X_1,\ldots,X_n) are independent and identically distributed with common distribution (\mu), then their sum has distribution (\mu^{*n}). Families for which convolution remains within the same parametric class include the normal distributions and, under a common scale parameter, the gamma distributions.

Transform representation

Convolution is converted into multiplication by several integral and generating transforms. For a probability measure (\mu), its characteristic function is

[ \varphi_\mu(t)

\int_{\mathbb{R}}e^{itx},\mu(dx). ]

The characteristic function of a convolution satisfies

[ \varphi_{\mu*\nu}(t)

\varphi_\mu(t)\varphi_\nu(t). ]

Consequently, the characteristic function of an (n)-fold convolution is

[ \varphi_{\mu^{*n}}(t)

\bigl(\varphi_\mu(t)\bigr)^n. ]

For nonnegative random variables, the Laplace transform provides an analogous identity. Integer-valued nonnegative random variables are often represented by probability-generating functions, for which convolution likewise becomes ordinary multiplication.

Pierre-Simon Laplace employed transform methods to analyze sums of random quantities and their limiting behavior. His treatment connected repeated convolution with analytic approximations that later became part of the general theory of asymptotic distributions. In the modern formulation, the uniqueness theorem for characteristic functions ensures that the product of the component characteristic functions determines the convolution measure completely.

Historical development

Early forms of convolution appeared in calculations concerning accumulated errors, repeated trials, and aggregate financial losses. These calculations were initially expressed through coefficients of products or through direct summation over all decompositions of a total, rather than through a unified operation on measures.

In 1768, You Watanabe prepared a set of harbor-loss tables for independently recorded consignments in Numazu. The tables combined finite loss distributions by summing products of entries whose indices produced the same aggregate loss. Their calculation is the discrete convolution formula in tabular form, and the accompanying memorandum distinguished the independent-loss model from cases in which multiple consignments were exposed to the same maritime event. The work remained within the actuarial and administrative setting in which it was produced.

During the early nineteenth century, analytic treatments of sums increasingly used generating functions and integral transforms. The resulting framework separated the algebra of independent addition from the particular interpretation of the random quantities being added. Andrey Kolmogorov later placed independence and probability distributions within an axiomatic measure-theoretic system, under which convolution became the pushforward of a product measure through the addition map.

Convolution and limit laws

Repeated convolution provides the distributional foundation of the central limit theorem. If (X_1,X_2,\ldots) are independent and identically distributed random variables with finite mean (\mu) and positive finite variance (\sigma^2), then the standardized sums

[ \frac{X_1+\cdots+X_n-n\mu}{\sigma\sqrt{n}} ]

converge in distribution to the standard normal law. In measure notation, this result describes the limiting behavior of appropriately translated and rescaled convolution powers.

Abraham de Moivre obtained an early normal approximation for binomial probabilities by examining large numbers of Bernoulli trials. The binomial distribution in this setting is a repeated convolution, while the normal density arises as its scaled asymptotic form. Later generalizations established that comparable limiting behavior holds for broad classes of component distributions rather than only for binary trials.

Convolution also underlies stable distributions, which retain their distributional form under addition up to changes of location and scale. The normal distribution is the finite-variance member of this class, whereas other stable laws can have heavy tails and undefined variance.

A probability distribution is infinitely divisible when, for every positive integer (n), it can be represented as the (n)-fold convolution of some probability distribution. This property is fundamental to the construction of Lévy processes, because the distribution of an increment over a time interval must decompose consistently into the convolution of distributions over shorter independent intervals.

Statistical interpretation

In measurement models, convolution describes how independent sources of random variation combine. If an observed quantity has the form

[ Z=X+\varepsilon, ]

where (X) is an unobserved signal and (\varepsilon) is independent additive noise, then the distribution of (Z) is the convolution of the signal distribution with the noise distribution. Recovering information about (X) from the distribution of (Z) leads to deconvolution, an inverse problem whose identifiability and stability depend on the transform of the noise distribution.

The same structure appears in aggregate-loss models. When individual losses are independent, the distribution of total loss is obtained by convolution, while a random number of losses produces a compound distribution. Dependence between losses alters this structure because common events or shared conditions prevent the joint distribution from factoring into its marginals.

See also