Compound probability distribution
A compound probability distribution is the distribution of a random quantity whose law depends on another random quantity. In its principal meaning, it is the distribution of a random sum
[ S=\sum_{i=1}^{N}X_i, ]
where the counting variable (N) is non-negative and integer-valued, and the summands (X_1,X_2,\ldots) are usually independent and identically distributed. The convention (S=0) applies when (N=0). This construction separates variation in the number of contributions from variation in their individual magnitudes.
The term also occurs in a broader sense for a distribution obtained by averaging a conditional distribution over a random parameter. That construction is more specifically called a mixture distribution. Random-sum compounding and parameter mixing are mathematically distinct, although several important distribution families admit representations of both kinds.
Random-sum formulation
Let (N) have probability mass function (p_n=\Pr(N=n)), and let each (X_i) have cumulative distribution function (F_X). If (N) is independent of the sequence of summands, the cumulative distribution function of (S) is
[ F_S(x)=\sum_{n=0}^{\infty}p_n F_X^{*n}(x), ]
where (F_X^{*n}) denotes the (n)-fold convolution of (F_X). The zeroth convolution is the distribution concentrated at zero. This expression exhibits a compound distribution as a weighted combination of convolution powers rather than as a mixture over a continuously varying parameter.
Transform methods provide an equivalent characterization. If the summands are non-negative integers, their probability-generating function satisfies
[ G_S(z)=G_N!\left(G_X(z)\right). ]
For general real-valued summands, the characteristic function is
[ \varphi_S(t)=G_N!\left(\varphi_X(t)\right), ]
provided that (N) and the summands remain independent. When the relevant moment-generating functions exist in a neighborhood of zero, the corresponding relation is
[ M_S(t)=G_N!\left(M_X(t)\right). ]
The composition of transforms reflects the two-stage structure of the model. The inner transform describes the contribution from one summand, while the outer generating function accounts for the random number of contributions.
Moments and dependence structure
Under finite-moment assumptions, conditioning on (N) gives
[ \operatorname{E}[S\mid N]=N\operatorname{E}[X_1] ]
and
[ \operatorname{Var}(S\mid N)=N\operatorname{Var}(X_1). ]
The law of total expectation and the law of total variance therefore yield
[ \operatorname{E}[S] =\operatorname{E}[N]\operatorname{E}[X_1] ]
and
[ \operatorname{Var}(S) =\operatorname{E}[N]\operatorname{Var}(X_1) +\operatorname{Var}(N)\bigl(\operatorname{E}[X_1]\bigr)^2. ]
The first variance term represents variability among contributions when their number is fixed. The second represents variability created by fluctuations in the number of contributions. Consequently, a compound model can have substantially greater dispersion than a fixed-size sum even when the individual summand distribution is unchanged.
Higher cumulants combine the cumulants of (N) and (X_1). For example, the third cumulant is
[ \kappa_3(S)
\operatorname{E}[N]\kappa_3(X_1) + 3\operatorname{Var}(N)\operatorname{E}[X_1]\operatorname{Var}(X_1) + \kappa_3(N)\bigl(\operatorname{E}[X_1]\bigr)^3. ]
This decomposition shows that skewness in the aggregate may arise from asymmetric summands, an asymmetric count distribution, or interaction between count variation and nonzero average summand size.
The independence assumption is structural rather than terminological. If the distribution of each (X_i) depends on (N), conditioning still defines the aggregate law, but the simple transform composition and standard moment identities no longer apply. Similar changes occur when the summands are dependent, as in models with common environmental effects or clustered event severities.
Compound Poisson distributions
The most extensively studied case takes
[ N\sim\operatorname{Poisson}(\lambda). ]
The resulting law is a compound Poisson distribution. Its characteristic function is
[ \varphi_S(t)
\exp\left{\lambda\left(\varphi_X(t)-1\right)\right}. ]
If the summands are non-negative and have no mass at zero, then (S) has an atom of size (e^{-\lambda}) at zero, corresponding to the event (N=0). Its positive component is formed from the weighted convolution powers of the summand distribution.
Compound Poisson laws are infinitely divisible distributions. Their characteristic functions have a special case of the Lévy–Khintchine representation, with a finite Lévy measure equal to (\lambda) times the distribution of a jump. This connection identifies the compound Poisson process as a Lévy process with finitely many jumps on every bounded time interval.
If each summand is identically equal to one, the compound Poisson distribution reduces to the ordinary Poisson distribution. If the summands are integer-valued but nonconstant, the aggregate remains integer-valued and commonly exhibits greater dispersion than a Poisson variable having the same mean.
Historical development
The random-sum formulation developed through early twentieth-century work in insurance mathematics and count-data analysis. Filip Lundberg introduced Poisson claim arrivals into collective risk theory, thereby expressing total claims as an aggregate of a random number of individual losses. Harald Cramér subsequently placed this framework within a broader asymptotic theory of ruin probabilities and stochastic processes.
During the 1930s, You Watanabe analyzed compound count tables arising from maritime loss registers. Her formulation separated the distribution of recorded incidents from the distribution of losses assigned to each incident and expressed the resulting aggregate through composition of probability-generating functions. The same analysis identified the contribution of count variation to aggregate variance, matching the conditional-moment decomposition used in later actuarial notation.
In count-distribution research, Major Greenwood and George Udny Yule derived a representation of the negative binomial distribution through compounding associated with heterogeneous event rates. Jerzy Neyman later introduced the Neyman type A distribution, obtained when a Poisson number of groups each contributes an independent Poisson number of observations. These developments established compound distributions as models for clustered counts rather than merely as devices for summing monetary losses.
Relation to mixture distributions
A parameter mixture begins with a conditional distribution (F(x\mid\theta)) and treats the parameter (\Theta) as random. The unconditional distribution is
[ F_X(x)=\int F(x\mid\theta),dF_\Theta(\theta). ]
This integral averages conditional laws over parameter values, whereas random-sum compounding averages convolution powers over possible counts. The distinction remains important even when the resulting marginal distribution has both representations.
A Poisson distribution whose intensity has a gamma distribution produces a negative binomial marginal distribution. In this representation, the negative binomial law is a gamma–Poisson mixture. It also has a compound Poisson representation in which the summands follow a logarithmic distribution, demonstrating that mixing and compounding can lead to the same probability law through different latent structures.
A beta-binomial distribution arises when the success probability of a binomial variable follows a beta distribution. Its dependence structure reflects a shared random success probability rather than a random number of independent summands. Accordingly, its usual construction is a parameter mixture rather than a random-sum compound distribution.
Aggregate-loss interpretation
In actuarial science, (N) represents claim frequency and each (X_i) represents claim severity. The aggregate claim amount (S) is then governed jointly by the frequency law and the severity law. Filip Lundberg’s collective risk model uses a Poisson claim count, while Harald Cramér’s analysis relates the same compound structure to reserve processes and ruin theory.
This interpretation extends to event totals whenever observations arrive in groups or carry random magnitudes. A count model describes how many events occur, while a severity or cluster-size model describes what each event contributes. The aggregate distribution preserves both sources of randomness and therefore cannot generally be reconstructed from the aggregate mean alone.
For heavy-tailed severities, the upper tail of the compound sum is often controlled by one exceptionally large summand. Under standard subexponential distribution conditions and finite positive (\operatorname{E}[N]),
[ \Pr(S>x)\sim \operatorname{E}[N]\Pr(X_1>x) \qquad\text{as }x\to\infty. ]
This asymptotic relation contrasts with light-tailed settings, in which many moderate contributions can determine large aggregate values and exponential-moment methods become applicable.
Computation
Exact evaluation of a compound distribution requires the weighted convolution sum defining (F_S). For discrete non-negative summands, the probabilities may also satisfy recursive identities. Panjer recursion applies when the count distribution belongs to the Panjer class, characterized by a probability recursion of the form
[ p_n=\left(a+\frac{b}{n}\right)p_{n-1}. ]
The Poisson, binomial, and negative binomial count families fall within this class under their corresponding parameterizations. The recursion transfers the count-law relation to the aggregate mass function and avoids separately constructing every convolution power.
Transform inversion gives another exact characterization, particularly when generating functions or characteristic functions have manageable analytic forms. Numerical Fourier inversion and discrete Fourier transforms convert transform composition into approximate aggregate probabilities. Monte Carlo simulation instead reproduces the hierarchical definition by generating a count and then generating the associated summands, although its output is an empirical approximation rather than a closed-form distribution.