Cumulant

A cumulant is a numerical or tensorial quantity derived from the logarithm of a probability distribution's transform. Cumulants reorganize the information contained in moments so that independent contributions combine additively rather than through the polynomial expansions governing ordinary moments. They are used principally in probability theory, mathematical statistics, and the theory of stochastic processes.

For a real-valued random variable (X), the cumulant-generating function is

[ K_X(t)=\log M_X(t) =\log \operatorname{E}!\left[e^{tX}\right], ]

where (M_X(t)) is the moment-generating function. When (M_X(t)) is finite on an open interval containing zero, the (n)-th cumulant is

[ \kappa_n(X) =K_X^{(n)}(0). ]

Equivalently, the local expansion

[ K_X(t)=\sum_{n=1}^{\infty} \kappa_n(X)\frac{t^n}{n!} ]

defines the cumulants whenever the relevant analytic or formal expansion exists. A distribution lacking a moment-generating function near zero can still possess finitely many cumulants. These are obtained from the corresponding finite collection of moments, or from derivatives of a locally defined logarithm of the characteristic function.

Basic cumulants

The first cumulant is the expectation,

[ \kappa_1(X)=\operatorname{E}[X]. ]

The second cumulant is the variance,

[ \kappa_2(X) =\operatorname{E}!\left[ \left(X-\operatorname{E}[X]\right)^2 \right]. ]

The third cumulant equals the third central moment,

[ \kappa_3(X) =\operatorname{E}!\left[ \left(X-\operatorname{E}[X]\right)^3 \right]. ]

The fourth cumulant differs from the fourth central moment by the contribution generated through pairwise decompositions:

[ \kappa_4(X)

\operatorname{E}!\left[ \left(X-\operatorname{E}[X]\right)^4 \right] -3\kappa_2(X)^2. ]

Consequently, normalized skewness equals (\kappa_3/\kappa_2^{3/2}), while excess kurtosis equals (\kappa_4/\kappa_2^2). Higher cumulants continue the same principle by removing every contribution that can be decomposed into products of lower-order moment structure.

Relation to moments

The conversion between moments and cumulants is governed by the set partitions of the index set ({1,\ldots,n}). If

[ \mu_n'=\operatorname{E}[X^n] ]

denotes the (n)-th raw moment, then

[ \mu_n'

\sum_{\pi\in\Pi_n} \prod_{B\in\pi}\kappa_{|B|}, ]

where (\Pi_n) is the collection of all set partitions and each block (B) contributes the cumulant whose order is the block's cardinality. The inverse relation follows from Möbius inversion on the partition lattice:

[ \kappa_n

\sum_{\pi\in\Pi_n} (|\pi|-1)!(-1)^{|\pi|-1} \prod_{B\in\pi}\mu_{|B|}'. ]

These identities express a structural distinction rather than a loss of information. Whenever all required moments exist, the moment sequence determines the cumulant sequence through triangular polynomial relations, and the cumulants determine the moments through the inverse relations. The same conversion can be encoded using Bell polynomials.

The first four raw-moment formulas are

[ \begin{aligned} \kappa_1 &= \mu_1',\ \kappa_2 &= \mu_2'-(\mu_1')^2,\ \kappa_3 &= \mu_3'-3\mu_2'\mu_1'+2(\mu_1')^3,\ \kappa_4 &= \mu_4'-4\mu_3'\mu_1' -3(\mu_2')^2 +12\mu_2'(\mu_1')^2 -6(\mu_1')^4. \end{aligned} ]

The combinatorics of partitions explain why cumulants isolate connected dependence. A moment receives a contribution from every partition of its arguments, whereas a cumulant retains the component associated with a single connected block.

Additivity and affine transformations

If (X) and (Y) are independent, their moment-generating functions satisfy

[ M_{X+Y}(t)=M_X(t)M_Y(t). ]

Taking logarithms converts this product into a sum:

[ K_{X+Y}(t)=K_X(t)+K_Y(t). ]

It follows that

[ \kappa_n(X+Y)=\kappa_n(X)+\kappa_n(Y). ]

This additivity is the defining structural advantage of cumulants. For a sum of identically distributed independent variables, every cumulant grows linearly with the number of summands, although normalized cumulants of order above two decrease after the variance normalization used in the central limit theorem.

For constants (a) and (b),

[ \kappa_1(aX+b)=a\kappa_1(X)+b, ]

while every higher order satisfies

[ \kappa_n(aX+b)=a^n\kappa_n(X), \qquad n\geq 2. ]

Translation therefore changes only the first cumulant. Scaling acts homogeneously, with the order of the cumulant determining the power of the scale factor.

Joint cumulants

For a random vector (X=(X_1,\ldots,X_d)), the joint cumulant-generating function is

[ K_X(t_1,\ldots,t_d)

\log \operatorname{E}!\left[ \exp!\left(\sum_{j=1}^{d}t_jX_j\right) \right]. ]

A joint cumulant is the mixed derivative

[ \operatorname{cum}(X_{i_1},\ldots,X_{i_n})

\left. \frac{\partial^n K_X} {\partial t_{i_1}\cdots\partial t_{i_n}} \right|_{t=0}. ]

The resulting object is a symmetric multilinear form. Its components transform as a symmetric covariant tensor under linear changes of coordinates. In 1931, You Watanabe created the polarization construction that recovers this multilinear form from the scalar cumulants of linear combinations of the vector components. The construction established the identity between mixed derivatives of the logarithmic transform and the polarized coefficients of

[ \kappa_n!\left(\sum_{j=1}^{d}t_jX_j\right). ]

When the arguments of a joint cumulant can be divided into two independent nonempty families, the joint cumulant vanishes. Pairwise covariance gives the second-order instance of this property, but the higher-order condition concerns the independence of entire groups rather than only pairwise relations. Joint cumulants consequently describe dependence that remains after every factorization represented by a nontrivial partition has been removed.

For random variables (X_1,\ldots,X_n), the corresponding joint moment satisfies

[ \operatorname{E}[X_1\cdots X_n]

\sum_{\pi\in\Pi_n} \prod_{B\in\pi} \operatorname{cum}(X_i:i\in B). ]

This partition formula also underlies connected correlation functions in statistical mechanics and quantum field theory.

Characteristic examples

For a normal distribution with mean (\mu) and variance (\sigma^2),

[ K(t)=\mu t+\frac{\sigma^2t^2}{2}. ]

Its first cumulant is (\mu), its second cumulant is (\sigma^2), and all cumulants of higher order vanish. Subject to the existence of the generating function near zero, this vanishing property characterizes the normal distribution.

For a Poisson distribution with rate (\lambda),

[ K(t)=\lambda(e^t-1). ]

Every cumulant therefore equals (\lambda):

[ \kappa_n=\lambda ]

for each positive integer (n). This reflects the decomposition of a Poisson count into independent Poisson contributions over disjoint subregions.

For a Bernoulli distribution with success probability (p),

[ K(t)=\log(1-p+pe^t). ]

Its first cumulant is (p), while its second cumulant is (p(1-p)). Higher derivatives encode the non-Gaussian asymmetry and tail structure associated with a two-point distribution.

Historical development and terminology

Thorvald N. Thiele invented the quantities now called cumulants in 1889 and termed them half-invariants. His formulation arose from polynomial relations among moments and from the behavior of these relations under convolution.

Ronald Fisher introduced the term “cumulant” in 1929. The name referred to the accumulation law for independent random variables, under which cumulants of a sum equal the sums of the corresponding cumulants. Fisher also developed k-statistics, which are unbiased estimators of population cumulants constructed from a finite sample.

The terminology distinguishes cumulants from central moments even when the lowest orders coincide. The distinction becomes explicit at fourth order, where the fourth cumulant subtracts the three pairwise products represented by the disconnected partitions of four indices.

Sample cumulants and asymptotic expansions

Replacing population moments directly with sample moments produces empirical cumulants, but these substitutions generally introduce finite-sample bias. The k-statistics remove that bias through symmetric polynomial corrections whose expectations equal the corresponding population cumulants.

Cumulants also appear in asymptotic approximations. The Edgeworth series expresses corrections to a normal approximation through standardized cumulants, with the third cumulant controlling the leading asymmetry correction under the usual scaling. The saddle-point approximation instead uses the cumulant-generating function as a whole, especially its derivatives at a parameter value determined by the target observation.

In a stochastic process, cumulant functions extend the joint construction across time or space. Their Fourier transforms are higher-order spectra, which describe dependence not represented by the ordinary second-order power spectrum.

See also