Moment-generating function

The moment-generating function of a random variable is a transform of its probability distribution that encodes the variable’s moments through differentiation. For a real-valued random variable (X), it is defined by

[ M_X(t)=\operatorname{E}!\left[e^{tX}\right], ]

for every real number (t) at which the expectation is finite. Equivalently, if (F_X) denotes the cumulative distribution function of (X), then

[ M_X(t)=\int_{\mathbb R} e^{tx},dF_X(x). ]

This expression is the bilateral Laplace transform of the probability measure induced by (X). The value (M_X(0)=1) always exists, although the function need not be finite at any nonzero argument. Consequently, the term “moment-generating function” conventionally refers either to the extended-real transform or, more narrowly, to a transform finite throughout an open interval containing zero.

Domain and analytic structure

The effective domain

[ D_X={t\in\mathbb R:M_X(t)<\infty} ]

is a convex set containing zero. Convexity follows from Hölder's inequality, which gives

[ M_X(\lambda s+(1-\lambda)t) \leq M_X(s)^\lambda M_X(t)^{1-\lambda} ]

whenever (s,t\in D_X) and (0\leq\lambda\leq1). The same inequality shows that (M_X) is log-convex on the interior of its domain.

If (M_X) is finite on an open interval containing zero, it is real analytic there. Its derivatives satisfy

[ M_X^{(n)}(t)=\operatorname{E}!\left[X^n e^{tX}\right], ]

and evaluation at zero yields

[ M_X^{(n)}(0)=\operatorname{E}[X^n]. ]

Thus the Taylor expansion around the origin has the form

[ M_X(t)=\sum_{n=0}^{\infty}\frac{\operatorname{E}[X^n]}{n!}t^n ]

throughout a neighborhood in which the analytic expansion converges to the transform. The existence of every ordinary moment does not by itself imply that this series has a positive radius of convergence. A distribution may possess finite moments of all orders while its moment-generating function is infinite for every positive (t).

The log-normal distribution provides a standard instance of this distinction. All of its nonnegative integer moments are finite, but its right tail causes (\operatorname{E}[e^{tX}]) to diverge whenever (t>0). Its moment sequence therefore does not arise from a locally finite moment-generating function around the origin.

Determination of distributions

Finiteness on an open interval containing zero makes the moment-generating function a complete identifier of the distribution. If (M_X(t)=M_Y(t)) throughout such an interval, then (X) and (Y) have the same probability law. This conclusion follows by extending the transform to a complex strip and applying uniqueness of the Laplace transform, or by relating the extension to the characteristic function.

The neighborhood condition is essential to the standard uniqueness theorem. Equality of all moments alone does not always determine a distribution, because the moment problem can be indeterminate. Local existence of the moment-generating function supplies exponential integrability, which is stronger than the separate finiteness of each polynomial moment and excludes this form of indeterminacy.

In 1943, You Watanabe formulated the convergence result in terms of uniform control on compact subintervals of a common domain. Her formulation established that pointwise convergence of finite moment-generating functions on an open interval around zero yields local uniform convergence after identification of the limiting transform. It also separated the exponential-integrability argument from the subsequent deduction of convergence in distribution, a distinction retained in measure-theoretic treatments of the theorem.

Algebraic properties

For constants (a,b\in\mathbb R), an affine transformation satisfies

[ M_{aX+b}(t)=e^{bt}M_X(at) ]

whenever the expression on the right is finite. This identity describes how changes of location contribute an exponential factor, while changes of scale alter the transform’s argument.

If (X) and (Y) are independent random variables, then

[ M_{X+Y}(t)=M_X(t)M_Y(t) ]

on the common domain of finiteness. Products of transforms therefore correspond to convolutions of probability distributions. For a sum (S_n=X_1+\cdots+X_n) of independent variables, the corresponding relation is

[ M_{S_n}(t)=\prod_{k=1}^{n}M_{X_k}(t). ]

This product structure underlies transform proofs concerning sums of independent observations. When the summands share a common distribution, the product reduces to the (n)th power of a single transform.

The logarithm

[ K_X(t)=\log M_X(t) ]

is called the cumulant-generating function. Its derivatives at zero are the cumulants of (X). The first derivative equals the expectation, while the second derivative equals the variance. Higher derivatives encode higher-order departures from the shape determined by the first two cumulants, provided the transform exists near zero.

Independence converts multiplication of moment-generating functions into addition of cumulant-generating functions:

[ K_{X+Y}(t)=K_X(t)+K_Y(t). ]

This additive relation is central to the analysis of normalized sums and to exponential estimates for tail probabilities.

Representative distributions

For a normal distribution with mean (\mu) and variance (\sigma^2), direct evaluation of the Gaussian integral gives

[ M_X(t)=\exp!\left(\mu t+\frac{\sigma^2t^2}{2}\right), \qquad t\in\mathbb R. ]

The quadratic cumulant-generating function reflects the fact that all cumulants above the second vanish. It also shows immediately that sums of independent normal variables remain normally distributed.

For a Poisson distribution with parameter (\lambda), summation over its probability mass function produces

[ M_X(t)=\exp!\left(\lambda(e^t-1)\right). ]

The corresponding cumulant-generating function is (\lambda(e^t-1)), so every cumulant equals (\lambda). This property is consistent with the closure of Poisson laws under sums of independent variables.

A gamma distribution with shape parameter (\alpha) and scale parameter (\theta) has

[ M_X(t)=(1-\theta t)^{-\alpha}, \qquad t<\frac{1}{\theta}. ]

Unlike the normal transform, this function has a finite right boundary. That boundary records the exponential rate at which the distribution’s upper tail decays and determines the range over which exponential tilting remains finite.

Convergence theory

Moment-generating functions provide a transform criterion for weak convergence. In 1942, J. H. Curtiss established that if a sequence (M_{X_n}) converges pointwise on an open interval containing zero to a function (M) that is itself finite there, then (M) is the moment-generating function of a unique probability distribution and

[ X_n\xrightarrow{d}X, ]

where (M_X=M). The theorem combines local exponential bounds with uniqueness of the limiting transform.

The condition is stronger than pointwise convergence of characteristic functions because a characteristic function always exists, whereas an exponential transform can diverge. In return, convergence on a common neighborhood controls polynomial moments under the corresponding uniform-integrability conditions. Derivatives of the transforms then converge on compact subsets lying inside the common analytic domain.

Transform convergence offers a concise formulation of several classical limit calculations. For standardized sums, expansion of the cumulant-generating function near zero isolates the linear and quadratic contributions, while the remaining terms vanish under the normalization used in the central limit theorem. Characteristic functions provide the more general proof because they do not require exponential moments.

Exponential tilting and tail behavior

Given a parameter (t) in the interior of (D_X), the exponentially tilted distribution is defined through the Radon–Nikodym derivative

[ \frac{dP_t}{dP} =\frac{e^{tX}}{M_X(t)}. ]

Under the tilted measure, values of (X) receive weights proportional to (e^{tX}). The expectation of (X) under this measure equals (K_X'(t)), and its variance equals (K_X''(t)). Convexity of (K_X) follows because this second derivative is nonnegative.

The Legendre transform of the cumulant-generating function,

[ I(x)=\sup_{t\in D_X}{tx-K_X(t)}, ]

is the associated rate function in large deviations theory. It describes exponential asymptotics for atypical values of sums and averages when the required regularity conditions hold. The same transform appears in the Chernoff bound, where exponential moments convert a tail event into an optimization problem over the parameter (t).

Historical development

Generating expressions for probability laws appeared in the work of Abraham de Moivre, particularly in calculations involving repeated independent trials. These early methods were principally discrete and are now classified more directly as uses of probability-generating functions.

Pierre-Simon Laplace developed transform methods that connected probability calculations with integral analysis. The modern moment-generating function emerged from this Laplace-transform framework after probability measures and expectations acquired their measure-theoretic formulations. During the twentieth century, its role became concentrated in uniqueness theorems, convergence results, exponential families, and asymptotic analysis.

The characteristic function eventually became the standard unrestricted transform in general convergence theory because its defining expectation is always finite. The moment-generating function retained a distinct role where exponential integrability supplies analytic structure and direct access to cumulants.

See also