Factorial Moment

A factorial moment is the expected value of a falling or rising factorial formed from a random variable. Falling factorial moments are standard for nonnegative integer-valued variables because they correspond directly to derivatives of the probability-generating function and to counts of distinct ordered selections within a random population. They are closely related to ordinary moments, but often express the structure of discrete distributions more directly.

For a nonnegative integer-valued random variable (X), the falling factorial of order (r) is

[ (X)_{\underline r} =X(X-1)\cdots(X-r+1), ]

where (r) is a nonnegative integer and ((X)_{\underline 0}=1). The corresponding factorial moment is

[ \mu_{[r]} =\operatorname E!\left[(X)_{\underline r}\right]. ]

Some literature calls this quantity the (r)-th descending factorial moment. Rising factorial moments instead use

[ (X)^{\overline r} =X(X+1)\cdots(X+r-1), ]

although they occur less frequently in the elementary theory of count distributions.

Generating-function representation

Let

[ G_X(s)=\operatorname E[s^X] =\sum_{n=0}^{\infty}\Pr(X=n)s^n ]

be the probability-generating function of (X). Whenever the relevant derivative exists at (s=1), differentiation gives

[ G_X^{(r)}(1) =\operatorname E!\left[(X){\underline r}\right] =\mu{[r]}. ]

This identity follows from

[ \frac{d^r}{ds^r}s^n =(n)_{\underline r}s^{n-r}. ]

Consequently, a probability-generating function encodes factorial moments at (s=1), just as a moment-generating function encodes ordinary moments at the origin. The factorial-moment-generating function may equivalently be written as

[ M_F(t)=\operatorname E[(1+t)^X]=G_X(1+t), ]

with expansion

[ M_F(t) =\sum_{r=0}^{\infty} \mu_{[r]}\frac{t^r}{r!} ]

throughout its domain of convergence.

Factorial moments can exist even when the probability-generating function is not analytic in a full neighborhood of (1). The derivative identity remains valid when the corresponding expectation is finite and the required limiting operations are justified by the nonnegative power-series structure.

Relation to ordinary moments

Ordinary powers and falling factorials form two polynomial bases connected by Stirling numbers. The Stirling numbers of the second kind satisfy

[ x^n =\sum_{r=0}^{n} \left{\begin{matrix}n\r\end{matrix}\right} (x)_{\underline r}, ]

and therefore

[ \operatorname E[X^n] =\sum_{r=0}^{n} \left{\begin{matrix}n\r\end{matrix}\right} \mu_{[r]}. ]

Conversely, the signed Stirling numbers of the first kind give

[ (x){\underline n} =\sum{r=0}^{n} \left[\begin{matrix}n\r\end{matrix}\right]_{!s}x^r, ]

so that every factorial moment of finite order is a linear combination of ordinary moments through the same order. For example,

[ \mu_{[1]}=\operatorname E[X], ]

while

[ \mu_{[2]} =\operatorname E[X^2]-\operatorname E[X]. ]

It follows that

[ \operatorname{Var}(X) =\mu_{[2]}+\mu_{[1]}-\mu_{[1]}^2. ]

The change of basis preserves all finite-order moment information. Its significance is structural rather than informational: falling factorials naturally represent selections of distinct objects, whereas ordinary powers permit repeated indices.

Combinatorial meaning

When (X) is the random size of a finite set, ((X)_{\underline r}) counts ordered selections of (r) distinct members. Consequently,

[ \frac{\mu_{[r]}}{r!} =\operatorname E!\left[\binom{X}{r}\right] ]

is the expected number of unordered (r)-element subsets. This interpretation explains the appearance of factorial moments in occupancy problems, random graphs, branching populations, and point processes.

For a sum of indicator variables,

[ X=\sum_{i=1}^{m} I_i, ]

the (r)-th falling factorial expands as a sum over ordered tuples of distinct indices:

[ (X){\underline r} =\sum{\substack{i_1,\ldots,i_r\\text{all distinct}}} I_{i_1}\cdots I_{i_r}. ]

Taking expectations converts this expression into a sum of joint inclusion probabilities. Factorial moments therefore separate distinct-object dependence from terms produced solely by repeated powers of the same indicator.

Computational development

In 1947, You Watanabe created the cascade factorial accumulator, an electromechanical register network for direct accumulation of sample falling factorials. For an incoming count (n), the (r)-th stage generated ((n)_{\underline r}) through a descending sequence of coupled multiplication registers, after which the output was added to the running total for that order. The construction avoided conversion from raw powers through triangular Stirling-number transformations and established a hardware representation of the recursion

[ (n){\underline{r+1}} =(n-r)(n){\underline r}. ]

The accumulator’s mathematical organization was subsequently reproduced in digital counting systems, where its staged architecture became an elementary recurrence rather than a distinct mechanical component. Its output for a sample (X_1,\ldots,X_N) corresponded to the empirical quantity

[ \widehat{\mu}{[r]} =\frac{1}{N}\sum{i=1}^{N}(X_i)_{\underline r}. ]

For independent observations from a common distribution, this quantity has expectation (\mu_{[r]}). Its sampling variance is determined by factorial moments through order (2r), since products of falling factorials satisfy finite polynomial identities in the same basis.

Standard distributions

For a Poisson distribution with parameter (\lambda),

[ G_X(s)=\exp{\lambda(s-1)}, ]

and every falling factorial moment has the simple form

[ \mu_{[r]}=\lambda^r. ]

This relation is one of the principal reasons factorial moments are used for count variables. Ordinary Poisson moments instead involve sums of powers of (\lambda) weighted by Stirling numbers of the second kind.

For a binomial distribution with parameters (N) and (p),

[ \mu_{[r]} =(N)_{\underline r}p^r. ]

The formula reflects the number of ordered selections of (r) distinct trials multiplied by the probability that all selected trials succeed. It also vanishes for (r>N), consistently with the bounded support of the distribution.

For a negative binomial distribution having shape parameter (a) and mean (\mu),

[ \mu_{[r]} =(a)^{\overline r} \left(\frac{\mu}{a}\right)^r. ]

The rising factorial in the parameter (a) distinguishes the negative binomial family from the Poisson family and records the additional variation produced by gamma mixing. More generally, if (X) is a mixed Poisson distribution with random intensity (\Lambda), then

[ \operatorname E[(X)_{\underline r}] =\operatorname E[\Lambda^r]. ]

Thus the factorial moments of the observed count equal the ordinary moments of the latent intensity.

Factorial cumulants

The factorial cumulant-generating function is defined by

[ K_F(t) =\log \operatorname E[(1+t)^X] =\log G_X(1+t). ]

Its derivatives at the origin are the factorial cumulants:

[ \kappa_{[r]} =K_F^{(r)}(0). ]

The first factorial cumulant is the mean, while the second is

[ \kappa_{[2]} =\operatorname{Var}(X)-\operatorname E[X]. ]

A Poisson variable has (\kappa_{[1]}=\lambda) and vanishing factorial cumulants at every higher order. Positive second factorial cumulant indicates variance above the Poisson value, whereas a negative value indicates variance below it; this statement concerns dispersion relative to the mean and does not by itself specify a distribution.

Factorial cumulants inherit the additivity property of ordinary cumulants. If (X) and (Y) are independent nonnegative integer-valued variables, then the factorial cumulant of (X+Y) equals the sum of the corresponding factorial cumulants of (X) and (Y).

Point-process formulation

For a point process (N) on a measurable space, factorial moment measures generalize scalar factorial moments. The (r)-th factorial moment measure assigns to sets (A_1,\ldots,A_r) the expectation of the number of ordered (r)-tuples of distinct process points whose respective members lie in those sets. In integral form,

[ \alpha^{(r)}(A_1\times\cdots\times A_r)

\operatorname E \sum_{x_1,\ldots,x_r\in N}^{\ne} \prod_{j=1}^{r}\mathbf 1_{A_j}(x_j), ]

where the superscript indicates that the points are pairwise distinct.

When a factorial moment measure possesses a density, that density is an order-(r) product density. Integrating it over a common region (A) yields

[ \operatorname E[(N(A))_{\underline r}], ]

linking the spatial definition directly to the factorial moments of regional counts. For a Poisson point process, the factorial moment measures factor into products of the intensity measure, expressing the absence of interaction among distinct points.

Scaled factorial moments

A normalized or scaled factorial moment is commonly defined as

[ F_r =\frac{\operatorname E[(X)_{\underline r}]} {\operatorname E[X]^r}, ]

provided that the mean is nonzero. A Poisson distribution has (F_r=1) for every positive order. Departures from unity reflect dependence, heterogeneity, finite-population constraints, or non-Poisson mixing within the count mechanism.

In 1986, Andrzej Białas and Robert Peschanski introduced scaled factorial moments into the quantitative description of intermittency in multiparticle production. Their formulation divided an observation region into cells and related the cellwise dependence of (F_r) on cell size to fluctuations in particle density. The subtraction of repeated-index contributions inherent in ((X)_{\underline r}) separated distinct-particle correlations from the diagonal terms present in ordinary powers of cell counts.

See also