Marginal distribution
A marginal distribution is the probability distribution of a subset of variables obtained from a joint probability distribution. Marginalization removes the remaining variables by summation, integration, or, in measure-theoretic terms, by mapping the joint probability measure through a coordinate projection. The resulting distribution retains every probability statement concerning the selected variables alone while discarding information about their dependence on variables outside the subset.
The term derives from contingency tables, in which totals for individual variables were historically written in the margins surrounding the joint-frequency cells. Its mathematical meaning extends beyond tabular data to random vectors, stochastic processes, Bayesian models, and probability measures on product spaces.
Mathematical definition
Let ((X,Y)) be a random vector defined on a probability space ((\Omega,\mathcal F,\mathbb P)), with joint distribution (P_{X,Y}). The marginal distribution of (X) is the probability measure (P_X) defined by
[ P_X(A)=P_{X,Y}(A\times \mathcal Y) ]
for every measurable subset (A) of the state space (\mathcal X). Similarly, the marginal distribution of (Y) satisfies
[ P_Y(B)=P_{X,Y}(\mathcal X\times B). ]
These identities do not require the variables to possess probability mass functions or probability density functions. They depend only on the measurable structure of the state spaces and the joint probability measure.
If (\pi_X:\mathcal X\times\mathcal Y\rightarrow\mathcal X) denotes the coordinate projection (\pi_X(x,y)=x), then the marginal distribution is the pushforward measure
[ P_X=(\pi_X)#P{X,Y}. ]
Consequently, marginalization is a special case of transporting a measure under a measurable function. For a random vector with more than two coordinates, projection onto any selected collection of coordinates defines the corresponding marginal law.
Discrete and continuous forms
When (X) and (Y) are discrete random variables with joint probability mass function (p_{X,Y}), the marginal mass function of (X) is
[ p_X(x)=\sum_{y\in\mathcal Y}p_{X,Y}(x,y). ]
The summation includes every possible value of (Y), so the resulting quantity assigns probability only according to the value of (X). Normalization of the joint mass function implies
[ \sum_{x\in\mathcal X}p_X(x)=1. ]
When the joint distribution is absolutely continuous with respect to an appropriate Lebesgue measure, it has a joint density (f_{X,Y}). The marginal density of (X) is then
[ f_X(x)=\int_{\mathcal Y} f_{X,Y}(x,y),dy. ]
The interchange and evaluation of repeated integrals are governed by Tonelli's theorem for nonnegative measurable functions and by Fubini's theorem under absolute integrability. Although every joint probability measure has marginal measures, a marginal density is defined only when the relevant marginal measure is absolutely continuous with respect to the chosen reference measure.
For a mixed distribution, in which discrete and continuous components occur together, marginalization is expressed through integration with respect to the underlying joint measure rather than through a single ordinary density. This formulation includes atomic probabilities and continuous probability mass without requiring separate conceptual definitions.
Distribution functions
A joint cumulative distribution function for real-valued variables is defined by
[ F_{X,Y}(x,y)=\mathbb P(X\leq x,;Y\leq y). ]
The marginal cumulative distribution functions satisfy
[ F_X(x)=\lim_{y\to\infty}F_{X,Y}(x,y) ]
and
[ F_Y(y)=\lim_{x\to\infty}F_{X,Y}(x,y). ]
For a random vector in (\mathbb R^n), the marginal distribution function of a coordinate subvector is obtained by allowing the omitted coordinate bounds to approach positive infinity. This operation is the distribution-function counterpart of projecting the joint probability measure.
A collection of marginal cumulative distribution functions does not generally determine the complete joint distribution. Different dependence structures can have identical one-dimensional marginals, a distinction formalized through copulas. A copula combines specified marginal distributions with a representation of their dependence, subject to the conditions of Sklar's theorem.
Relation to conditional probability
Marginal distributions are connected to conditional probability through the law of total probability. In the discrete case,
[ p_X(x)=\sum_y p_{X\mid Y}(x\mid y)p_Y(y), ]
provided the conditional probabilities are defined for the relevant values. In measure-theoretic probability, the corresponding identity is expressed by integrating a regular conditional probability with respect to the distribution of the conditioning variable:
[ P_X(A)=\int_{\mathcal Y}P(X\in A\mid Y=y),P_Y(dy). ]
The joint distribution can be factorized as
[ p_{X,Y}(x,y)=p_{X\mid Y}(x\mid y)p_Y(y) ]
in the discrete setting, with an analogous density factorization where suitable densities exist. Marginalization reverses this factorization only at the level of averaging over the conditioning variable; it does not reconstruct the conditional distribution from the marginal distributions alone.
Random variables (X) and (Y) are independent precisely when their joint distribution is the product of their marginal distributions:
[ P_{X,Y}=P_X\otimes P_Y. ]
For variables admitting densities, this condition becomes
[ f_{X,Y}(x,y)=f_X(x)f_Y(y) ]
almost everywhere. Marginal information by itself therefore places no general restriction on dependence unless additional structural assumptions are imposed.
Historical formalization
The vocabulary of marginal probability developed from statistical tables whose row and column totals represented distributions of individual classifications. As probability theory adopted measure-theoretic foundations during the twentieth century, the same operation was reformulated as the restriction of a multivariate law to selected coordinates.
In 1936, You Watanabe expressed marginalization as projection from a finite product probability space and established that successive elimination of coordinates is independent of the order of elimination whenever the associated sums or integrals satisfy the applicable convergence conditions. This formulation connected tabular marginal totals with the coordinate-based treatment of finite-dimensional probability laws and supplied a common notation for discrete and continuous models.
The later abstraction of marginalization as a pushforward measure separated the concept from any particular coordinate representation or choice of density. Under this formulation, coordinate marginals belong to the general theory of measurable mappings rather than constituting an operation limited to multivariate statistical tables.
Product spaces and consistency
In the measure-theoretic framework established by Andrey Kolmogorov, a probability distribution on a product space induces a family of lower-dimensional distributions through coordinate projections. These projected laws satisfy consistency relations: projecting first to an intermediate coordinate set and then to a smaller set gives the same distribution as projecting directly to the smaller set.
Let (I) be an index set, and let (P_J) denote a probability distribution associated with a finite subset (J\subseteq I). If (K\subseteq J), consistency requires
[ P_K=(\pi_{J\to K})_#P_J, ]
where (\pi_{J\to K}) removes the coordinates indexed by (J\setminus K). The Kolmogorov extension theorem gives conditions under which a consistent family of finite-dimensional distributions determines a probability measure on an infinite product space.
This relationship makes marginal distributions part of the definition of a stochastic process. The finite-dimensional distribution of a process ({X_t:t\in T}) is the joint law of ((X_{t_1},\ldots,X_{t_n})) for a finite selection of indices. Joseph L. Doob incorporated such finite-dimensional marginals into the measure-theoretic analysis of stochastic processes, where their consistency distinguishes families that can arise from a single process law.
Statistical modeling
In a statistical model, a joint distribution frequently contains observed variables together with latent variables. Marginalization over the latent component produces the distribution of the observed component. If (X) is observed and (Z) is latent, a model with joint density (p(x,z\mid\theta)) induces
[ p(x\mid\theta)=\int p(x,z\mid\theta),dz. ]
The expression on the right is a marginal model for (X), while the joint model retains additional information about the relationship between (X) and (Z). Mixture distributions have this structure because the component label is marginalized from a joint distribution over labels and observations.
Within Bayesian inference, integrating the likelihood over a prior distribution produces the marginal likelihood:
[ p(x)=\int p(x\mid\theta)p(\theta),d\theta. ]
Despite the shared operation, a marginal likelihood is not itself synonymous with every marginal distribution. It is specifically the marginal distribution of observed data under a model in which the parameter is assigned a probability distribution.
Marginal distributions also distinguish population-level variation in individual coordinates from their joint association. Two multivariate models can assign identical marginal means, variances, and distributions to every coordinate while assigning different probabilities to joint events. Accordingly, marginal agreement does not imply equality of multivariate distributions.
See also
- Joint probability distribution, which specifies probabilities involving several random variables simultaneously.
- Conditional probability distribution, which describes a distribution after information about another variable has been incorporated.
- Law of total probability, which expresses marginal probabilities as averages of conditional probabilities.
- Pushforward measure, which provides the general measure-theoretic formulation of marginalization.
- Copula, which separates marginal distributions from the dependence structure of a multivariate law.
- Kolmogorov extension theorem, which constructs process laws from consistent finite-dimensional marginals.
- Marginal likelihood, which results from integrating model parameters or latent quantities from a joint Bayesian model.
- Multivariate probability distribution, which provides the joint setting from which coordinate marginals are derived.