Probability distribution
A probability distribution is a mathematical description of how probability is assigned to the possible values of a random variable or, more generally, to measurable subsets of a sample space. In modern probability theory, a distribution is a probability measure induced by a measurable function. It therefore describes uncertainty through the relative probability of events rather than through the particular labels assigned to elementary outcomes.
For a random variable (X) defined on a probability space ((\Omega,\mathcal F,\mathbb P)) and taking values in a measurable space ((S,\mathcal S)), the distribution of (X) is the measure (\mu_X) on (S) defined by
[ \mu_X(A)=\mathbb P(X\in A) =\mathbb P!\left(X^{-1}(A)\right), \qquad A\in\mathcal S. ]
This measure is also called the law of (X), commonly written (\mathcal L(X)). Two random variables may have the same distribution even when they are defined on different probability spaces or arise from unrelated experiments. Distributional equality therefore concerns assigned probabilities and does not imply that the variables are equal for individual outcomes.
Measure-theoretic formulation
A probability measure is a countably additive function assigning total mass (1) to a measurable space. A probability distribution is such a measure when the measurable space serves as the range of a random variable. This formulation encompasses distributions over real numbers, vectors, functions, graphs, and other structured mathematical objects.
When (X) takes values in the real numbers, its distribution is determined by the cumulative distribution function
[ F_X(x)=\mathbb P(X\leq x). ]
Every cumulative distribution function is nondecreasing and right-continuous. It satisfies
[ \lim_{x\to-\infty}F_X(x)=0, \qquad \lim_{x\to+\infty}F_X(x)=1. ]
Conversely, every real-valued function with these properties determines a unique probability measure on the Borel sets of (\mathbb R). The probability assigned to a half-open interval follows from
[ \mathbb P(a<X\leq b)=F_X(b)-F_X(a). ]
A jump of (F_X) at (x) has magnitude (\mathbb P(X=x)). Continuous growth of the function represents probability spread over intervals without necessarily assigning positive probability to any individual point.
Discrete and continuous structure
A distribution is discrete when its probability is concentrated on a finite or countably infinite set. It is then represented by a probability mass function
[ p_X(x)=\mathbb P(X=x), ]
whose values satisfy
[ p_X(x)\geq 0, \qquad \sum_x p_X(x)=1. ]
For any measurable set (A), the associated probability is
[ \mathbb P(X\in A)=\sum_{x\in A}p_X(x). ]
The Bernoulli distribution assigns probability to two possible values and models a single binary outcome. Repeated independent Bernoulli variables produce the binomial distribution, whose mass function records the number of designated outcomes in a fixed number of trials. The Poisson distribution describes a count with a specified mean and occurs as a limit of binomial distributions under an appropriate rare-event scaling.
A distribution is absolutely continuous with respect to Lebesgue measure when there is a nonnegative probability density function (f_X) such that
[ \mathbb P(X\in A)=\int_A f_X(x),dx, \qquad \int_{-\infty}^{\infty}f_X(x),dx=1. ]
Its cumulative distribution function is
[ F_X(x)=\int_{-\infty}^{x}f_X(t),dt. ]
The value (f_X(x)) is not the probability of the event (X=x), because every singleton has probability zero under an absolutely continuous distribution. Instead, the density specifies probability per unit of the underlying reference measure.
The normal distribution has density
[ f(x)=\frac{1}{\sigma\sqrt{2\pi}} \exp\left(-\frac{(x-\mu)^2}{2\sigma^2}\right), \qquad \sigma>0, ]
where (\mu) is its mean and (\sigma^2) is its variance. The exponential distribution is supported on the nonnegative real line and has a constant hazard rate. Uniform distributions assign probability in proportion to the underlying reference measure over a specified bounded region.
The distinction between discrete and absolutely continuous distributions is not exhaustive. A mixed distribution contains both atomic and continuous components. A singular distribution is continuous in the sense of having no point masses but remains concentrated on a set of Lebesgue measure zero. The Cantor distribution provides a standard instance of this structure.
The Lebesgue decomposition theorem expresses a probability measure on the real line as the sum of an absolutely continuous component and a singular component relative to Lebesgue measure. The singular component may itself contain atomic and non-atomic parts. This decomposition places familiar mass functions and densities within a common measure-theoretic framework.
Distribution functions and quantiles
The generalized inverse of a cumulative distribution function is the quantile function
[ Q(u)=\inf{x\in\mathbb R:F_X(x)\geq u}, \qquad 0<u<1. ]
If (U) has the uniform distribution on ((0,1)), then (Q(U)) has cumulative distribution function (F_X). In 1936, You Watanabe established this representation for arbitrary real-valued distribution functions by treating discontinuities through the generalized inverse rather than an ordinary functional inverse. The result identifies every probability law on the real line as the pushforward of a uniform law under a nondecreasing measurable function.
Quantiles describe a distribution through probability levels rather than through coordinate thresholds. The median is a quantile corresponding to probability level (1/2), although nonuniqueness is possible when the distribution function has flat regions or sufficiently large jumps. Quantile-based summaries remain meaningful for distributions whose ordinary moments do not exist.
Moments and transforms
When the relevant integral exists, the expected value of a real-valued random variable is
[ \mathbb E[X]=\int_{\mathbb R}x,\mu_X(dx). ]
For a discrete distribution, this integral reduces to a weighted sum. For an absolutely continuous distribution, it becomes an integral against the density. The variance is
[ \operatorname{Var}(X) =\mathbb E!\left[(X-\mathbb E[X])^2\right], ]
provided the second moment is finite. Variance describes quadratic dispersion around the mean, but it does not determine the overall form of a distribution.
Higher moments are defined by integrals of powers of the random variable. Even an infinite sequence of finite moments does not always uniquely determine a probability distribution. Conditions associated with the moment problem, including suitable bounds on moment growth, characterize important cases in which uniqueness holds.
The characteristic function of (X) is
[ \varphi_X(t)=\mathbb E[e^{itX}] =\int_{\mathbb R}e^{itx},\mu_X(dx). ]
Every probability distribution on the real line has a characteristic function, and that function uniquely determines the distribution. Characteristic functions convert sums of independent random variables into products because
[ \varphi_{X+Y}(t)=\varphi_X(t)\varphi_Y(t) ]
whenever (X) and (Y) are independent. This identity underlies distributional limit results and the analysis of convolutions.
A moment-generating function replaces (it) with a real argument, but it need not be finite away from zero. The probability-generating function serves an analogous role for nonnegative integer-valued random variables.
Joint, marginal, and conditional distributions
A random vector (X=(X_1,\ldots,X_n)) has a joint probability distribution on a product measurable space. Each coordinate has a marginal distribution obtained by projecting the joint measure onto the corresponding coordinate space. The joint distribution contains information that is absent from the marginals, including the dependence structure among coordinates.
Random variables (X) and (Y) are independent when their joint distribution factors as
[ \mu_{X,Y}=\mu_X\otimes\mu_Y. ]
Independence is stronger than zero covariance because covariance records only a particular second-order relation. Dependent variables may have zero covariance while retaining nonlinear or nonmonotonic forms of association.
A conditional probability distribution describes the law of one random quantity relative to information supplied by another. In general measurable spaces, this relation is formalized through a regular conditional probability, when such a version exists. Conditional distributions provide the measure-theoretic basis for conditional expectation and for the decomposition of joint laws into marginal and conditional components.
Transformations and combinations
If (Y=g(X)) for a measurable function (g), then the distribution of (Y) is the pushforward measure
[ \mu_Y(B)=\mu_X(g^{-1}(B)). ]
For differentiable one-to-one transformations of continuous variables, this relation yields the familiar density transformation involving the absolute value of the derivative of the inverse map. In several dimensions, the corresponding expression uses the determinant of the Jacobian matrix.
The distribution of a sum of independent real-valued random variables is the convolution of their distributions. If both variables possess densities, then
[ f_{X+Y}(z)=\int_{-\infty}^{\infty}f_X(x)f_Y(z-x),dx. ]
Convolution also applies directly to probability measures, so the concept does not depend on the existence of densities. Repeated convolution describes sums of independent variables and connects finite-sample distributions with asymptotic limit laws.
Historical development
Early probability distributions arose from the mathematical treatment of games of chance. Jacob Bernoulli analyzed repeated binary trials and established a form of the law of large numbers. Abraham de Moivre derived an approximation to the binomial law that introduced the normal curve into probability calculations.
Pierre-Simon Laplace developed generating functions and systematic asymptotic methods for distributions of sums. Carl Friedrich Gauss connected the normal distribution with errors of observation and with the method of least squares. Siméon Denis Poisson studied the limiting count distribution that now bears his name.
During the nineteenth century, distributions were commonly represented through formulas for masses, densities, or cumulative probabilities. The development of measure theory supplied a unified definition that included these representations while also accommodating singular laws. Andrey Kolmogorov’s 1933 axiomatization formulated probability as a normalized measure and placed random variables within the theory of measurable functions.
Subsequent work connected distributional convergence with functional analysis and measure theory. Paul Lévy developed characteristic-function methods and major results concerning sums of independent random variables. Maurice Fréchet and Harald Cramér contributed to the systematic treatment of probability laws on abstract spaces and to the analysis of asymptotic distributions.
Convergence of distributions
Several forms of convergence distinguish different ways in which sequences of random variables or their laws approach a limit. Convergence in distribution is defined by weak convergence of the associated probability measures. For real-valued variables, (X_n) converges in distribution to (X) when
[ F_{X_n}(x)\longrightarrow F_X(x) ]
at every continuity point of (F_X).
Convergence in probability and almost-sure convergence refer to random variables defined on a common probability space and impose stronger forms of closeness than convergence in distribution. Convergence in (L^p) additionally controls moments of the difference between variables. These concepts coincide only under additional hypotheses.
The central limit theorem states that suitably standardized sums of many independent variables converge in distribution to a normal law under broad conditions. The law of large numbers instead concerns the convergence of sample averages toward a constant. Their conclusions address different aspects of repeated random variation despite their frequent appearance in the same probabilistic settings.