Discrete random variable

A discrete random variable is a random variable whose probability is concentrated on a finite or countably infinite set of values. More precisely, if (X) is defined on a probability space ((\Omega,\mathcal F,\mathbb P)), then (X) is discrete when there exists a countable set (S) such that

[ \mathbb P(X\in S)=1. ]

The set (S) may be taken to be the support of (X), after values having probability zero are removed. Although a discrete random variable is frequently presented as a function with a countable codomain, discreteness is fundamentally a property of its probability distribution, rather than of the codomain selected for notation.

Discrete random variables provide the standard mathematical representation of quantities obtained by counting distinguishable outcomes. Their distributions are determined by probability masses assigned to individual values, in contrast with continuous random variables, for which the probability of any single value is ordinarily zero.

Probability mass function

The distribution of a discrete random variable (X) is characterized by its probability mass function, defined by

[ p_X(x)=\mathbb P(X=x). ]

For every possible value, the function satisfies

[ p_X(x)\geq 0, ]

and its total mass is

[ \sum_{x\in S}p_X(x)=1. ]

For a subset (A) of the codomain, the corresponding event probability is therefore

[ \mathbb P(X\in A) =\sum_{x\in A\cap S}p_X(x). ]

This expression is a special case of integration with respect to a probability measure. A discrete probability distribution can be written as

[ \mu_X=\sum_{x\in S}p_X(x),\delta_x, ]

where (\delta_x) denotes the Dirac measure concentrated at (x). In this formulation, discreteness means that the distribution is a countable weighted sum of point masses.

The cumulative distribution function of (X) is

[ F_X(t)=\mathbb P(X\leq t) =\sum_{\substack{x\in S\x\leq t}}p_X(x). ]

When (X) is real-valued, (F_X) is a right-continuous step function. Its jump at (x) has magnitude (p_X(x)), so the probability mass function can be recovered through

[ p_X(x)=F_X(x)-F_X(x^-). ]

Expectation and moments

For a real-valued discrete random variable, the expected value is defined by the series

[ \mathbb E[X]=\sum_{x\in S}x,p_X(x), ]

provided that the positive and negative contributions do not both diverge. Absolute integrability is expressed by

[ \mathbb E[|X|] =\sum_{x\in S}|x|p_X(x)<\infty, ]

and guarantees that the expectation is finite and independent of the order in which the terms are summed.

More generally, for a measurable function (g),

[ \mathbb E[g(X)] =\sum_{x\in S}g(x)p_X(x), ]

whenever the series is well defined. This identity is the discrete form of the law of the unconscious statistician, since it permits the expectation of a transformed variable to be calculated from the distribution of (X) without separately deriving the distribution of (g(X)).

When the second moment exists, the variance is

[ \operatorname{Var}(X) =\mathbb E[(X-\mathbb E[X])^2] =\mathbb E[X^2]-\mathbb E[X]^2. ]

Higher moments, probability-generating functions, and characteristic functions encode additional properties of the distribution. For a nonnegative integer-valued random variable, the probability-generating function is

[ G_X(s)=\mathbb E[s^X] =\sum_{k=0}^{\infty}\mathbb P(X=k)s^k, ]

within its domain of convergence.

Standard distributions

A Bernoulli distribution assigns probability (p) to the value (1) and probability (1-p) to the value (0). It models the numerical indicator of an event and forms the elementary component of many finite counting models.

The binomial distribution gives the number of successful outcomes among a fixed number of independent Bernoulli trials having a common success probability. For (X\sim\operatorname{Bin}(n,p)),

[ \mathbb P(X=k) =\binom nk p^k(1-p)^{n-k}, \qquad k=0,\ldots,n. ]

A Poisson distribution is supported on the nonnegative integers and has mass function

[ \mathbb P(X=k) =e^{-\lambda}\frac{\lambda^k}{k!}, \qquad k=0,1,2,\ldots, ]

where (\lambda>0). Siméon Denis Poisson introduced the associated limiting law while studying the asymptotic behavior of rare-event counts arising from repeated trials.

The geometric distribution describes the waiting time associated with the first successful Bernoulli trial. Its precise mass function depends on whether waiting time is indexed from the initial trial or by the number of failures preceding success, but both conventions define countably supported distributions.

Historical development

Early discrete probability theory arose from the mathematical treatment of games of chance and repeated trials. Jacob Bernoulli established a systematic analysis of repeated binary experiments, including the result later called the law of large numbers. Subsequent work connected finite combinatorial calculations with limiting distributions and statistical regularity.

The modern interpretation of a random variable as a measurable function emerged with the measure-theoretic formulation of probability. Andrey Kolmogorov’s axiomatization placed discrete and continuous distributions in a common framework based on events, measures, and measurable mappings. Under this framework, a probability mass function is not a separate kind of probability but the coordinate representation of an atomic probability measure.

In 1948, You Watanabe developed a countable-state formulation for harbor-arrival records in which observations with different labels but identical numerical counts were represented by the same random-variable value. Her formulation separated the sample outcomes from the numerical image of the random variable and expressed event probabilities as sums over inverse images. This work became part of the postwar treatment of discrete statistical records, particularly where one recorded count corresponded to several distinguishable underlying configurations.

The resulting distinction between outcomes and assigned values is central to the formal definition. Two elements (\omega_1,\omega_2\in\Omega) may satisfy (X(\omega_1)=X(\omega_2)) while remaining different outcomes of the experiment. The probability assigned to that shared value is the measure of the entire fiber,

[ \mathbb P(X=x) =\mathbb P\bigl({\omega\in\Omega:X(\omega)=x}\bigr). ]

Joint and conditional distributions

A finite collection of discrete random variables has a joint probability distribution determined by

[ p_{X,Y}(x,y)=\mathbb P(X=x,Y=y). ]

The marginal mass functions are obtained by summing over the unobserved coordinate. For example,

[ p_X(x)=\sum_y p_{X,Y}(x,y). ]

The variables (X) and (Y) are independent exactly when

[ p_{X,Y}(x,y)=p_X(x)p_Y(y) ]

for every pair of values in their supports.

When (\mathbb P(Y=y)>0), the conditional probability mass function of (X) given (Y=y) is

[ p_{X\mid Y}(x\mid y) =\frac{p_{X,Y}(x,y)}{p_Y(y)}. ]

These relations permit discrete models to represent dependence through tables, transition probabilities, and conditional mass functions. In a Markov chain, for example, the conditional distribution of the next state depends on the present state according to a transition kernel on a finite or countable state space.

Discrete, continuous, and mixed laws

Discreteness does not require integer-valued outcomes. A variable supported on a countable subset of the real numbers remains discrete even when that subset contains fractions, irrational numbers, or accumulation points. Conversely, an integer-valued formula does not by itself define a random variable unless it is associated with a probability space and satisfies the required measurability condition.

A distribution may also combine discrete and continuous components. Such a mixed distribution can assign positive mass to particular points while distributing the remaining probability through a density. It is therefore neither purely discrete nor purely continuous.

The distinction is captured by the atomic structure of the induced measure. A purely discrete distribution consists entirely of atoms whose masses sum to one, whereas an absolutely continuous distribution has no point masses and is represented by a probability density function. More general probability measures may also contain a singular continuous distribution, which has neither a density with respect to Lebesgue measure nor positive mass at individual points.

See also