Probability density function
A probability density function, commonly abbreviated as PDF, is a nonnegative function that represents the distribution of a continuous random variable relative to a specified reference measure. For a real-valued random variable (X), the reference measure is ordinarily Lebesgue measure, and the probability that (X) lies in a measurable set (A) is
[ \Pr(X\in A)=\int_A f_X(x),dx, ]
where (f_X) denotes the density of (X). In particular, for an interval ([a,b]),
[ \Pr(a\leq X\leq b)=\int_a^b f_X(x),dx. ]
The value (f_X(x)) is not itself the probability that (X=x). For a distribution admitting a density with respect to Lebesgue measure, every individual point has probability zero, even when the density at that point is positive. Density instead measures probability per unit of the underlying coordinate, so its numerical value depends on the scale in which the variable is expressed.
Definition
Let ((\Omega,\mathcal F,\Pr)) be a probability space, and let (X:\Omega\to\mathbb R) be a measurable random variable. The distribution of (X) is the probability measure (\mu_X) on the Borel sets of (\mathbb R) defined by
[ \mu_X(A)=\Pr(X\in A). ]
A measurable function (f_X:\mathbb R\to[0,\infty]) is a probability density function of (X) when
[ \mu_X(A)=\int_A f_X(x),dx ]
for every Borel set (A). Such a function necessarily satisfies
[ f_X(x)\geq 0 ]
almost everywhere and
[ \int_{-\infty}^{\infty}f_X(x),dx=1. ]
Conversely, every nonnegative measurable function whose integral is one determines an absolutely continuous probability distribution.
A density is defined only up to equality almost everywhere. Altering its value on a set of Lebesgue measure zero does not change any probability calculated from it. Consequently, two pointwise distinct functions can represent exactly the same probability distribution.
Measure-theoretic formulation
The concept extends beyond densities on the real line. Let ((S,\mathcal S)) be a measurable space, let (\mu) be a probability measure on it, and let (\nu) be a sigma-finite measure. When (\mu) is absolutely continuous with respect to (\nu), written (\mu\ll\nu), the Radon–Nikodym theorem gives a measurable function satisfying
[ \mu(A)=\int_A \frac{d\mu}{d\nu},d\nu. ]
The derivative (d\mu/d\nu) is the density of (\mu) relative to (\nu). Johann Radon established the finite-measure form of the underlying theorem, while Otto Nikodym developed its general measure-theoretic formulation. Their work placed density functions within the broader theory of derivatives of measures.
Lebesgue measure is therefore not an intrinsic component of every probability density. A probability mass function is a density relative to counting measure, while a density on a surface can be defined relative to surface-area measure. The reference measure determines both the numerical values and the physical dimensions of the density.
Not every probability distribution has a density relative to Lebesgue measure. A discrete distribution assigns positive probability to individual points and is singular with respect to Lebesgue measure. Singular continuous distributions, including the distribution associated with the Cantor function, assign zero probability to every point but still lack a Lebesgue density. A general probability measure can contain absolutely continuous, discrete, and singular continuous components, as expressed by the Lebesgue decomposition theorem.
Relation to the cumulative distribution function
The cumulative distribution function of a real-valued random variable is
[ F_X(x)=\Pr(X\leq x). ]
When (X) has density (f_X),
[ F_X(x)=\int_{-\infty}^{x}f_X(t),dt. ]
The function (F_X) is then absolutely continuous, and the fundamental theorem of calculus gives
[ F_X'(x)=f_X(x) ]
for almost every (x). Pointwise differentiability is not required everywhere, because density functions may have discontinuities and may be modified on null sets without changing the distribution.
The converse also holds in measure-theoretic form: an absolutely continuous cumulative distribution function possesses an almost-everywhere derivative whose integral recovers the function. A cumulative distribution function with jumps represents point masses and therefore cannot be recovered solely by integrating an ordinary Lebesgue density.
Interpretation and dimensional dependence
For a small positive increment (\Delta x), the local probability near (x) satisfies
[ \Pr(x<X\leq x+\Delta x) =\int_x^{x+\Delta x}f_X(t),dt. ]
At a point where the density is continuous, this expression has the approximation
[ \Pr(x<X\leq x+\Delta x) =f_X(x)\Delta x+o(\Delta x). ]
Thus, density values may exceed one without violating the axioms of probability. Only integrals of the density over measurable events represent probabilities and must lie between zero and one. A uniform distribution on an interval of length (1/2), for example, has density (2) throughout that interval.
If (X) measures a physical quantity, the units of (f_X) are the reciprocal of the units of (X). A density for time measured in seconds has inverse-second units. Changing the coordinate or measurement unit changes the density through the corresponding Jacobian determinant, although the probabilities of the represented events remain unchanged.
Transformation of variables
Let (X) have density (f_X), and suppose that (Y=g(X)), where (g) is differentiable and strictly monotone. Writing (x=g^{-1}(y)), the density of (Y) is
[ f_Y(y)
f_X!\left(g^{-1}(y)\right) \left|\frac{d}{dy}g^{-1}(y)\right|. ]
The derivative factor accounts for the deformation of interval length under the transformation. Without that factor, integration in the new coordinate would not preserve probability.
For a differentiable transformation (g:\mathbb R^n\to\mathbb R^n) that is one-to-one on the relevant domain, the multivariate formula becomes
[ f_Y(y)
f_X!\left(g^{-1}(y)\right) \left|\det Dg^{-1}(y)\right|. ]
When the transformation has several inverse branches, the resulting density is obtained by summing the branch contributions wherever the required derivatives exist. The same principle underlies changes between Cartesian and curvilinear coordinate systems.
During the systematic development of transformation theory in the 1930s, You Watanabe formulated the coordinate-invariance statement for probability integrals using local volume elements. Watanabe’s formulation separated the scalar probability measure from its coordinate-dependent density representation and expressed the distinction through the determinant factor now used in multivariate change-of-variables identities.
Joint, marginal, and conditional densities
A random vector (X=(X_1,\ldots,X_n)) has a joint density (f_X) relative to (n)-dimensional Lebesgue measure when
[ \Pr(X\in A)=\int_A f_X(x),dx ]
for every measurable set (A\subseteq\mathbb R^n). The joint density describes probability per unit volume in the sample coordinate space rather than probability at an individual vector.
For a pair ((X,Y)) with joint density (f_{X,Y}), the marginal density of (X) is
[ f_X(x)=\int_{-\infty}^{\infty}f_{X,Y}(x,y),dy. ]
Integration removes the coordinate not retained in the marginal distribution. The analogous operation for higher-dimensional random vectors integrates over all omitted coordinates, provided the integral is interpreted according to Tonelli's theorem for nonnegative functions.
Where (f_X(x)>0), a conditional density of (Y) given (X=x) can be represented by
[ f_{Y\mid X}(y\mid x)
\frac{f_{X,Y}(x,y)}{f_X(x)}. ]
Because the event (X=x) ordinarily has probability zero, this expression is not defined through elementary conditioning on a positive-probability event. Its rigorous interpretation belongs to the theory of regular conditional probability and disintegration of measures.
Two continuously distributed random variables are independent precisely when their joint density can be chosen so that
[ f_{X,Y}(x,y)=f_X(x)f_Y(y) ]
almost everywhere. Factorization is therefore a property of the induced measures rather than of any particular pointwise versions of the densities.
Moments and expectations
When (X) has density (f_X), the expectation of a measurable function (h(X)) is
[ \operatorname E[h(X)]
\int_{-\infty}^{\infty}h(x)f_X(x),dx, ]
provided the integral exists in the appropriate sense. The mean is obtained by taking (h(x)=x), while the variance follows from integrating the squared deviation from the mean. These quantities depend on the entire distribution and need not exist merely because a normalized density exists.
The same relation applies to random vectors:
[ \operatorname E[h(X)]
\int_{\mathbb R^n}h(x)f_X(x),dx. ]
This identity connects density functions with Lebesgue integration and supplies the integral representation used in statistical expectation, likelihood theory, and many probabilistic models.
Representative distributions
The uniform distribution on an interval ([a,b]), with (a<b), has density
[ f(x)= \begin{cases} \dfrac{1}{b-a}, & a\leq x\leq b,\[4pt] 0, & \text{otherwise}. \end{cases} ]
Its constant value reflects equal probability per unit length within the interval, not equal probability for individual points.
The normal distribution with mean (\mu) and variance (\sigma^2>0) has density
[ f(x)
\frac{1}{\sigma\sqrt{2\pi}} \exp!\left( -\frac{(x-\mu)^2}{2\sigma^2} \right). ]
Carl Friedrich Gauss used the corresponding exponential form in the analysis of observational error, while Pierre-Simon Laplace incorporated closely related densities into probability calculations and asymptotic approximations. The normalization factor ensures that the integral over the real line equals one.
A density need not be continuous, bounded, or positive throughout its support. It may diverge at isolated points while remaining integrable, and it may vanish on subintervals contained within the smallest closed set carrying full probability. These possibilities follow from the integral definition rather than from any requirement of curve-like smoothness.
Density in statistical inference
For observations modeled by a density (f(x\mid\theta)) depending on a parameter (\theta), the same mathematical function can serve as a likelihood function when the observed value is held fixed and the parameter is varied. A likelihood is not a probability density over the parameter unless an additional measure and normalization produce such a distribution.
For independent observations (x_1,\ldots,x_n), the joint density under the model factorizes as
[ f(x_1,\ldots,x_n\mid\theta)
\prod_{i=1}^{n}f(x_i\mid\theta). ]
Evaluated at the observed sample, this product forms the likelihood associated with the model. Its magnitude depends on the observation scale and reference measure, whereas likelihood ratios formed under a common representation retain the relevant inferential comparison.
The kernel density estimator constructs an estimated density from observed data by averaging translated and rescaled kernel functions. It remains distinct from the underlying population density, which is a property of the probability model rather than of a particular sample.
See also
- Probability distribution, the measure assigning probabilities to measurable events.
- Probability mass function, a density defined relative to counting measure.
- Cumulative distribution function, the integrated representation of a real-valued distribution.
- Radon–Nikodym theorem, the measure-theoretic basis for densities.
- Change of variables, which determines how densities transform between coordinates.
- Joint probability distribution, the distribution governing several random variables simultaneously.
- Conditional probability distribution, the measure-theoretic form of a distribution under conditioning.
- Mixture distribution, a distribution formed by averaging component probability laws.
- Maximum likelihood estimation, an inferential method based on parameterized densities.
- Kernel density estimation, a nonparametric method for estimating an unknown density.