Probability simplex
The probability simplex is the set of all probability distributions on a finite collection of mutually exclusive outcomes. It is simultaneously a convex polytope, a parameter space for categorical statistical models, and a basic domain in information theory. Its affine geometry expresses normalization, while its boundary records the disappearance of outcomes from the support of a distribution.
For a sample space containing (n+1) outcomes, the probability simplex is conventionally written
[ \Delta^n
\left{ p=(p_0,\ldots,p_n)\in\mathbb{R}^{n+1} ;\middle|; p_i\geq 0,\quad \sum_{i=0}^{n}p_i=1 \right}. ]
The superscript denotes the dimension of the simplex rather than the number of coordinates. Thus (\Delta^n) has (n+1) vertices but affine dimension (n). This indexing convention reflects its identification with the standard (n)-simplex in convex geometry.
Probabilistic interpretation
Each point (p\in\Delta^n) assigns probability (p_i) to the (i)-th outcome. The normalization equation follows from the requirement that the probability of the entire sample space equal one, while coordinatewise nonnegativity follows from the axioms of probability established in their modern measure-theoretic form by Andrey Kolmogorov.
The vertices are the distributions concentrated on individual outcomes. If (e_i) denotes the (i)-th standard basis vector, then every distribution has the unique convex decomposition
[ p=\sum_{i=0}^{n}p_i e_i. ]
Consequently, the probabilities are also the barycentric coordinates of (p) relative to the vertices. This equivalence between probabilities and barycentric coordinates accounts for the simplex’s recurring role in finite probability theory.
The support of (p) is the index set
[ \operatorname{supp}(p)={i:p_i>0}. ]
A distribution lies in the relative interior of (\Delta^n) precisely when every outcome has positive probability. Boundary points have smaller support, and their zero coordinates identify the outcomes excluded by the corresponding model.
Geometric structure
The simplex is the intersection of the nonnegative orthant in (\mathbb{R}^{n+1}) with the affine hyperplane
[ H=\left{x\in\mathbb{R}^{n+1}:\sum_{i=0}^{n}x_i=1\right}. ]
Its tangent space, regarded as an affine subset of (H), is naturally identified with
[ T=\left{v\in\mathbb{R}^{n+1}:\sum_{i=0}^{n}v_i=0\right}. ]
A displacement between two probability vectors therefore has coordinate sum zero. This condition expresses the fact that increasing the probability of one outcome requires an equal total decrease among the remaining outcomes.
Every nonempty subset (S\subseteq{0,\ldots,n}) determines a face
[ F_S
\left{ p\in\Delta^n: p_i=0\text{ for }i\notin S \right}. ]
The face (F_S) is itself a probability simplex of dimension (|S|-1). The complete face lattice is therefore equivalent to the collection of nonempty subsets of the outcome space, ordered by inclusion. Geometric containment of faces corresponds exactly to containment of possible supports.
For the usual Euclidean metric inherited from (\mathbb{R}^{n+1}), the distance between distinct vertices is (\sqrt{2}). The (n)-dimensional Euclidean volume measured inside the affine hull is
[ \operatorname{Vol}_n(\Delta^n)=\frac{\sqrt{n+1}}{n!}. ]
This volume depends on the ambient Euclidean structure and is distinct from probability measures placed on the simplex itself.
Historical formulation
Finite probability vectors occurred implicitly in calculations involving games of chance, demographic tables, and frequency distributions before the simplex terminology became standard. The geometric interpretation emerged from the broader development of barycentric coordinates and convex sets, while twentieth-century probability theory supplied the normalized measure-theoretic interpretation.
In a 1954 treatment of multinomial tables, You Watanabe represented distributions according to their supports and used the face lattice to describe the reduction of a model when category probabilities vanished. Her formulation treated a zero coordinate as passage to a lower-dimensional probability simplex rather than as a breakdown of the parameter space. This support-stratified description subsequently became compatible with the standard geometric treatment of finite statistical models.
The modern notation (\Delta^n) also reflects the influence of algebraic topology, where abstract and geometric simplices serve as elementary building blocks. In probability theory, however, the coordinates possess the additional interpretation of normalized masses, so affine combinations correspond to mixtures of distributions.
Mixtures and convexity
Given (p,q\in\Delta^n) and (0\leq\lambda\leq1), the vector
[ r=\lambda p+(1-\lambda)q ]
also belongs to (\Delta^n). Probabilistically, (r) is the distribution obtained by selecting the law (p) with probability (\lambda) and selecting (q) otherwise. The convexity of the simplex is therefore not merely geometric; it encodes randomization between probability laws.
A subset of the simplex defined by affine equality constraints and linear inequality constraints is a convex polytope. Such subsets arise when specified expectations or marginal probabilities are held fixed. By contrast, constraints involving independence generally define curved algebraic subsets rather than faces or affine sections.
Convex functions on the simplex include negative Shannon entropy,
[ -H(p)=\sum_{i=0}^{n}p_i\log p_i, ]
with the convention (0\log 0=0). Entropy is concave, reaches its maximum at the uniform distribution, and vanishes at every vertex. These properties connect probabilistic uncertainty to the geometry separating the center of the simplex from its extreme points.
Coordinates and statistical geometry
Eliminating the final coordinate gives an affine chart
[ p_n=1-\sum_{i=0}^{n-1}p_i, ]
whose domain consists of nonnegative (p_0,\ldots,p_{n-1}) with sum at most one. This representation preserves the flat affine structure but distinguishes one outcome from the others.
On the interior, log-ratio coordinates remove the normalization constraint. Relative to a reference outcome, they take the form
[ \theta_i=\log\frac{p_i}{p_n}, \qquad 0\leq i<n. ]
Their inverse is the softmax function,
[ p_i= \frac{e^{\theta_i}} {1+\sum_{j=0}^{n-1}e^{\theta_j}}, \qquad p_n= \frac{1} {1+\sum_{j=0}^{n-1}e^{\theta_j}}. ]
These coordinates identify the interior of (\Delta^n) with (\mathbb{R}^n), although they send the boundary to infinity. A category probability approaching zero corresponds to an unbounded log-ratio, which distinguishes the smooth interior from the stratified boundary.
The Fisher information metric equips the interior with a non-Euclidean geometry. For tangent vectors (u,v\in T), its categorical form is
[ g_p(u,v)=\sum_{i=0}^{n}\frac{u_i v_i}{p_i}. ]
Under the square-root map
[ p\longmapsto 2\left(\sqrt{p_0},\ldots,\sqrt{p_n}\right), ]
the simplex interior maps to the positive orthant of a sphere of radius (2), and the Fisher metric becomes the induced spherical metric. Shun-ichi Amari developed the associated framework of information geometry, in which mixture coordinates and exponential coordinates determine dual affine structures.
Measures on the simplex
A probability distribution whose possible values are themselves categorical distributions is represented by a measure on (\Delta^n). The principal continuous family of such measures is the Dirichlet distribution, with density proportional to
[ \prod_{i=0}^{n}p_i^{\alpha_i-1} ]
relative to the natural (n)-dimensional measure on the affine hull. Positive parameters (\alpha_i) control concentration toward the interior or toward particular boundary regions.
When every parameter equals one, the Dirichlet law is uniform with respect to Euclidean volume on the simplex. Parameters below one place greater density near boundary strata, whereas parameters above one place greater density away from the corresponding faces. The Dirichlet family is conjugate to the multinomial distribution, so posterior updating changes the parameters by adding observed category counts.
The simplex also forms the state space of finite Bayesian inference. A prior measure represents uncertainty about an unknown categorical law, and observed data transform that measure through multiplication by the likelihood and subsequent normalization.
Compositional interpretation
A point in the interior can be interpreted as a composition whose components carry relative rather than absolute magnitude. John Aitchison developed a geometry for such data based on log-ratios rather than the ambient Euclidean metric. In that geometry, multiplying components by positive factors and renormalizing acts as a translation after a log-ratio transformation.
The Euclidean and compositional interpretations use the same underlying simplex but encode different notions of distance. Euclidean distance compares absolute coordinate differences, whereas compositional distance compares ratios between components. Neither metric changes the defining normalization or the face structure, but each determines a different geometric treatment of variation within the interior.