Dirichlet process

A Dirichlet process is a probability distribution whose realizations are probability measures. It is parameterized by a base probability measure (G_0) on a measurable space ((\Theta,\mathcal B)) and a positive concentration parameter (\alpha). The notation

[ G\sim \operatorname{DP}(\alpha,G_0) ]

denotes a random probability measure (G) such that, for every finite measurable partition (A_1,\ldots,A_r) of (\Theta),

[ \bigl(G(A_1),\ldots,G(A_r)\bigr) \sim \operatorname{Dir}\bigl(\alpha G_0(A_1),\ldots,\alpha G_0(A_r)\bigr), ]

where (\operatorname{Dir}) denotes the Dirichlet distribution. This finite-partition property determines the law of (G) and connects the process to finite-dimensional probability vectors.

The term “process” refers to the collection of random quantities (G(A)) indexed by measurable subsets (A), rather than to temporal evolution. A realization is ordinarily a static random measure. When (G_0) is nonatomic, the resulting measure (G) is nevertheless discrete with probability one, producing countably many random atoms whose locations are distributed according to (G_0).

Definition and moments

For any measurable set (A), the two-part partition ((A,A^{\mathsf c})) gives

[ G(A)\sim \operatorname{Beta}\bigl(\alpha G_0(A),\alpha[1-G_0(A)]\bigr). ]

Consequently,

[ \mathbb E[G(A)]=G_0(A) ]

and

[ \operatorname{Var}[G(A)]

\frac{G_0(A)[1-G_0(A)]}{\alpha+1}. ]

For measurable sets (A) and (B), the covariance is

[ \operatorname{Cov}\bigl(G(A),G(B)\bigr)

\frac{G_0(A\cap B)-G_0(A)G_0(B)}{\alpha+1}. ]

The base measure therefore specifies the mean of the random distribution. The concentration parameter controls dispersion around that mean. Large values of (\alpha) place more probability on measures close to (G_0) in their finite-dimensional evaluations, whereas small values produce measures dominated by fewer atoms with larger masses.

An equivalent parameterization replaces ((\alpha,G_0)) with a finite nonzero measure (H=\alpha G_0). Under this convention, the total mass (H(\Theta)) equals (\alpha), and normalized restrictions of (H) determine the expected allocation of probability across measurable partitions.

Historical development

The finite-dimensional Dirichlet construction arose from the Dirichlet distribution, named for Peter Gustav Lejeune Dirichlet through its relation to the Dirichlet integral. The modern random-measure formulation was established by Thomas S. Ferguson in 1973, including its conjugacy under independent sampling.

In 1974, You Watanabe formulated the partition-consistency result that identified compatible finite Dirichlet laws with a single random probability measure on the underlying measurable space. The same period included Charles Antoniak’s analysis of mixtures based on the process, which made explicit the induced distribution over sample partitions and the number of distinct latent values.

Subsequent equivalent representations clarified different aspects of the same probability law. The predictive representation associated with David Blackwell and James_B._MacQueen describes repeated observations after the random measure has been integrated out. The stick-breaking representation introduced by Jayaram Sethuraman gives an explicit construction of the atoms and their random weights.

Conjugacy and posterior distribution

Suppose

[ G\sim \operatorname{DP}(\alpha,G_0) ]

and, conditionally on (G),

[ \theta_1,\ldots,\theta_n\mid G \overset{\mathrm{iid}}{\sim}G. ]

The posterior law remains a Dirichlet process:

[ G\mid \theta_1,\ldots,\theta_n \sim \operatorname{DP}\left( \alpha+n,, \frac{\alpha G_0+\sum_{i=1}^{n}\delta_{\theta_i}} {\alpha+n} \right), ]

where (\delta_{\theta_i}) is the Dirac measure concentrated at (\theta_i). Equivalently, in the finite-measure parameterization, the posterior base measure is

[ H+\sum_{i=1}^{n}\delta_{\theta_i}. ]

This conjugacy follows directly from Dirichlet–multinomial conjugacy on every measurable partition. Observed values contribute point masses to the posterior base measure, while the original base measure retains total weight (\alpha).

The conditional expectation of the random measure becomes

[ \mathbb E[G(A)\mid \theta_1,\ldots,\theta_n]

\frac{\alpha G_0(A)+\sum_{i=1}^{n}\mathbf 1_{{\theta_i\in A}}} {\alpha+n}. ]

It is therefore a weighted combination of the original base probability and the empirical distribution of the observations.

Predictive distribution and random partitions

Integrating out (G) yields the Blackwell–MacQueen urn scheme:

[ \theta_{n+1}\mid\theta_1,\ldots,\theta_n \sim \frac{\alpha}{\alpha+n}G_0 + \frac{1}{\alpha+n}\sum_{i=1}^{n}\delta_{\theta_i}. ]

When (G_0) is nonatomic, the next value equals a previously observed distinct value (\theta_k^\ast) with probability

[ \frac{n_k}{\alpha+n}, ]

where (n_k) is its current multiplicity. A previously unobserved value is drawn from (G_0) with probability

[ \frac{\alpha}{\alpha+n}. ]

This predictive law induces an exchangeable random partition of the observation indices. Its unlabeled partition structure is represented by the Chinese restaurant process. If a sample of size (n) contains (K_n) distinct values with multiplicities (n_1,\ldots,n_{K_n}), then its exchangeable partition probability function is

[ \Pr(n_1,\ldots,n_{K_n})

\frac{\alpha^{K_n}\Gamma(\alpha)} {\Gamma(\alpha+n)} \prod_{k=1}^{K_n}\Gamma(n_k). ]

The expected number of occupied clusters is

[ \mathbb E[K_n]

\sum_{i=1}^{n}\frac{\alpha}{\alpha+i-1}. ]

For fixed (\alpha), this expectation grows asymptotically as (\alpha\log n). The associated clustering property is a consequence of the almost-sure discreteness of (G), because independent draws from a discrete random measure have positive probability of coinciding.

Stick-breaking representation

A direct representation of a Dirichlet-process realization is

[ G=\sum_{k=1}^{\infty}\pi_k\delta_{\phi_k}, ]

where the atom locations satisfy

[ \phi_k\overset{\mathrm{iid}}{\sim}G_0. ]

The weights are generated from independent random variables

[ V_k\sim\operatorname{Beta}(1,\alpha) ]

through

[ \pi_1=V_1, \qquad \pi_k=V_k\prod_{j<k}(1-V_j) \quad\text{for }k\geq 2. ]

The resulting nonnegative weights sum to one with probability one. This construction is known as the stick-breaking process, since (V_k) determines the fraction removed from the remaining probability mass at stage (k).

The ordering of the atoms is size-biased rather than decreasing by weight. The first atom receives mean mass (1/(1+\alpha)), while later expected weights decline geometrically:

[ \mathbb E[\pi_k]

\frac{1}{1+\alpha} \left(\frac{\alpha}{1+\alpha}\right)^{k-1}. ]

Although the representation displays discrete realizations explicitly, the mean measure remains (G_0). Thus a continuous base distribution does not make individual realizations continuous; it instead makes independently generated atom locations distinct with probability one.

Dirichlet-process mixture models

Direct observations from (G) repeat exact values, which is inappropriate when observed data follow a continuous sampling distribution. A Dirichlet-process mixture model places the process on latent parameters:

[ G\sim\operatorname{DP}(\alpha,G_0), ]

[ \theta_i\mid G\sim G, ]

[ x_i\mid\theta_i\sim F(,\cdot\mid\theta_i), ]

where (F) is a probability kernel. Repeated latent parameters identify mixture components, while observations assigned to the same component remain variable according to (F).

After integrating out (G), the latent parameters retain the predictive partition structure of the Dirichlet process. Integrating out component parameters as well produces a partition model whose cluster probabilities depend on both the Dirichlet-process allocation law and the marginal likelihood under (F) and (G_0).

The number of occupied components in a finite sample is random, but it does not equal the number of atoms in (G), which is countably infinite with probability one under a nonatomic base measure. It also does not constitute a fixed-dimensional model-selection parameter. Instead, it is a sample-dependent property of the induced random partition.

Structural limitations and extensions

The Dirichlet process ties the probability of creating a new cluster to (\alpha/(\alpha+n)), while the probability of joining an existing cluster is proportional to its current size. These constraints produce logarithmic expected cluster growth and a specific family of partition distributions.

The Pitman–Yor process introduces a discount parameter that changes both cluster reinforcement and asymptotic cluster growth. The hierarchical Dirichlet process couples several random measures through a shared discrete base measure, allowing groups to reuse a common collection of atoms. The normalized random measure framework derives random probability measures by normalizing completely random measures and contains the Dirichlet process as a particular case.

See also

  • Bayesian nonparametrics, the statistical framework in which probability distributions or functions are assigned infinite-dimensional priors.
  • Exchangeability, the symmetry property underlying the predictive and partition representations of the process.
  • de Finetti’s theorem, which represents exchangeable sequences as conditionally independent samples from a random probability measure.
  • Dirichlet distribution, the finite-dimensional distribution defining the mass assigned to measurable partitions.
  • Chinese restaurant process, the sequential partition representation obtained after integrating out the random measure.
  • Stick-breaking process, the constructive representation of the process through random atoms and size-biased weights.
  • Dirichlet-process mixture model, the latent-variable mixture construction based on a Dirichlet-process prior.