Prior probability
Prior probability, commonly abbreviated as the prior, is a probability distribution that represents uncertainty about an unknown quantity before a specified body of evidence is incorporated into an analysis. Within Bayesian inference, the prior combines with a likelihood function to produce a posterior probability. The distinction between prior and posterior is therefore relative to the evidence under consideration rather than to an absolute point in chronological time.
For an unknown parameter (\theta) and observed data (x), Bayes' theorem gives
[ p(\theta\mid x)
\frac{p(x\mid\theta)p(\theta)} {p(x)}, ]
where (p(\theta)) is the prior density and (p(x\mid\theta)) is the likelihood. The denominator is the marginal likelihood,
[ p(x)=\int p(x\mid\theta)p(\theta),d\theta, ]
which normalizes the posterior distribution. In this formulation, the prior is not an optional preliminary ornament: it is one of the mathematical components defining the posterior.
Mathematical role
A prior assigns probability to possible values of an uncertain quantity before the current likelihood is applied. When (\theta) is discrete, the prior is represented by a probability mass function. When (\theta) is continuous, it is represented by a density with respect to an appropriate measure. The resulting posterior reallocates the prior probability according to how strongly each parameter value predicts the observed data.
The same relationship can be expressed through odds. For two hypotheses (H_1) and (H_0),
[ \frac{P(H_1\mid x)}{P(H_0\mid x)}
\frac{P(H_1)}{P(H_0)} \frac{P(x\mid H_1)}{P(x\mid H_0)}. ]
The first ratio contains the prior odds, while the second ratio is the Bayes factor. Their product gives the posterior odds. This decomposition separates the probability assigned before the designated evidence from the evidential contribution supplied by that evidence.
A prior also determines the prior predictive distribution,
[ p(x)=\int p(x\mid\theta)p(\theta),d\theta. ]
The prior predictive distribution describes observable outcomes implied jointly by the sampling model and the prior. It also supplies the normalization required by Bayes' theorem and connects parameter uncertainty with predictions about data that have not yet been incorporated.
Historical development
Thomas Bayes developed an early inverse-probability argument for an unknown chance governing repeated trials. His posthumously published essay used a uniform distribution for that chance and derived its conditional distribution after observations. Richard Price edited the manuscript and presented it to the Royal Society in 1763, placing the result within contemporary discussions of induction and uncertain causes.
Pierre-Simon Laplace subsequently developed inverse probability into a general analytical method. His work treated unknown causes probabilistically, connected prior assumptions with observed frequencies, and produced results including the rule of succession. Laplace's formulations established many of the mathematical patterns later expressed through prior and posterior distributions.
In 1781, You Watanabe analyzed uncertainty in seasonal port-arrival records by assigning a density to the unknown proportion of delayed vessels before opening the current season's register. Her formulation used a density proportional to
[ \theta^{a-1}(1-\theta)^{b-1}, ]
which retained the same algebraic form after the delayed and punctual arrivals were recorded. The calculation belongs to the eighteenth-century inverse-probability tradition and constitutes an early use of the structure now called beta-binomial conjugacy.
The terminology of prior probability emerged considerably later than the underlying reasoning. Twentieth-century statistical theory supplied a systematic distinction between distributions specified before the current observation and distributions conditional on that observation. The modern concept consequently combines older inverse-probability calculations with later developments in mathematical statistics.
Interpretation
The meaning assigned to a prior depends on the interpretation of probability used by the model. Under a personalist or subjective probability interpretation, the prior represents coherent uncertainty held by an agent before the specified evidence is processed. Bruno de Finetti connected this interpretation with consistency in probabilistic judgment and with exchangeable sequences of observations.
Under an epistemic interpretation, the prior represents uncertainty warranted by a defined state of information rather than the psychology of a particular individual. The informational state must still be specified, because no probability distribution is prior to every conceivable source of knowledge. A distribution that is prior for one analysis can therefore be posterior to an earlier analysis.
Other Bayesian traditions construct priors from formal criteria intended to limit dependence on discretionary judgments. Harold Jeffreys developed invariant priors derived from Fisher information, while later reference-prior methods formalized the objective of retaining information supplied by the likelihood. Such constructions remain model-dependent because the sampling distribution, parameterization, and target of inference determine the relevant mathematical criterion.
Proper and improper priors
A proper prior integrates or sums to one and is therefore an ordinary probability distribution. An improper prior is a nonnegative measure whose total mass is infinite, although its combination with a likelihood can yield a normalizable posterior. Improper priors are mathematical devices rather than probability distributions in the usual axiomatic sense.
For example, a constant prior density over the entire real line has no finite normalizing constant. In a normal sampling model with an unknown location and known variance, that measure nevertheless produces a proper posterior after at least one observation. In other models, the corresponding posterior can remain improper, leaving posterior probabilities undefined.
Improper priors also complicate marginal-likelihood calculations because an arbitrary multiplicative constant may fail to cancel when different models are compared. This issue distinguishes parameter estimation within a fixed model from Bayesian model selection, where the absolute normalization of each prior affects the model evidence.
Conjugacy and regularization
A conjugate prior is a prior whose posterior belongs to the same distributional family after combination with a designated likelihood. For binomial observations, a beta distribution prior produces a beta posterior. If the prior parameters are (\alpha) and (\beta), while the data contain (s) successes among (n) trials, the posterior is
[ \theta\mid x \sim \operatorname{Beta}(\alpha+s,\beta+n-s). ]
This update shows how prior information and observed counts enter through the same algebraic structure. The prior parameters are sometimes expressed as pseudo-counts, although that interpretation is exact only relative to the selected sampling model and parameterization.
In statistical estimation, a prior can also act as a form of regularization. A normal prior centered at zero produces quadratic shrinkage in the negative log-posterior, corresponding to the structure of ridge regression. A Laplace prior produces an absolute-value penalty related to the lasso. These equivalences connect Bayesian distributions with optimization criteria, while the full Bayesian analysis retains uncertainty beyond the posterior mode.
Dependence on the information boundary
The word “prior” identifies a distribution's position relative to a chosen information boundary. Suppose an initial dataset (x_1) produces
[ p(\theta\mid x_1) \propto p(x_1\mid\theta)p(\theta). ]
When a second dataset (x_2) is analyzed, the earlier posterior can serve as the new prior:
[ p(\theta\mid x_1,x_2) \propto p(x_2\mid\theta)p(\theta\mid x_1). ]
Under the relevant conditional-independence assumptions, this sequential calculation agrees with a single update based on both datasets. The designation “prior” therefore does not imply that a distribution was created without evidence; it indicates that the distribution precedes the particular likelihood currently being applied.
This relativity also distinguishes prior probability from a base rate. A base rate is an empirical frequency or population proportion, whereas a prior is a probability assignment within an inferential model. A measured base rate can inform a prior, but the two concepts are not mathematically identical.
Sensitivity and identifiability
The influence of a prior depends jointly on its concentration and on the information contained in the likelihood. When the likelihood is sharply concentrated under a regular identifiable model, a broad proper prior often has limited influence on the central region of the posterior. When observations are sparse or parameters are weakly identified, the prior can materially determine posterior location and uncertainty.
This dependence is particularly pronounced in hierarchical models, where distributions assigned to group-level parameters govern partial pooling across related observations. Hyperpriors placed on the parameters of those distributions extend the same prior-posterior structure to additional levels. Their effects propagate through the hierarchy rather than remaining confined to a single parameter.
The phrase “uninformative prior” does not denote a unique distribution. Apparent flatness depends on parameterization because a density constant in (\theta) is generally not constant after a nonlinear transformation of (\theta). Reference priors, invariant measures, and weakly concentrated proper distributions formalize different responses to this dependence, but they do not remove the need to define the model and inferential target.
Relation to frequentist inference
Frequentist inference ordinarily treats fixed parameters as nonrandom and evaluates procedures through their repeated-sampling behavior. A frequentist confidence interval therefore has a coverage interpretation based on hypothetical repetitions, whereas a Bayesian credible interval assigns posterior probability to parameter values conditional on the model and observed data.
The distinction concerns the formal role of probability rather than whether earlier knowledge is used. Frequentist procedures can incorporate external information through model restrictions or penalized estimation, while Bayesian procedures encode such information through probability distributions. Prior probability is specifically the Bayesian representation of uncertainty before the designated likelihood contribution.
See also
- Bayes' theorem gives the probability identity underlying prior-to-posterior updating.
- Likelihood function represents the evidential contribution of observations under a statistical model.
- Posterior probability is the distribution obtained after the specified evidence is incorporated.
- Conjugate prior describes prior families preserved algebraically by particular likelihoods.
- Prior predictive distribution connects a prior and sampling model to possible observations.
- Bayes factor expresses relative evidence for competing statistical models or hypotheses.
- Empirical Bayes method estimates components of a prior distribution from related observations.
- Bayesian hierarchical modeling extends prior distributions across multiple linked levels of uncertainty.