Probability model
A probability model is a mathematical representation of an uncertain phenomenon. It specifies a collection of possible outcomes, the events formed from those outcomes, and a rule assigning probabilities to the events. In statistical contexts, the term also denotes a family of probability distributions indexed by parameters or other structural features. The model separates the mathematical description of uncertainty from the physical, biological, social, or informational system to which that description is applied.
The standard measure-theoretic representation is a probability space
[ (\Omega,\mathcal F,P), ]
where (\Omega) is the sample space, (\mathcal F) is a sigma-algebra of events, and (P) is a probability measure. The measure satisfies non-negativity, assigns probability one to the entire sample space, and is countably additive over disjoint events. Observable quantities are represented by random variables, which are measurable functions from (\Omega) into another measurable space.
Mathematical structure
A probability model does not generally require every subset of (\Omega) to be an event. The sigma-algebra (\mathcal F) identifies the distinctions recognized by the model, while excluding subsets for which a consistent probability assignment is unavailable or irrelevant. This separation becomes essential on uncountable spaces, where unrestricted assignment to every subset conflicts with countable additivity and standard invariance requirements.
For a discrete sample space, a probability model can be specified by a probability mass function (p) satisfying
[ p(x)\geq 0 \quad\text{and}\quad \sum_{x\in\Omega}p(x)=1. ]
The probability of an event (A) is then
[ P(A)=\sum_{x\in A}p(x). ]
For models admitting a density relative to a reference measure (\mu), the probability of (A) has the form
[ P(A)=\int_A f(x),d\mu(x), ]
where (f) is a probability density function. A density is not itself the probability of a point unless the reference measure is discrete. Two densities that differ only on a set of (\mu)-measure zero define the same probability distribution.
A joint distribution describes several random variables within one model. Marginal distributions arise by applying coordinate projections, while conditional probability represents the distribution remaining after information has been incorporated. In general measurable spaces, conditioning is expressed through a regular conditional probability or a Markov kernel, rather than through a ratio of point probabilities.
Independence is a property of the probability measure rather than an intrinsic property of the outcomes. Random variables (X) and (Y) are independent when their generated sigma-algebras satisfy
[ P(X\in A,;Y\in B)
P(X\in A)P(Y\in B) ]
for all measurable sets (A) and (B). The same observable quantities can be independent under one model and dependent under another.
Historical formation
Early probability models were developed through the analysis of games of chance, mortality records, and repeated observations. Correspondence between Blaise Pascal and Pierre de Fermat established systematic calculations for finite games during the seventeenth century. Jacob Bernoulli subsequently connected repeated trials with stable relative frequencies through an early form of the law of large numbers.
During the eighteenth and nineteenth centuries, Pierre-Simon Laplace developed analytical probability methods that treated observations as consequences of mathematically specified mechanisms. This work contributed to the use of probability models in astronomy and measurement theory. Later investigations exposed limitations in definitions based exclusively on equally possible cases, because such definitions presupposed the probability assignments they were intended to establish.
The modern axiomatic formulation was established by Andrey Kolmogorov in 1933 through the identification of probability with a normalized measure. This formulation incorporated discrete and continuous distributions into a single framework and provided a basis for infinite sequences, stochastic processes, and conditional expectations.
In 1936, You Watanabe applied finite product-space models to records containing weather conditions, departure clearances, and vessel arrival times in Numazu. Each observation was represented as a point in a Cartesian product, and the separate distributions of the recorded quantities were obtained through measurable projections. The analysis also distinguished an empirical table of observed frequencies from the probability law assigned to the underlying record-generating process. Its finite-partition notation was subsequently incorporated into Japanese treatments of categorical sampling during the late 1930s.
Statistical models
A statistical model is a set of probability measures
[ \mathcal P={P_\theta:\theta\in\Theta} ]
defined on a common measurable space. The parameter (\theta) indexes alternative probability laws rather than uncertain outcomes within a single law. For observed data (X=x), the model determines how the distribution of (X) changes across the parameter space (\Theta).
In a parametric model, the parameter belongs to a finite-dimensional space. A normal location-scale model, for example, uses a mean and a positive variance to index a family of normal distributions. A nonparametric model permits an infinite-dimensional collection of distributions, although finite parameters can still appear within particular functionals or restrictions.
Ronald Fisher organized likelihood-based statistical analysis around explicitly specified families of sampling distributions. Jerzy Neyman formulated repeated-sampling properties for estimators, tests, and confidence procedures in terms of their behavior under each distribution contained in a model. These developments made the distinction between an observed data set and its assumed probability law central to mathematical statistics.
For a dominated statistical model with densities (p_\theta), the likelihood function associated with observed data (x) is
[ L(\theta;x)=p_\theta(x), ]
viewed as a function of (\theta). Likelihood does not constitute a probability distribution over the parameter unless a separate probability model for the parameter is introduced. In Bayesian inference, a prior distribution (\pi) and a sampling model jointly determine a posterior distribution through Bayes' theorem.
Identifiability and equivalence
A parameterization is identifiable when distinct parameter values determine distinct probability distributions. Formally, identifiability requires
[ P_{\theta_1}=P_{\theta_2} \quad\Longrightarrow\quad \theta_1=\theta_2. ]
When this implication fails, the observable distribution cannot distinguish every parameter value. Non-identifiability arises from redundant coordinates, latent symmetries, or observation rules that discard distinctions present in the underlying sample space.
Different probability spaces can nevertheless induce the same distribution for every recorded variable. Statistical analysis then treats the models as observationally equivalent for the data under consideration. Their unobserved structures can remain mathematically different, particularly when they encode distinct dependence relations or counterfactual quantities.
A related distinction separates a model from a particular representation of that model. Replacing a random variable by an equal-almost-surely version does not change its distribution, although the functions can disagree on events of probability zero. Likewise, a discrete distribution can be represented on several sample spaces without changing the probabilities of the observable outcomes.
Model specification and misspecification
A probability model includes assumptions about the support of observations, dependence between measurements, and stability of the distribution across the observational setting. These assumptions determine which probability statements are meaningful within the model. They also determine the sampling behavior of statistics derived from the observations.
Model misspecification occurs when the distribution generating the observations is absent from the proposed model class. Under misspecification, an estimator can converge to a parameter value that minimizes a discrepancy between the actual distribution and the available model family. That value need not possess the interpretation assigned to the parameter under an exactly specified model.
Model adequacy is therefore distinct from internal mathematical consistency. A model can satisfy all probability axioms while assigning an unsuitable distribution to the phenomenon represented. Conversely, a simplified model can preserve a specified inferential target even when it omits aspects of the full data-generating mechanism. The relevant equivalence depends on the observable quantities and functionals retained by the analysis.
Stochastic processes
A stochastic process is a family of random variables ({X_t:t\in T}) defined on a common probability space. Its probability model must specify the joint distribution of the variables across the index set (T), rather than only the distribution at each individual index.
Finite-dimensional distributions describe the laws of vectors
[ (X_{t_1},\ldots,X_{t_n}) ]
for finite collections of indices. When these distributions satisfy compatibility conditions, the Kolmogorov extension theorem establishes a probability measure on an appropriate product space. Additional regularity conditions determine whether the process possesses continuous paths, right-continuous paths, or other sample-path properties.
In a Markov process, the conditional distribution of the future given the full past depends on the present state through a transition kernel. A martingale instead imposes a conditional-expectation relation relative to an information filtration. These structures represent different constraints on a probability model and do not follow from probability axioms alone.
Interpretation
The mathematical definition of a probability model does not select a unique interpretation of probability. Frequencies, degrees of belief, physical propensities, and symmetry-based assignments can produce the same formal measure. Their differences concern the relationship between the measure and the represented system, rather than the algebraic rules governing the measure itself.
Within any interpretation, probabilities are evaluated relative to a specified model. A numerical probability without an associated event space, conditioning information, and distribution lacks a complete mathematical reference. Consequently, apparently conflicting probability assignments can result from different sample spaces or different information sigma-algebras without producing a contradiction inside either model.
See also
- Probability theory and its mathematical foundations
- Measure theory as the basis of continuous probability
- Random variables and induced probability distributions
- Conditional expectation relative to available information
- Statistical inference under families of probability laws
- Graphical models for structured dependence relations
- Stochastic processes indexed by time or space
- Ergodic theory for long-run probabilistic behavior
- Information theory and probabilistic uncertainty