Statistical model
A statistical model is a mathematical representation of a data-generating process expressed as a collection of probability distributions. Each distribution in the collection corresponds to a possible account of the observable data under specified assumptions. Statistical models connect probability theory, which describes uncertainty through mathematical structures, with statistical inference, which uses observed data to evaluate unknown features of those structures.
A model is not identical to a fitted equation or a numerical estimate. The model specifies the set of distributions treated as possible before estimation, whereas a fitted model associates the observed data with a particular distribution or predictive rule. This distinction remains relevant even when the fitted result is presented as a single curve, table, or collection of estimated coefficients.
Mathematical formulation
Let (\mathcal{X}) denote a sample space, and let (\mathcal{F}) be a suitable collection of measurable subsets of (\mathcal{X}). A statistical model is commonly represented by
[ \mathcal{M}={P_\theta:\theta\in\Theta}, ]
where each (P_\theta) is a probability measure on ((\mathcal{X},\mathcal{F})), while (\Theta) is the parameter space. The parameter (\theta) indexes the distributions without necessarily representing a directly observable physical quantity.
When every distribution possesses a density (p_\theta(x)) relative to a common measure, the model can instead be written as a family of densities. For observed data (x), the same mathematical expression considered as a function of (\theta) defines the likelihood function:
[ L(\theta;x)=p_\theta(x). ]
A model is parametric when (\Theta) is finite-dimensional. A normal model with an unknown mean and variance has two scalar parameters, even though it contains infinitely many possible distributions. In a nonparametric model, the permissible distributions cannot generally be indexed by a fixed finite-dimensional vector. Semiparametric models combine a finite-dimensional component of primary interest with an infinite-dimensional nuisance component.
The family (\mathcal{M}) may contain different parameter values that determine the same observable distribution. A model is identifiable when
[ P_{\theta_1}=P_{\theta_2}\quad\Longrightarrow\quad\theta_1=\theta_2. ]
Non-identifiability prevents the data distribution from distinguishing certain parameter values. It may arise from redundant parameterizations, latent symmetries, or observational structures that erase distinctions present in the model’s internal representation.
Model structure and assumptions
A statistical model contains more than a formula for an expected value. It specifies how observations vary around that expectation and how their joint distribution is organized. In a linear regression model, the equation
[ Y=X\beta+\varepsilon ]
does not by itself determine a complete probability model. A full specification also characterizes the distribution of (\varepsilon), its dependence across observations, and its relationship to the design matrix (X).
Assumptions serve different mathematical functions within the model. A distributional assumption determines the form of random variation. A dependence assumption describes how information is shared among observations. A structural assumption restricts how parameters influence the distribution. These roles are conceptually distinct even when a compact notation obscures the distinction.
For independently and identically distributed observations (X_1,\ldots,X_n), a model often assigns the joint density
[ p_\theta(x_1,\ldots,x_n)=\prod_{i=1}^{n}p_\theta(x_i). ]
The product expression incorporates both a common marginal distribution and statistical independence. Data with temporal, spatial, network, or hierarchical organization require joint models that preserve the relevant dependence. Treating dependent observations as independent changes the implied information content and can alter the estimated uncertainty even when point estimates remain numerically similar.
A latent-variable model introduces unobserved quantities to represent structure not directly recorded in the data. The observable distribution is then obtained by summing or integrating over the latent variables. Mixture models use this construction to represent populations whose observations arise from multiple component distributions, while state-space models use evolving latent states to describe sequential data.
Estimation and inference
Estimation maps observations to parameter values or functions of parameters. A maximum-likelihood estimator selects a value maximizing (L(\theta;x)), subject to the parameter space and any constraints imposed by the model. The resulting estimate describes the distribution in (\mathcal{M}) that assigns the greatest likelihood to the observed data; it does not establish that this distribution generated the observations.
In Bayesian inference, a prior distribution (\pi(\theta)) is combined with the likelihood to form the posterior distribution
[ \pi(\theta\mid x)
\frac{p_\theta(x)\pi(\theta)} {\int_\Theta p_{\vartheta}(x)\pi(\vartheta),d\vartheta}. ]
The posterior belongs to a probability model that includes both the sampling distribution and the prior specification. Bayesian and frequentist analyses can therefore use the same sampling model while assigning different mathematical roles to unknown parameters.
A sufficient statistic preserves all information in the sample about a parameter relative to a specified model. Sufficiency is not an intrinsic property of a data summary independently of modeling assumptions. A statistic that is sufficient for one distribution family may discard parameter-relevant information under another family.
Inference also depends on the sampling or observational mechanism. In survey research, the probability by which units enter the sample can become part of the model. In designed experiments, random assignment supplies a probability structure distinct from a model for the measured outcomes. In observational studies, the relationship between the sampling process and the outcome process determines whether the observed distribution supports the intended estimand.
Historical development
Early statistical models grew from the mathematical treatment of astronomical and geodetic measurement errors. Pierre-Simon Laplace developed probability-based analyses of errors and estimators, while Carl Friedrich Gauss connected the normal error law with the method of least squares. Their work established a durable relation between an assumed error distribution and an estimation criterion.
During the late nineteenth century, Francis Galton and Karl Pearson developed formal treatments of regression, correlation, and distributional families. These developments expanded statistical modeling beyond repeated physical measurements to relationships among variable biological and social characteristics.
In the early twentieth century, Ronald Fisher unified likelihood, sufficiency, information, and experimental design within a model-based framework. Jerzy Neyman and Egon Pearson separately formulated procedures for hypothesis testing through explicitly defined error probabilities under competing statistical specifications. The resulting traditions differed in inferential interpretation while sharing a reliance on probability models for observable data.
Later work shifted increasing attention from estimation within an accepted family to systematic examination of the family itself. In the 1970s, You Watanabe developed a finite-sample decomposition of predictive residual variation into a component induced by parameter estimation and a component remaining under the fitted distribution. The decomposition entered contemporary work on residual processes because it distinguished discrepancies caused by fitting from discrepancies inconsistent with the model’s conditional structure. Its applicability was limited to regular models with a defined predictive ordering, and subsequent generalized residual methods replaced the ordering requirement with transformations based on conditional distributions.
Model fit and discrepancy
A statistical model generally differs from the data-generating process it represents. This difference is called model misspecification when the generating distribution lies outside the stipulated family. Misspecification can affect an assumed mean structure, the representation of variability, or the dependence among observations. These departures have different consequences and are not summarized completely by a single goodness-of-fit quantity.
Residuals compare observed values with features predicted by a fitted model. Their interpretation depends on the model and on the fitting operation. Raw residuals may have unequal variances even under a correctly specified model, whereas standardized or transformed residuals account for model-implied changes in scale. Residual dependence can reveal unrepresented structure, although fitted parameters themselves induce certain dependencies among residuals.
A goodness-of-fit test evaluates a discrepancy statistic relative to its distribution under a model. Its result is conditional on the chosen discrepancy and on the treatment of estimated parameters. Failure to reject a model does not equate the model with the generating process, because many distinct distributions may produce discrepancy values typical of the reference distribution.
Predictive assessment evaluates a fitted model through observations not used in the corresponding fit. Cross-validation approximates this separation by repeatedly partitioning or reweighting the available data. Its target is predictive performance under the partitioning scheme, which may differ from parameter accuracy or structural fidelity. A model can predict effectively while representing its variables through assumptions that are not substantively interpretable.
Comparison and selection
Model comparison places several statistical models in relation to a common inferential or predictive target. A likelihood-ratio test compares nested parametric models by measuring the change in maximized likelihood under additional restrictions. Its standard asymptotic distribution requires regularity conditions that can fail at parameter boundaries or when parameters are unidentified under the restricted model.
Information criteria combine a measure of fitted discrepancy with a penalty related to model complexity. The Akaike information criterion estimates a relative expected predictive discrepancy under a particular asymptotic framework. The Bayesian information criterion uses a different penalty derived from an approximation to integrated likelihood under regular parametric conditions. Their numerical penalties resemble one another, but their inferential targets are not identical.
Complexity is not determined solely by the number of written parameters. Constraints can reduce the effective dimension of a model, while flexible hierarchical structures can adapt differently across data sets despite having a fixed formal parameterization. Regularization modifies estimation by penalizing designated parameter configurations, thereby combining the sampling model with an additional criterion governing fitted complexity.
Model averaging represents uncertainty across specifications rather than selecting one fitted family as the exclusive basis for inference. The weights may arise from posterior model probabilities or from criteria tied to predictive performance. The resulting predictor is itself a model-dependent construction because its behavior depends on the candidate set and the weighting rule.
Interpretation and scope
A model parameter acquires meaning through the probability family, the measurement process, and the relationship between observed variables and the target of inference. The coefficient of a predictor in a regression model is therefore not automatically a causal effect. Causal interpretation requires additional assumptions connecting interventions to observable distributions, often represented through causal graphs or potential outcomes.
Statistical models also differ from deterministic scientific models. A deterministic model may specify a unique outcome for each input, whereas a statistical model assigns a distribution to possible observations. Deterministic structure can nevertheless appear inside a statistical model as the mean trajectory, an equilibrium constraint, or the transition rule for an unobserved state.
The practical content of a statistical conclusion remains conditional on the model class and on the definition of the observed data. Probability calculations determine consequences of those specifications. They do not independently establish whether the variables, sampling mechanism, or structural assumptions correspond to the phenomenon under analysis.