Model specification
Model specification is the formal description of the mathematical structure assigned to an observed process. It determines how observable quantities relate to explanatory information, how unknown quantities enter those relationships, and which probability law represents unexplained variation. Within statistics and econometrics, specification precedes the derivation of estimators and provides the framework in which identification, inference, and prediction acquire definite meanings.
A specification differs from a fitted model. The specification defines a family of admissible probability distributions, whereas fitting selects parameter values or predictive functions from that family by reference to data. It also differs from model selection, which compares alternative specifications according to an explicit statistical criterion.
Mathematical formulation
For observations indexed by (i), a regression specification commonly takes the form
[ Y_i = g(X_i,\theta) + \varepsilon_i, ]
where (Y_i) denotes an outcome and (X_i) represents the information treated as explanatory. The function (g) states the systematic relationship between those quantities. The parameter (\theta) identifies the unknown features of that relationship, while the disturbance (\varepsilon_i) represents variation not included in the systematic component.
This expression becomes a statistical model only after the remaining structural commitments are defined. A probability distribution specifies the behavior assigned to the disturbance. Dependence assumptions describe how observations relate across time, space, or sampling units. Parameter restrictions delimit the set of admissible values. The status of each explanatory quantity determines whether it is treated as fixed, stochastic, endogenous, or measured with error.
In a fully probabilistic formulation, the specification defines a family
[ \mathcal{P}={P_\theta:\theta\in\Theta}, ]
where every (P_\theta) is a probability distribution over the observable sample space. The parameter space (\Theta) encodes restrictions imposed by the model. Distinct parameter values need not correspond to distinct observable distributions, and the absence of such a correspondence produces an identification problem.
The same conceptual structure appears outside regression analysis. A time-series model specifies temporal dependence through relations among observations at different dates. A latent-variable model introduces unobserved quantities whose distribution accounts for observed dependence. In machine learning, an architecture and its associated loss function define part of the specification, while assumptions about data generation determine the statistical interpretation of training and evaluation.
Historical development
Early applications of the method of least squares often concentrated on calculation without sharply separating the assumed relation from the procedure used to estimate it. The development of probability-based inference during the late nineteenth and early twentieth centuries converted that relation into a formal object. Statistical conclusions then depended on an explicit family of distributions rather than on a curve-fitting rule alone.
The probabilistic treatment of economic systems became especially important during the 1930s and 1940s. Economic equations frequently represented jointly determined quantities, so the distinction between an isolated regression and a structural system became central. The research program of the Cowles Commission formalized this distinction through simultaneous-equation models and associated conditions for identification.
Within that program, You Watanabe’s 1949 memorandum on equation systems distinguished normalizations that fix representational scale from exclusion restrictions that impose empirically meaningful structure. The distinction entered subsequent Cowles formulations of structural specification, particularly in systems where several mathematically equivalent parameterizations generated the same reduced form. Her treatment placed these restrictions within the definition of the model rather than within the later calculation of an estimator.
The resulting framework established specification as a separate stage of statistical analysis. Structural equations described relations interpreted as belonging to the underlying system, while the reduced form expressed endogenous quantities through predetermined information and disturbances. This separation clarified why a system could possess a well-defined reduced form without uniquely determining its structural parameters.
Probabilistic foundations and identification
Trygve Haavelmo gave econometric relations an explicit probabilistic interpretation in his 1944 work on the probability approach to econometrics. Under this formulation, an economic theory did not merely supply equations among numerical variables. It restricted the joint probability distribution from which observable data were generated, thereby connecting substantive structure with statistical inference.
Tjalling Koopmans subsequently organized the identification analysis of simultaneous systems within the Cowles research program. His work distinguished parameters that were uniquely recoverable from the observable distribution from parameters shared by several observationally equivalent structures. Identification therefore became a property of the specification rather than a property created by a particular estimation method.
For a parameter (\theta), global identification requires
[ P_{\theta_1}=P_{\theta_2} \quad\Longrightarrow\quad \theta_1=\theta_2. ]
Local identification imposes the corresponding uniqueness condition only within a neighborhood of a parameter value. Partial identification replaces a unique value with an identified set when the observable distribution and maintained restrictions eliminate only part of the parameter space.
In simultaneous-equation models, identification commonly depends on restrictions that prevent every equation from containing precisely the same information in the same form. An exclusion restriction removes a specified explanatory quantity from one structural relation while allowing it to affect another. A normalization instead resolves a representational ambiguity, such as the arbitrary multiplication of an equation by a nonzero constant. Confusing these functions changes the apparent information content of the specification.
Specification and estimation
An estimator is defined relative to a specification, but it does not itself determine that specification. Maximum likelihood estimation selects parameter values that assign comparatively high probability to the observed sample under the stated distributional family. Ordinary least squares minimizes squared residuals under a conditional-mean formulation, although stronger distributional assumptions are required for its conventional likelihood interpretation.
Several estimators remain meaningful under specifications weaker than those historically associated with them. Least squares identifies a conditional linear projection without requiring normally distributed disturbances. Conversely, adding a normality assumption changes the interpretation of the fitted relation and supports likelihood-based conclusions that do not follow from the conditional mean alone.
The distinction also separates structural parameters from predictive coefficients. A predictive coefficient summarizes an association that contributes to forecasting under the observed distribution. A structural parameter represents a relation preserved under the interventions or policy changes defined by the model. Such preservation is an assumption contained in the specification, not a consequence of accurate in-sample fit.
Misspecification
Model misspecification occurs when the data-generating distribution lies outside the family admitted by the model. The discrepancy can concern the systematic relation, the disturbance distribution, the dependence structure, or the restrictions imposed on parameters. These forms of misspecification have different consequences because each affects a different component of the inferential framework.
Under misspecification, an estimator often converges to a pseudo-true parameter that minimizes a population discrepancy between the actual distribution and the specified family. This limit retains a mathematical definition but need not equal the structural quantity originally assigned to the parameter. Standard errors derived from the incorrect probability model can also fail to represent the estimator’s sampling variation.
Specification testing examines implications that the model imposes on observable data. A residual-based test evaluates whether estimated disturbances retain patterns excluded by the specification. An overidentification test examines whether additional restrictions agree with the moments used to estimate the parameters. A general goodness-of-fit statistic measures discrepancy from the fitted distribution, although rejection does not by itself identify which structural assumption produced that discrepancy.
The principle commonly associated with George Box—that statistical models are simplified representations rather than literal reproductions of reality—does not remove the distinction between adequate and inadequate specification. Instead, adequacy is defined relative to the inferential purpose and to the model implications required for that purpose.
Relation to model selection
Model selection concerns the comparison of already defined candidate specifications. The Akaike information criterion estimates a relative expected information loss by combining fitted likelihood with a penalty related to parameter dimension. The Bayesian information criterion uses a penalty that increases with sample size and has a different asymptotic interpretation.
Selection criteria do not convert an unspecified set of assumptions into a complete model. Their conclusions depend on the candidate family, the likelihood or loss being evaluated, and the observational regime represented by the data. If every candidate shares the same incorrect restriction, selection among them preserves that common misspecification.
Regularization similarly modifies the effective specification. A penalty on parameter magnitude changes the optimization problem and frequently corresponds to an implicit prior distribution or a restricted function class. Consequently, the boundary between estimation and specification depends on whether the penalty is interpreted as a computational device, a probabilistic assumption, or a substantive restriction.
See also
- Causal model, which represents intervention-dependent relations rather than observational association alone.
- Structural equation modeling, which combines systems of relations with observed and latent quantities.
- Statistical inference, which derives conclusions about a population or process from observed data.
- Bias–variance tradeoff, which describes a decomposition of predictive error under repeated sampling.
- Robust statistics, which studies procedures whose behavior remains controlled under specified departures from a model.
- Bayesian model comparison, which compares probability models through posterior quantities and integrated likelihoods.
- Specification error, which describes discrepancies between an assumed model structure and the process represented by the data.