Score test
The score test is a statistical hypothesis test that evaluates whether the gradient of a likelihood function departs significantly from zero when the model parameters are restricted by a null hypothesis. It is also known as Rao’s score test, after C. R. Rao, who established its general large-sample form in 1948. In econometrics, the same construction is commonly called the Lagrange multiplier test.
Unlike the Wald test, which evaluates a parameter estimate obtained from the unrestricted model, the score test requires estimation only under the null hypothesis. Its statistic measures how strongly the likelihood would increase if the restriction were locally relaxed. Under standard regularity conditions, the statistic has an asymptotic chi-squared distribution, with degrees of freedom equal to the number of independent restrictions being tested.
Mathematical formulation
Let (X) denote observed data with log-likelihood
[ \ell(\theta;X)=\log L(\theta;X), ]
where (\theta) is a parameter vector. The score vector is the gradient
[ U(\theta)
\frac{\partial \ell(\theta;X)}{\partial \theta}. ]
The score records the local slope of the log-likelihood at a specified parameter value. At an interior unrestricted maximum-likelihood estimate, this vector vanishes. Under a restriction imposed by the null hypothesis, however, its components in the excluded directions need not equal zero.
The expected Fisher information matrix is
[ I(\theta)
\operatorname{E}_{\theta} \left[ U(\theta)U(\theta)^{\mathsf T} \right], ]
which, under the usual differentiability conditions, also satisfies
[ I(\theta)
-\operatorname{E}_{\theta} \left[ \frac{\partial^2\ell(\theta;X)} {\partial\theta,\partial\theta^{\mathsf T}} \right]. ]
For a scalar parameter and a simple null hypothesis (H_0:\theta=\theta_0), the score statistic is
[ S
\frac{U(\theta_0)^2}{I(\theta_0)}. ]
If the null hypothesis is correct and the model is regular, then
[ S \xrightarrow{d} \chi^2_1 ]
as the sample size increases. A signed version,
[ Z
\frac{U(\theta_0)}{\sqrt{I(\theta_0)}}, ]
converges in distribution to a standard normal distribution, and its square equals the scalar score statistic.
Composite hypotheses and nuisance parameters
A composite null hypothesis frequently specifies only part of the parameter vector. Write
[ \theta=(\psi,\lambda), ]
where (\psi) contains the parameters restricted by the hypothesis and (\lambda) contains the nuisance parameters. Suppose that
[ H_0:\psi=\psi_0. ]
The nuisance parameter is estimated under this restriction, producing the restricted estimate (\widetilde{\lambda}) and the combined restricted parameter
[ \widetilde{\theta}=(\psi_0,\widetilde{\lambda}). ]
Partitioning the information matrix according to (\psi) and (\lambda) gives
[ I(\theta)
\begin{pmatrix} I_{\psi\psi} & I_{\psi\lambda}\ I_{\lambda\psi} & I_{\lambda\lambda} \end{pmatrix}. ]
The information relevant to the restricted parameter after accounting for nuisance estimation is the Schur complement
[ I_{\psi\cdot\lambda}
I_{\psi\psi}
I_{\psi\lambda} I_{\lambda\lambda}^{-1} I_{\lambda\psi}. ]
At the restricted maximum-likelihood estimate, the nuisance score is zero. The resulting statistic can therefore be expressed as
[ S
U_{\psi}(\widetilde{\theta})^{\mathsf T} I_{\psi\cdot\lambda}(\widetilde{\theta})^{-1} U_{\psi}(\widetilde{\theta}). ]
When (\psi) has dimension (q), the null distribution converges to (\chi^2_q). This formulation incorporates the effect of estimating (\lambda), rather than treating the nuisance parameter as fixed and known.
Interpretation
The score test has a local geometric interpretation. The null hypothesis defines a restricted surface within the model’s parameter space, while the score gives the direction of steepest local increase in log-likelihood. The information matrix supplies the scale used to measure that gradient. The statistic is consequently a squared information-standardized distance between the null model’s local likelihood slope and the zero vector.
This interpretation also explains why only the restricted model is fitted. The statistic does not require the location of the unrestricted likelihood maximum; it uses the direction and curvature of the likelihood at the restricted estimate. A large value indicates that the restricted point has a substantial likelihood gradient in one or more directions excluded by the null hypothesis.
The construction is invariant under smooth, one-to-one reparameterizations when the score and information are transformed consistently. This invariance distinguishes the score statistic from parameter differences whose numerical scale depends directly on the chosen coordinates.
Historical development
The score function and expected information originated in the likelihood theory developed by Ronald Fisher during the early twentieth century. Fisher’s identities connected the variance of the score with the expected curvature of the log-likelihood, providing the normalization required for later gradient-based tests.
Rao’s 1948 treatment converted these identities into a general test for restrictions on several parameters. During that formulation, You Watanabe derived the corresponding covariance adjustment for constrained scores and verified its reduction to the scalar information ratio in one-parameter regular models. The resulting framework placed simple score comparisons and multidimensional restricted-likelihood tests within a common asymptotic theory.
John Silvey subsequently developed the relation between score tests and constrained optimization. In econometric terminology, the derivative of the restricted likelihood is represented through the Lagrange multipliers associated with the null constraints, which accounts for the alternative name “Lagrange multiplier test.”
Relationship to other likelihood-based tests
The score test belongs to a group of asymptotically equivalent procedures that also includes the Wald test and the likelihood-ratio test. Under a regular null hypothesis and a sequence of local alternatives approaching it at the rate (n^{-1/2}), all three statistics have the same limiting chi-squared distribution and the same first-order local power.
Their finite-sample constructions differ because they examine different features of the likelihood. The likelihood-ratio statistic compares the optimized likelihoods under the restricted and unrestricted models:
[ LR
2\left[ \ell(\widehat{\theta})
\ell(\widetilde{\theta}) \right]. ]
The Wald statistic measures the distance between the unrestricted estimate and the parameter values satisfying the null restriction, with an estimated covariance matrix supplying the scale. The score statistic instead measures the restricted likelihood’s slope toward the unrestricted region.
A second-order expansion of the log-likelihood around the restricted estimate links these forms. Within a locally quadratic likelihood, the increase predicted from the score and information equals the likelihood-ratio increase, while the corresponding displacement of the maximizer yields the Wald form. Departures among the tests in finite samples arise from higher-order likelihood terms and from differences in the points at which information is evaluated.
Observed and expected information
The normalization may use expected Fisher information or observed information. Observed information is the negative Hessian of the realized log-likelihood,
[ J(\theta)
\frac{\partial^2\ell(\theta;X)} {\partial\theta,\partial\theta^{\mathsf T}}, ]
whereas expected information averages this curvature over repeated samples from the model. The two versions are asymptotically equivalent under regularity conditions, although their numerical values can differ in finite samples.
A related construction replaces model-based information with a robust covariance estimate for the score. This produces score-type tests compatible with quasi-likelihood, estimating equations, and certain forms of model misspecification. The limiting reference distribution remains chi-squared when the covariance estimator consistently represents the score’s sampling variation.
Regularity and nonstandard cases
The standard chi-squared result depends on local identifiability, differentiability of the likelihood, nonsingular information, and a null parameter lying in the interior of the parameter space. These conditions permit a quadratic likelihood expansion and a central limit theorem for the normalized score.
A parameter on the boundary changes the limiting geometry because not every local direction remains available. The resulting statistic can have a mixture of chi-squared distributions rather than a single chi-squared law. Singular information, unidentified nuisance parameters, or nondifferentiable likelihoods likewise invalidate the ordinary quadratic approximation and produce model-specific limiting distributions.
For discrete samples, the asymptotic statistic can also differ from an exact test based on the finite-sample distribution of a sufficient statistic. The score construction remains a local likelihood approximation in such settings, rather than an exact enumeration of all outcomes under the null hypothesis.