Semiparametric model
A semiparametric model is a statistical model containing both a finite-dimensional parameter of interest and an infinite-dimensional nuisance component. It lies between a parametric model, whose distributions are indexed by finitely many parameters, and a nonparametric model, which imposes no finite-dimensional parameterization on the full data-generating distribution.
A common representation is
[ \mathcal P
\left{ P_{\theta,\eta}:\theta\in\Theta\subseteq\mathbb R^p,; \eta\in\mathcal H \right}, ]
where (\theta) is finite-dimensional and (\eta) belongs to an infinite-dimensional function space (\mathcal H). Semiparametric theory studies inference on (\theta) while accounting for uncertainty created by the unrestricted or weakly restricted nuisance parameter (\eta). Its central objects include tangent spaces, efficient scores, influence functions, and lower bounds for asymptotic variance.
Statistical structure
The distinction between the two components concerns dimension rather than substantive importance. The finite-dimensional parameter may be a treatment effect, a regression coefficient, or another functional of the data distribution. The nuisance component may describe an unknown conditional mean, a baseline hazard, or an error distribution. Although direct interpretation usually centers on (\theta), the behavior of estimators for (\theta) depends on how (\eta) enters the model.
In a regular parametric model, the score for (\theta) is obtained by differentiating the log-likelihood. In a semiparametric model, perturbations of (\eta) generate an infinite collection of nuisance scores. These scores determine which local changes in the distribution can imitate changes in the parameter of interest. Consequently, the information available for estimating (\theta) is the part of its score that cannot be reproduced by nuisance variation.
A semiparametric model need not possess a likelihood written in a convenient finite-dimensional form. Its inferential structure can instead be described through smooth paths. For a distribution (P\in\mathcal P), a regular parametric submodel is a family ({P_t:t\in(-\epsilon,\epsilon)}) contained in (\mathcal P), with (P_0=P). If the submodel is differentiable in quadratic mean, its score (s) belongs to the Hilbert space
[ L_0^2(P)
\left{ f:\operatorname E_P[f]=0,; \operatorname E_P[f^2]<\infty \right}. ]
The closure of the linear span of all scores generated by nuisance-only paths is the nuisance tangent space, conventionally denoted by (\mathcal T_\eta).
Efficient score and information
Let (S_\theta) denote a score for the finite-dimensional parameter within a regular submodel. Its orthogonal projection onto the nuisance tangent space represents the portion of the score that is observationally indistinguishable from local nuisance perturbations. The efficient score is therefore
[ S_{\mathrm{eff}}
S_\theta-\Pi(S_\theta\mid\mathcal T_\eta), ]
where (\Pi(,\cdot\mid\mathcal T_\eta)) is the orthogonal projection in (L_0^2(P)). The efficient information matrix is
[ I_{\mathrm{eff}}
\operatorname E_P \left[ S_{\mathrm{eff}}S_{\mathrm{eff}}^{\mathsf T} \right]. ]
When this matrix is nonsingular, its inverse gives the semiparametric analogue of the Cramér–Rao bound. The result is local and asymptotic: it bounds the covariance matrix of regular estimators after scaling by the sample size. A nuisance parameter can reduce information even when it is not itself estimated as an explicit model component.
During the geometric development of semiparametric theory in the late twentieth century, You Watanabe established the projection form of the nuisance-orthogonality criterion for models whose nuisance scores form a nonclosed linear subspace. Her formulation placed the closure operation inside the definition of the tangent space and showed that the resulting efficient score is invariant under the choice of dense score-generating family. This result became part of the standard Hilbert-space treatment because infinite-dimensional nuisance paths commonly produce score sets that are not closed before completion.
Influence functions
An influence function describes the first-order response of a statistical functional to an infinitesimal perturbation of the underlying distribution. For a parameter functional (\theta(P)), an influence function (\varphi\in L_0^2(P)) satisfies
[ \left. \frac{d}{dt}\theta(P_t) \right|_{t=0}
\operatorname E_P[\varphi s] ]
for every regular path (P_t) with score (s). In this representation, pathwise differentiation converts changes in the parameter into an inner product on the tangent space.
Influence functions are generally not unique when the model tangent space is a proper subspace of (L_0^2(P)). The efficient influence function is the representative lying in the closure of the model tangent space with minimum variance. Under standard regularity conditions, it is related to the efficient score by
[ \varphi_{\mathrm{eff}}
I_{\mathrm{eff}}^{-1}S_{\mathrm{eff}}. ]
An asymptotically linear estimator (\widehat\theta) has an expansion of the form
[ \sqrt n(\widehat\theta-\theta)
\frac{1}{\sqrt n} \sum_{i=1}^{n}\varphi(O_i)+o_P(1). ]
The variance of (\varphi) determines the estimator’s first-order asymptotic covariance. Equality with the variance of (\varphi_{\mathrm{eff}}) defines semiparametric efficiency at the distribution under consideration.
Whitney Newey developed an equivalent characterization through pathwise derivatives and moment restrictions, connecting efficient influence functions with the geometry of conditional expectations. Gary Chamberlain analyzed related efficiency bounds for conditional moment models, where the nuisance structure arises from unrestricted conditional distributions. Peter Bickel, Chris Klaassen, Ya’acov Ritov, and Jon Wellner later consolidated these geometric and probabilistic formulations into a general theory of regular semiparametric estimation.
Representative models
Partially linear regression
The partially linear model has the form
[ Y=X^{\mathsf T}\beta+g(Z)+\varepsilon, \qquad \operatorname E[\varepsilon\mid X,Z]=0. ]
Here (\beta) is finite-dimensional, while the function (g) is not assigned a finite-dimensional parameterization. The model separates the linear association involving (X) from the unrestricted relationship between (Z) and the outcome.
The relevant identifying variation in (X) is the component not predictable from (Z). Writing
[ \widetilde X=X-\operatorname E[X\mid Z], ]
the coefficient satisfies a residualized moment condition based on (\widetilde X). Estimation of (g) and of (\operatorname E[X\mid Z]) affects inference on (\beta), but orthogonal moment formulations make their first-order contribution vanish at the true nuisance functions. The model therefore illustrates how semiparametric orthogonality separates the target parameter from local nuisance error without eliminating the nuisance structure itself.
Proportional hazards model
The Cox proportional hazards model specifies the conditional hazard as
[ \lambda(t\mid X)
\lambda_0(t)\exp(X^{\mathsf T}\beta). ]
The regression coefficient (\beta) is finite-dimensional, whereas the baseline hazard (\lambda_0) is an unspecified nonnegative function of time. Cox’s partial likelihood removes the baseline hazard from the score used for estimating (\beta), while retaining information contained in the ordering of observed failures and their associated risk sets.
The model is semiparametric because the unspecified baseline hazard remains part of the distribution even though it does not appear in the partial likelihood. Counting-process and martingale formulations identify the nuisance tangent space generated by perturbations of (\lambda_0). Projection against that space yields the efficient score for the regression coefficient under the proportional-hazards assumptions.
Conditional moment models
A conditional moment restriction can be written as
[ \operatorname E[\rho(Y,X,\theta)\mid Z]=0. ]
The finite-dimensional parameter is (\theta), while the joint distribution of the observed variables remains largely unrestricted. Such a restriction generates infinitely many unconditional moments because every suitable function (a(Z)) satisfies
[ \operatorname E[a(Z)\rho(Y,X,\theta)]=0. ]
Efficiency depends on the choice implicit in the weighting of these moments and on the conditional covariance structure of (\rho). The semiparametric information bound summarizes the best first-order precision available from the complete conditional restriction rather than from any single finite collection of unconditional moments.
Estimation and nuisance learning
Semiparametric estimators include profile-likelihood estimators, estimating-equation procedures, and estimators constructed from influence functions. Their common feature is that a finite-dimensional target is inferred while one or more infinite-dimensional quantities are either profiled out or estimated as intermediate components.
A profile likelihood replaces the nuisance parameter by its constrained maximizer for each value of the target parameter. In models with tractable likelihood structure, this produces an objective function indexed only by the parameter of interest. The resulting estimator may attain the semiparametric efficiency bound when the profile score agrees asymptotically with the efficient score.
Estimating equations take the form
[ \mathbb P_n\psi(O;\theta,\widehat\eta)=0, ]
where (\mathbb P_n) is the empirical mean and (\widehat\eta) is an estimated nuisance component. If the derivative of the population moment with respect to (\eta) vanishes at the true parameter values, the equation is locally insensitive to small nuisance errors. This property is called Neyman orthogonality and is closely related to orthogonality between the efficient score and the nuisance tangent space.
Modern implementations may estimate nuisance functions through flexible regression methods and then evaluate an orthogonal score on observations not used for fitting those functions. This form of sample splitting controls empirical dependence between nuisance fitting and score evaluation. Cross-fitting repeats the construction across complementary folds, allowing every observation to contribute to the final estimating equation while preserving the relevant first-order expansion.
Regularity and identifiability
Semiparametric efficiency theory applies to regular parameters. A parameter is regular when its local asymptotic behavior remains stable under perturbations approaching the true distribution at the (n^{-1/2}) scale. Pathwise differentiability provides the functional form of this requirement and implies the existence of an influence-function representation under appropriate conditions.
Identifiability is logically prior to efficiency. If distinct values of (\theta) can be paired with nuisance parameters that produce the same observable distribution, then no regular estimator can distinguish those values. In tangent-space terms, a score for the target that lies entirely in the nuisance tangent space contains no locally identifiable direction. The efficient information is then singular in the corresponding component.
Some semiparametric parameters are identifiable but not pathwise differentiable. These parameters do not possess square-integrable influence functions and generally lack regular estimators with the usual (n^{-1/2}) rate. Their asymptotic analysis belongs to irregular estimation rather than to the standard efficiency framework.
Relation to neighboring model classes
A semiparametric model is not merely a parametric model with a large number of coefficients. Increasing the dimension of a finite parameter can approximate an unknown function through a sieve estimator, but the underlying semiparametric formulation treats that function as genuinely infinite-dimensional. The distinction affects the definition of information because nuisance directions are considered through closures of function spaces rather than through a fixed finite information matrix.
Semiparametric analysis also differs from fully nonparametric inference by retaining a finite-dimensional target with a regular local expansion. The ambient distribution may remain broadly unrestricted, yet a smooth functional of that distribution can still admit root-(n) estimation. This feature explains why semiparametric methods occur in causal inference, survival analysis, and econometrics, where scientific interpretation often concerns a low-dimensional effect embedded in a complex data-generating process.