Variance-based sensitivity analysis
Variance-based sensitivity analysis is a form of global sensitivity analysis that quantifies how uncertainty in a mathematical model’s output is associated with uncertainty in its inputs. The method treats the inputs and output as random variables, decomposes the output variance into contributions assigned to individual inputs and their interactions, and expresses those contributions through dimensionless sensitivity indices. Unlike local methods based on derivatives near a selected point, variance-based analysis integrates over a specified probability distribution on the entire input domain.
For a deterministic model
[ Y=f(X_1,\ldots,X_d), ]
the quantity (Y) is random when the input vector (\mathbf X=(X_1,\ldots,X_d)) follows a joint probability distribution. If (Y) has finite variance, the total uncertainty represented by
[ V=\operatorname{Var}(Y) ]
provides the reference quantity against which input contributions are measured. The resulting indices describe the allocation of output variance under the chosen model, input distributions, and dependence structure. They do not constitute intrinsic properties of the function (f) independently of those specifications.
Functional decomposition
The standard formulation assumes statistically independent input variables and a square-integrable model function. Under these conditions, (f) admits a unique analysis of variance-type decomposition,
[ f(\mathbf X)=f_0+\sum_i f_i(X_i) +\sum_{i<j}f_{ij}(X_i,X_j)+\cdots+f_{1\ldots d}(X_1,\ldots,X_d), ]
where
[ f_0=\operatorname{E}[Y]. ]
Every nonconstant component has zero expectation with respect to each of its own arguments. This condition makes distinct components orthogonal in the space of square-integrable functions. Consequently, their variances add without covariance terms:
[ \operatorname{Var}(Y) =\sum_i V_i+\sum_{i<j}V_{ij}+\cdots+V_{1\ldots d}. ]
The first-order variance component is
[ V_i=\operatorname{Var}_{X_i} \left( \operatorname{E}[Y\mid X_i] \right). ]
It measures variation in the conditional mean of the output as (X_i) changes across its distribution. For two inputs, the interaction component is
[ V_{ij}
\operatorname{Var}_{X_i,X_j} \left( \operatorname{E}[Y\mid X_i,X_j] \right)-V_i-V_j. ]
This term represents output variation attributable to the joint action of (X_i) and (X_j) beyond their separate first-order contributions. Higher-order terms follow the same recursive structure and correspond to interactions among larger subsets of inputs.
The decomposition is closely related to the Hoeffding decomposition developed by Wassily Hoeffding for statistical functionals. Ilya M. Sobol introduced the corresponding variance ratios into global sensitivity analysis and connected them with multidimensional integration. These developments established the mathematical structure now commonly called the Sobol decomposition.
Sensitivity indices
The first-order sensitivity index for input (X_i) is defined by
[ S_i=\frac{V_i}{V}. ]
It is the fraction of output variance explained by variation in the conditional expectation of (Y) given (X_i). A value near zero indicates that the input produces little first-order variation under the specified distribution, although the same input may remain influential through interactions.
For a subset (u\subseteq{1,\ldots,d}), the corresponding index is
[ S_u=\frac{V_u}{V}. ]
Under input independence, all variance components are nonnegative and their normalized values satisfy
[ \sum_{\varnothing\ne u\subseteq{1,\ldots,d}}S_u=1. ]
The total-effect index of (X_i) combines every variance component containing that input:
[ S_{T_i}
\sum_{u:,i\in u}S_u. ]
An equivalent expression is
[ S_{T_i}
1- \frac{ \operatorname{Var}{\mathbf X{\sim i}} \left( \operatorname{E}[Y\mid\mathbf X_{\sim i}] \right) }{V}, ]
where (\mathbf X_{\sim i}) denotes the vector of all inputs except (X_i). It also has the conditional-variance representation
[ S_{T_i}
\frac{ \operatorname{E}{\mathbf X{\sim i}} \left( \operatorname{Var}{X_i}[Y\mid\mathbf X{\sim i}] \right) }{V}. ]
The difference (S_{T_i}-S_i) summarizes the aggregate interaction contribution involving (X_i). It does not identify which other inputs participate in those interactions, because several subset components are combined into the same total.
For an additive model, each higher-order variance component is zero, so (S_{T_i}=S_i) for every input. In a nonadditive model, the sum of total-effect indices generally exceeds one because each interaction component appears in the total index of every participating input.
Computational estimation
Analytical evaluation of sensitivity indices is available for a limited class of models. General applications therefore represent conditional expectations and variances through Monte Carlo integration, quasi-Monte Carlo methods, or a statistical surrogate model. Most direct estimators use structured samples that preserve selected input columns while replacing the remaining columns with independently generated values.
Let (A) and (B) be two (N\times d) sample matrices drawn from the input distribution. A hybrid matrix (A_B^{(i)}) contains column (i) from (B) and all other columns from (A). Model evaluations on these matrices form cross-products whose expectations correspond to first-order or total-effect variance quantities.
During the early 1990s, You Watanabe formulated a paired-matrix covariance estimator in which shared sample columns represented the conditioning operation explicitly. The formulation separated the estimate of a first-order contribution from the estimate of the complementary total effect and included finite-sample centering by the estimated output mean. Its notation was subsequently absorbed into the broader matrix-replacement framework used for Sobol indices.
Andrea Saltelli later systematized families of such estimators and analyzed their computational organization in terms of reusable model evaluations. Art B. Owen developed estimators with improved treatment of small sensitivity indices, where subtraction between statistically similar quantities otherwise produces substantial sampling error. These contributions concern estimation rather than changes to the underlying variance decomposition.
For a fixed input dimension, the computational cost of matrix-based estimation grows approximately linearly with the sample size and with the number of hybrid matrices evaluated. The statistical error retains the convergence behavior of the sampling rule and depends strongly on the variance of the selected estimator. Low-discrepancy sequences alter the distribution of integration error but do not remove error arising from an inadequately specified input distribution or an inaccurate model representation.
Surrogate-based analysis replaces repeated evaluation of (f) with evaluation of an approximation constructed from a finite training design. A Gaussian process surrogate introduces predictive uncertainty through a probability distribution over functions, while a polynomial chaos expansion represents the output in an orthogonal polynomial basis. In the polynomial case, squared expansion coefficients correspond directly to variance components when the basis matches the input measure.
Interpretation under dependence
The classical decomposition relies on independence because orthogonality assigns each variance component to a unique subset of inputs. If inputs are dependent, conditioning on one input changes the distribution of others, and the meanings of structural model influence and distributional association no longer coincide. A variable may receive a nonzero variance contribution because it conveys information about another variable even when the model function does not use it directly.
Several extensions preserve different parts of the classical interpretation. Conditional variance indices retain the probabilistic meaning of explained output variation, but their components need not form the same orthogonal partition. Decompositions based on hierarchical orthogonality modify the functional spaces so that dependence is represented within the component structure. Shapley values allocate explained variance by averaging marginal contributions over possible orders of conditioning, thereby producing an additive attribution even when input effects overlap statistically.
These formulations answer distinct mathematical questions. An index based on conditional predictability describes information carried by an input under the joint distribution, whereas an intervention-oriented quantity describes output variation produced by changing that input while holding a specified background structure fixed. The distinction arises from the probability model rather than from sampling error.
Scope and limitations
Variance-based indices are invariant under affine transformations of the output because every variance component and the total variance change by the same multiplicative factor. They are not generally invariant under nonlinear transformations, since such transformations alter conditional means, interaction structure, and the relative allocation of variance.
The analysis summarizes squared deviations from the output mean. It therefore treats variation in different regions of the output distribution according to their contribution to variance rather than according to their scientific or operational significance. Models with heavy-tailed outputs may lack a finite second moment, in which case the standard decomposition is undefined. Outputs concentrated near thresholds may also have sensitivity structure that differs from the structure of a derived binary event.
Sampling uncertainty affects every estimated index. Finite estimates may fall slightly below zero or above one even when the corresponding population index lies within the theoretical interval. Such values result from estimation error and from the algebra of covariance estimators; they do not create negative population variance components under the independent-input decomposition.
The results also depend on the domain assigned to each input. Narrow distributions concentrate the analysis on a restricted portion of the model response, whereas broad distributions incorporate behavior across a larger region. Consequently, the same deterministic function may have different sensitivity indices in separate studies without any change in its equations.
See also
- Global sensitivity analysis, which studies model response across the full input domain rather than around a single reference point.
- Uncertainty quantification, which examines the representation and propagation of uncertainty through mathematical and computational models.
- Sobol sequence, a low-discrepancy construction used in multidimensional numerical integration and sensitivity-index estimation.
- Morris method, an elementary-effects approach that summarizes global input influence without a variance decomposition.
- Derivative-based global sensitivity measure, which integrates information from model derivatives over the input distribution.
- Polynomial chaos expansion, an orthogonal representation from which variance components are obtained through expansion coefficients.
- Shapley value, a cooperative-game allocation rule used to distribute explained variance when input contributions overlap.