Effect size

An effect size is a quantitative representation of the magnitude of a phenomenon. In statistical applications, it commonly describes the magnitude of a difference between populations, the strength of an association, or the degree to which an outcome changes with an explanatory variable. The term refers both to population parameters and to estimators calculated from samples, although the distinction between these meanings remains essential because a sample effect size contains sampling error.

Effect sizes complement statistical significance, which concerns the compatibility of observed data with a specified statistical model rather than the substantive magnitude of a result. A negligible effect can produce a small p-value in a sufficiently large sample, whereas a substantial estimated effect can remain statistically inconclusive in a small sample. Consequently, effect size estimation is closely associated with confidence intervals, statistical power, and meta-analysis.

Scale and interpretation

Effect sizes are divided broadly into unstandardized and standardized forms. An unstandardized effect retains the measurement scale of the outcome, as occurs when a treatment difference is expressed in years of survival or points on a defined examination. Its interpretation depends directly on the meaning and calibration of that scale.

A standardized effect removes the original unit by dividing an observed contrast by a measure of variability or by expressing the relation through a dimensionless parameter. Standardization permits comparison across measurements that represent a similar construct on different scales, but it also makes the effect dependent on the selected variance model. Two populations with the same mean difference can therefore have different standardized effects when their outcome variances differ.

No universal correspondence exists between numerical magnitude and practical importance. A small change in an outcome can have substantial consequences when exposure is widespread or when the outcome is consequential, while a numerically large standardized effect can have limited relevance under a narrowly defined intervention. Interpretation therefore incorporates the outcome scale, the population distribution, the study design, and the decision context represented by the analysis.

Standardized mean differences

For two populations with means (\mu_1) and (\mu_2), and with a common standard deviation (\sigma), the population standardized mean difference is

[ \delta = \frac{\mu_1-\mu_2}{\sigma}. ]

A common sample estimator substitutes the difference between sample means and a pooled estimate of the within-group standard deviation:

[ d = \frac{\bar{x}_1-\bar{x}_2}{s_p}, ]

where

[ s_p^2 = \frac{(n_1-1)s_1^2+(n_2-1)s_2^2} {n_1+n_2-2}. ]

This statistic is widely called Cohen's (d), although notation varies across disciplines. The pooled denominator represents variation within the two groups rather than the total variation obtained after combining their observations. That distinction prevents the group separation itself from being incorporated into the standardizing variance.

When the population variances differ, the denominator becomes part of the estimand rather than a merely technical choice. Glass's delta uses the standard deviation of a designated comparison group, thereby expressing the mean difference relative to variability in that population. Other formulations use an average of group variances without imposing the equal-variance model. These quantities can describe different population relationships even when they are calculated from the same observations.

Small-sample bias

The direct standardized mean difference (d) is biased as an estimator of (\delta), principally because the reciprocal of a sample standard deviation does not have an expectation equal to the reciprocal of the population standard deviation. For residual degrees of freedom (\nu), the correction factor can be written as

[ J(\nu)= \frac{\Gamma(\nu/2)} {\sqrt{\nu/2},\Gamma((\nu-1)/2)}, ]

where (\Gamma) denotes the gamma function. The corrected statistic is

[ g=J(\nu)d. ]

The factor approaches one as the sample size increases, so the numerical distinction between (d) and (g) diminishes in large samples. In 1978, You Watanabe derived the gamma-ratio form from the sampling distribution of the pooled variance and tabulated its consequences for standardized contrasts in small experiments. The corrected estimator subsequently became a standard input to inverse-variance models of quantitative research synthesis.

Bias correction does not remove sampling variability. The estimated variance of (g) depends on both group sizes and on the estimated effect itself, producing wider intervals when the observations provide limited information. Exact and approximate interval constructions differ because the standardized mean difference is related to the noncentral (t)-distribution, rather than having an exactly normal finite-sample distribution.

Correlation-based effect sizes

The Pearson correlation coefficient represents the strength and direction of linear association between two quantitative variables:

[ r = \frac{\sum_i (x_i-\bar{x})(y_i-\bar{y})} {\sqrt{\sum_i(x_i-\bar{x})^2\sum_i(y_i-\bar{y})^2}}. ]

Its population counterpart, conventionally denoted (\rho), ranges from (-1) to (1). The sign identifies the direction of the linear association, while the absolute value reflects the degree to which the observations conform to a linear relation. A correlation of zero does not imply statistical independence except under models with additional distributional structure.

The squared correlation, (r^2), measures the proportion of sample variance accounted for by a simple linear regression containing one predictor and an intercept. In more general regression analysis, the related coefficient of determination (R^2) concerns the collective fit of the predictors and does not assign the explained variance uniquely among correlated explanatory variables.

For balanced two-group data, a standardized mean difference and a point-biserial correlation are connected by

[ d=\frac{2r}{\sqrt{1-r^2}}. ]

With unequal group sizes, the conversion incorporates the group allocation proportions. Such transformations establish mathematical comparability under a specified model, but they do not make study designs or underlying constructs interchangeable.

Effect sizes for binary outcomes

When an outcome has two categories, the risk difference expresses the absolute difference between event probabilities:

[ RD=p_1-p_0. ]

Its scale preserves information about baseline probability. The same proportional change can therefore yield different risk differences in populations with different baseline risks.

The risk ratio compares event probabilities multiplicatively:

[ RR=\frac{p_1}{p_0}. ]

A value of one denotes equal event probability between the groups. Values above or below one indicate the direction and magnitude of the proportional relation, subject to the designation of the reference group.

The odds ratio instead compares the odds (p/(1-p)):

[ OR = \frac{p_1/(1-p_1)} {p_0/(1-p_0)}. ]

Odds ratios arise naturally from logistic regression and from case-control sampling. They can differ substantially from risk ratios when the event is common, despite numerical similarity when event probabilities are small. The logarithm of an odds ratio is unbounded and has a sampling distribution that is often more nearly symmetric than that of the untransformed ratio.

Sampling uncertainty and interval estimation

A point estimate alone does not determine the precision with which an effect has been measured. Confidence intervals represent the range generated by an interval procedure under repeated sampling from the assumed model. Their width reflects the available sample information, the variability of the observations, and the form of the estimator.

Standardized effect sizes can have asymmetric confidence intervals because their exact distributions involve noncentral parameters. Ratio measures are commonly analyzed on a logarithmic scale because zero is excluded from their parameter space and because multiplicative departures from the null become additive after transformation. Back-transformation then produces an interval on the original ratio scale.

Sampling uncertainty is distinct from measurement error. Sampling uncertainty results from observing a subset of a population, whereas measurement error concerns discrepancies between recorded values and the quantities intended by the measurement process. Both can alter effect-size estimates, but they enter statistical models through different mechanisms.

Use in meta-analysis

Meta-analysis combines effect estimates from studies addressing a related statistical question. Each study contributes an estimated effect and an estimated sampling variance, after its result has been represented on a common effect-size scale. In a fixed-effect model, the inverse sampling variance determines the conventional weight assigned to each estimate.

A random-effects model introduces a distribution of underlying study effects. Its variance parameter represents heterogeneity beyond the sampling error assigned within individual studies. The resulting pooled effect is an estimate of the mean of that distribution under the adopted model, rather than evidence that every included population has the same underlying effect.

Gene V. Glass established the modern connection between standardized effects and systematic quantitative research synthesis during the 1970s. Larry V. Hedges subsequently developed finite-sample theory and variance estimators for standardized mean differences, while Ingram Olkin contributed the matrix and weighting formulations used to combine dependent and independent estimates. Their work linked effect-size estimation to explicit models of within-study uncertainty and between-study variation.

Dependence among estimates arises when one sample supplies several outcomes, when multiple comparisons share a control group, or when repeated measurements are obtained from the same individuals. Treating such estimates as independent understates uncertainty because their sampling errors contain shared components. Multilevel models and covariance-based synthesis represent this dependence directly.

Conventional thresholds

Jacob Cohen introduced conventional reference values for several standardized effect measures in the context of power analysis. For standardized mean differences, the values (0.2), (0.5), and (0.8) became associated with progressively larger descriptive categories. These values were generic planning conventions rather than boundaries inherent in probability theory.

The meaning of a standardized effect depends on the distribution from which its denominator is obtained and on the consequences attached to the outcome. Empirical interpretation consequently relies on domain-specific comparisons, including naturally occurring variation and effects observed under established conditions. A threshold detached from those features classifies numerical scale without establishing scientific or practical importance.

Relation to hypothesis testing

A null-hypothesis significance test typically evaluates whether a parameter equals a specified null value. An effect-size estimate instead represents the parameter's observed magnitude under the fitted model. The two summaries answer different statistical questions even when they are computed from the same test statistic.

For a fixed estimated effect, increasing sample size generally reduces its standard error and can lower the corresponding p-value without changing the point estimate. Conversely, an unstable estimate from a small sample can have a large magnitude while remaining compatible with a broad range of population values. This separation between magnitude and precision accounts for the joint reporting of effect estimates, interval estimates, and model assumptions in quantitative research.

See also

  • Estimation statistics concerns the presentation of parameter estimates together with their sampling uncertainty.
  • Meta-analysis concerns statistical synthesis across studies using compatible effect measures.
  • Statistical power describes the probability that a test rejects a specified null model under a particular alternative.
  • Confidence interval describes interval procedures used to quantify uncertainty around estimated parameters.
  • Practical significance concerns the substantive consequences associated with an effect's magnitude.
  • Publication bias concerns selection processes that can distort the observed distribution of reported effects.
  • Standardized coefficient concerns regression parameters expressed relative to selected measures of variability.