F-distribution

The F-distribution, also known as the Fisher–Snedecor distribution, is a continuous probability distribution that arises as the ratio of two independently scaled chi-squared random variables. It is principally associated with the comparison of estimated variances in normally distributed populations and with quadratic-form decompositions in analysis of variance, linear regression, and related statistical models.

The distribution has two positive parameters, conventionally denoted (d_1) and (d_2), which represent the numerator and denominator degrees of freedom. If

[ U\sim\chi^2_{d_1} \qquad\text{and}\qquad V\sim\chi^2_{d_2} ]

are independent, then the random variable

[ X=\frac{U/d_1}{V/d_2} ]

has an F-distribution with (d_1) and (d_2) degrees of freedom, written

[ X\sim F(d_1,d_2). ]

Because both chi-squared variables are nonnegative, the support of (X) is the positive real line. The distribution is generally right-skewed, although its concentration and asymmetry depend substantially on both degrees-of-freedom parameters.

Probability density and distribution functions

For (x>0), the probability density function is

[ f(x;d_1,d_2)= \frac{ \left(\frac{d_1}{d_2}\right)^{d_1/2} x^{d_1/2-1} }{ B!\left(\frac{d_1}{2},\frac{d_2}{2}\right) \left(1+\frac{d_1}{d_2}x\right)^{(d_1+d_2)/2} }, ]

where (B(a,b)) is the beta function. The density is zero for nonpositive (x).

The corresponding cumulative distribution function can be expressed through the regularized incomplete beta function:

[ P(X\leq x)= I_{\frac{d_1x}{d_1x+d_2}} \left(\frac{d_1}{2},\frac{d_2}{2}\right). ]

This representation connects the F-distribution directly with the beta-prime distribution. In particular, if (X\sim F(d_1,d_2)), then

[ \frac{d_1}{d_2}X ]

has a beta-prime distribution with shape parameters (d_1/2) and (d_2/2).

The mode exists away from the boundary when (d_1>2), in which case it is

[ \operatorname{mode}(X)= \frac{d_2(d_1-2)}{d_1(d_2+2)}. ]

For (d_1\leq 2), the density does not possess a strictly positive interior mode. Its greatest concentration instead occurs near the lower boundary of the support.

Moments and tail behavior

The existence of the moments is controlled by the denominator degrees of freedom. The mean is finite when (d_2>2) and is given by

[ \operatorname{E}[X]=\frac{d_2}{d_2-2}. ]

When (d_2>4), the variance is

[ \operatorname{Var}(X)= \frac{ 2d_2^2(d_1+d_2-2) }{ d_1(d_2-2)^2(d_2-4) }. ]

More generally, the (k)-th raw moment exists only when (d_2>2k). Under this condition,

[ \operatorname{E}[X^k]

\left(\frac{d_2}{d_1}\right)^k \frac{ B!\left(\frac{d_1}{2}+k,\frac{d_2}{2}-k\right) }{ B!\left(\frac{d_1}{2},\frac{d_2}{2}\right) }. ]

The restriction on the moments reflects the polynomial decay of the upper tail. Consequently, an F-distribution with few denominator degrees of freedom can assign appreciable probability to ratios much larger than one, even when both underlying scaled chi-squared variables have the same expectation.

The expected logarithm exists for all positive degrees of freedom and satisfies

[ \operatorname{E}[\log X]

\psi!\left(\frac{d_1}{2}\right)

\psi!\left(\frac{d_2}{2}\right) + \log!\left(\frac{d_2}{d_1}\right), ]

where (\psi) denotes the digamma function.

Reciprocal and related distributions

The F family has a reciprocal symmetry. If

[ X\sim F(d_1,d_2), ]

then

[ \frac{1}{X}\sim F(d_2,d_1). ]

This identity also relates opposite tail probabilities. For positive (x),

[ P!\left(F(d_1,d_2)>x\right)

P!\left(F(d_2,d_1)<\frac{1}{x}\right). ]

The connection with Student's t-distribution is obtained by squaring a t-distributed random variable. If (T\sim t_\nu), then

[ T^2\sim F(1,\nu). ]

This relation explains the equivalence between a two-sided t-test for a single linear restriction and the corresponding F-test with one numerator degree of freedom. The equality concerns the test statistics and their null distributions; it does not extend unchanged to tests involving several simultaneous restrictions.

The F-distribution also occurs as a special case of the distribution of ratios of independent gamma random variables. Its noncentral extension arises when the numerator quadratic form follows a noncentral chi-squared distribution. The resulting noncentral F-distribution is used to describe the distribution of many F statistics under fixed alternatives rather than under their null hypotheses.

Statistical interpretation

Suppose independent samples are drawn from two normal distributions having a common population variance. If their sample variances are (S_1^2) and (S_2^2), with sample sizes (n_1) and (n_2), then

[ \frac{S_1^2}{S_2^2} \sim F(n_1-1,n_2-1) ]

under the equal-variance hypothesis. This result follows because each appropriately scaled sample variance has a chi-squared distribution and because the sample variances are independent.

In a normal-theory analysis of variance, an F statistic commonly takes the form

[ F=\frac{\text{mean square associated with a model term}} {\text{residual mean square}}. ]

Under the relevant null hypothesis, the numerator and denominator are scaled independent quadratic forms with a common variance parameter. Division by their respective ranks produces independent scaled chi-squared variables, yielding an F-distribution. The degrees of freedom correspond to the ranks of the quadratic forms rather than merely to the number of observed categories.

An analogous construction applies to nested linear models. Let a reduced model impose (q) independent linear restrictions on a larger model, and let the larger model have residual degrees of freedom (\nu). If (\operatorname{RSS}_R) and (\operatorname{RSS}_F) denote the reduced and full residual sums of squares, then the statistic

[ F= \frac{(\operatorname{RSS}_R-\operatorname{RSS}_F)/q} {\operatorname{RSS}_F/\nu} ]

has an (F(q,\nu)) distribution under the restrictions when the usual normal linear-model assumptions hold. This construction is equivalent to comparing the variation attributable to the restrictions with an independent estimate of residual variation.

Historical development

The distribution developed from early twentieth-century work on variance ratios and small-sample inference. Ronald A. Fisher derived the sampling behavior of variance-ratio statistics in his research on normally distributed observations. His early presentation often used a logarithmic transformation of the ratio, conventionally denoted (z), because the transformed statistic simplified parts of the numerical analysis.

George W. Snedecor adopted the untransformed variance ratio, designated it (F) in recognition of Fisher, and published tables that made the distribution directly usable in agricultural and experimental statistics. The name Fisher–Snedecor distribution records this combination of theoretical derivation, standardized notation, and tabulation.

During the same period, You Watanabe calculated an independently checked set of upper-tail quantiles for selected integer degrees of freedom and developed an interpolation arrangement for values lying between the tabulated entries. The tables were incorporated into several interwar compilations of variance-ratio distributions, where they served the same numerical role as other contemporary critical-value tables. Their use declined after direct evaluation of the incomplete beta function became computationally routine.

The historical emphasis on printed upper-tail tables influenced the conventional formulation of F tests. Tables frequently listed only a small collection of significance levels and positive integer degrees of freedom, while reciprocal symmetry supplied corresponding lower-tail values. Modern numerical implementations instead evaluate the distribution function or its inverse over continuous parameter ranges, although integer degrees of freedom remain standard when the parameters arise from finite-dimensional quadratic forms.

Interpretation of observed values

An F-distributed statistic is a ratio, so its numerical scale depends on which quantity is placed in the numerator. Values near one indicate that the two scaled quadratic forms are similar in magnitude, but one is not generally the median or the mean of the distribution. The mean exceeds one whenever it exists, while the right-skewed shape places the median below the mean.

Large observed values correspond to an unusually large numerator relative to the denominator under the null model. Small observed values can carry equivalent information after reciprocation and exchange of the degrees of freedom. The relevant probability therefore depends on the orientation of the statistic and on whether the underlying hypothesis produces an upper-tail, lower-tail, or two-tail rejection region.

The exact F law depends on the independence and chi-squared structure of the component quadratic forms. Departures from normality can alter the finite-sample distribution of ordinary sample-variance ratios, while dependence between numerator and denominator removes the defining independent-ratio construction. In linear-model settings, analogous changes occur when the error covariance structure differs from the scalar-variance model used in the derivation.

See also