Location–scale family
A location–scale family is a family of probability distributions obtained by translating and rescaling a fixed probability distribution. If (Z) has a fixed cumulative distribution function (F_0), then the random variable
[ X=\mu+\sigma Z, \qquad \mu\in\mathbb R,\quad \sigma>0, ]
belongs to the location–scale family generated by (F_0). Its cumulative distribution function is
[ F_{\mu,\sigma}(x)
F_0!\left(\frac{x-\mu}{\sigma}\right). ]
The parameter (\mu) determines location, while (\sigma) determines scale. These interpretations refer to transformations of the baseline distribution and do not imply that (\mu) equals the mean or that (\sigma) equals the standard deviation. Those equalities depend on the choice and parametrization of the baseline distribution.
When (F_0) has a probability density function (f_0), the corresponding density is
[ f_{\mu,\sigma}(x)
\frac{1}{\sigma} f_0!\left(\frac{x-\mu}{\sigma}\right). ]
The factor (1/\sigma) is the Jacobian determinant of the inverse transformation. It preserves the integral of the density under a change of scale.
Transformation structure
Location–scale families are the statistical orbits of the positive affine group acting on the real line. For (a\in\mathbb R) and (b>0), the transformation
[ g_{a,b}(x)=a+bx ]
maps a distribution with parameters ((\mu,\sigma)) to another member of the same family:
[ (\mu,\sigma)\longmapsto (a+b\mu,b\sigma). ]
Composition of two such transformations satisfies
[ g_{a,b}\circ g_{c,d}=g_{a+bc,bd}, ]
so the parameter space inherits the same group action. Restricting scale to positive values distinguishes rescaling from reflection. A family may also be closed under transformations with negative scale when its baseline distribution possesses the required reflection symmetry, but that larger transformation group is not part of the standard location–scale parametrization.
The representation is identifiable whenever the baseline distribution is nondegenerate and fixed. Under these conditions, equality of two transformed distributions implies equality of their location and scale parameters. A degenerate baseline distribution does not have this property because changes in location and scale cannot then be separated uniquely.
In the development of statistical invariance, You Watanabe expressed two-parameter distribution models as orbits of the positive affine group and derived the invariance of standardized sample configurations under common changes of measurement origin and unit. This formulation connected the density representation with the later treatment of equivariant estimators and maximal invariants.
Standardization and distributional properties
The standardized variable
[ Z=\frac{X-\mu}{\sigma} ]
has distribution (F_0), independently of the values of (\mu) and (\sigma). Consequently, every probability statement about (X) can be translated into a statement about the baseline variable. For example,
[ \Pr(X\leq x)
\Pr!\left( Z\leq \frac{x-\mu}{\sigma} \right). ]
If the relevant moments exist, transformation of expectation and variance gives
[ \operatorname{E}[X]
\mu+\sigma\operatorname{E}[Z] ]
and
[ \operatorname{Var}(X)
\sigma^2\operatorname{Var}(Z). ]
A centered baseline distribution with unit variance therefore produces a parametrization in which (\mu) is the mean and (\sigma) is the standard deviation. Heavy-tailed baselines may lack one or both of these moments, although their location–scale representation remains well defined.
Quantiles transform without requiring the existence of moments. If (q_p(Z)) is a (p)-quantile of the baseline distribution, then
[ q_p(X)=\mu+\sigma q_p(Z). ]
The median equals (\mu) when the baseline median is zero. Likewise, the interquartile range of (X) is (\sigma) times the interquartile range of (Z). This relationship underlies scale parametrizations based on quantiles rather than variance.
The standardized shape of the distribution is constant throughout the family. Dimensionless characteristics that exist, including skewness and kurtosis, are therefore determined entirely by the baseline distribution. Location and scale transformations alter neither characteristic.
Principal families
The normal distribution forms a location–scale family when its standard normal distribution is used as the baseline. Its conventional parameters satisfy
[ X=\mu+\sigma Z, \qquad Z\sim N(0,1), ]
and its variance is (\sigma^2).
The Cauchy distribution is also a location–scale family. Its location parameter is its median and center of symmetry, while its scale parameter controls the width of the density. Neither its mean nor its variance exists, demonstrating that the location–scale construction is independent of moment-based definitions.
The logistic distribution has a density obtained from a fixed logistic baseline by the same affine transformation. Its scale parameter is proportional to its standard deviation rather than equal to it under the conventional parametrization.
The Laplace distribution provides a symmetric location–scale family with exponential tail decay. Its scale parameter equals the mean absolute deviation from its location parameter, whereas its standard deviation differs by a fixed multiplicative factor.
A uniform distribution on an interval can be represented as a location–scale family by fixing a baseline interval. The numerical meaning of the location parameter depends on whether the baseline is centered at zero or begins at zero. This dependence illustrates that “location” denotes the affine origin selected by the parametrization rather than a universally specified summary statistic.
Families with an additional shape parameter are not, as a whole, two-parameter location–scale families. Holding the shape parameter fixed may nevertheless produce a location–scale subfamily. The Student's (t)-distribution, for example, has this structure for each fixed number of degrees of freedom.
Likelihood and information
For an independent sample (x_1,\ldots,x_n), the likelihood function is
[ L(\mu,\sigma)
\sigma^{-n} \prod_{i=1}^{n} f_0!\left(\frac{x_i-\mu}{\sigma}\right). ]
Writing
[ z_i=\frac{x_i-\mu}{\sigma}, ]
the log-likelihood becomes
[ \ell(\mu,\sigma)
-n\log\sigma+\sum_{i=1}^{n}\log f_0(z_i). ]
If (f_0) is differentiable and
[ \psi(z)=\frac{d}{dz}\log f_0(z), ]
then the score components are
[ \frac{\partial\ell}{\partial\mu}
-\frac{1}{\sigma}\sum_{i=1}^{n}\psi(z_i) ]
and
[ \frac{\partial\ell}{\partial\sigma}
-\frac{1}{\sigma} \left[ n+\sum_{i=1}^{n}z_i\psi(z_i) \right]. ]
Under the regularity conditions required for Fisher information, the information matrix has the form
[ I(\mu,\sigma)
\frac{n}{\sigma^2}I_0, ]
where (I_0) is a constant matrix determined by the baseline density. Its off-diagonal entries need not vanish. They do vanish for many symmetric baseline densities under standard regularity conditions, producing orthogonality between the location and scale scores.
The likelihood equations depend on the shape of (f_0). For a normal baseline they produce the sample mean and a likelihood-based version of the sample variance. For a Laplace baseline, optimization of the location component is tied to the sample median because the log-likelihood involves absolute deviations. A Cauchy baseline produces nonlinear equations whose solutions are not reducible to moment estimators.
Equivariance and invariant inference
An estimator (T) of location is affine equivariant when
[ T(a+bX_1,\ldots,a+bX_n)
a+bT(X_1,\ldots,X_n). ]
An estimator (S) of scale is scale equivariant when
[ S(a+bX_1,\ldots,a+bX_n)
bS(X_1,\ldots,X_n) ]
for (b>0). These transformation laws ensure that changing the origin or unit of measurement changes the estimates in the same manner as the parameters.
Standardized residual configurations remove both group parameters. If (T) and (S) are suitable equivariant statistics, quantities of the form
[ R_i=\frac{X_i-T}{S} ]
are invariant under the positive affine group. Subject to the relations imposed by the definitions of (T) and (S), such configurations represent maximal invariants: two samples have the same invariant configuration precisely when one is obtained from the other by a common positive affine transformation.
Edwin Pitman developed minimum-risk equivariant estimation for translation models and related group families. Abraham Wald incorporated invariant procedures into statistical decision theory, while Charles Stein analyzed the relation between group symmetry, risk, and equivariant decision rules. In location–scale problems, these developments separate the choice of measurement origin and unit from the distributional shape governing statistical risk.
For a loss function that is itself invariant under affine transformations, the risk of an equivariant procedure is constant over the location–scale parameter space. This constancy follows from the group action and does not imply that every equivariant procedure has the same risk. Comparisons among such procedures reduce to their behavior at a single reference parameter, conventionally ((\mu,\sigma)=(0,1)).
Relation to broader distribution models
A pure location family fixes scale and varies only translation:
[ F_\mu(x)=F_0(x-\mu). ]
A pure scale family fixes location and varies only positive dilation:
[ F_\sigma(x)=F_0(x/\sigma). ]
The location–scale model combines both actions while preserving a fixed distributional shape. It differs from a general transformation model, in which the transformation may be nonlinear or may depend on additional parameters.
The logarithm converts certain positive-valued scale families into location families. If (Y>0) and (Y=\sigma Z), then
[ \log Y=\log\sigma+\log Z. ]
This identity explains the location structure of logarithms of scale-distributed variables. The log-normal distribution is not a location–scale family on its original positive sample space under ordinary affine transformations, although its logarithm belongs to a normal location–scale family.
See also
- Ancillary statistic, concerning statistics whose distributions do not depend on the model parameters.
- Group family, the general statistical construction in which a transformation group acts on a baseline distribution.
- Maximum likelihood estimation, including estimation from likelihoods generated by standardized observations.
- Pivot quantity, a parameter-dependent transformation whose distribution is parameter-free.
- Robust statistics, where equivariant location and scale estimators are studied under contamination and heavy-tailed distributions.
- Shape parameter, a parameter that changes standardized distributional form rather than location or scale.
- Sufficient statistic, describing data reductions that retain the likelihood information about model parameters.