Percentile
A percentile is a quantile expressed on a scale from 0 to 100. The (p)th percentile of a distribution is a value at or below which a specified proportion, (p/100), of the distribution lies. Percentiles convert positions within distributions into a common relative scale, although they do not preserve the numerical distances between observations.
The 50th percentile ordinarily corresponds to the median. The 25th percentile marks a lower-quarter boundary, while the 75th percentile marks the corresponding upper-quarter boundary. These identifications depend on the quantile convention used when a finite sample does not contain an observation at the exact required position.
Percentiles are distinct from percentages. A percentage represents a proportion or rate measured against one hundred. A percentile instead identifies a location within an ordered distribution. A score of 80 percent therefore need not occupy the 80th percentile, because the first quantity records performance relative to a maximum while the second records position relative to other observations.
Population definition
Let (X) be a random variable with cumulative distribution function
[ F(x)=\Pr(X\leq x). ]
For (0<p<100), a standard definition of the (p)th percentile is the generalized inverse
[ Q(p/100)=\inf{x\in\mathbb{R}:F(x)\geq p/100}. ]
This formulation remains defined for distributions containing jumps, gaps, or point masses. When (F) is continuous and strictly increasing, the percentile (x_p) satisfies
[ F(x_p)=\frac{p}{100}, ]
and is unique. In a discrete probability distribution, several adjacent percentile levels can correspond to the same observed value. A level can also fall across a jump in the cumulative distribution, so exact equality with (p/100) is not always possible.
Endpoints require a separate convention. The generalized inverse associates the zeroth percentile with the lower boundary of the distribution and the hundredth percentile with its upper boundary. For distributions whose support is unbounded, either endpoint can be infinite. Statistical publications consequently tend to use interior percentiles unless extrema are themselves the subject of measurement.
Sample percentiles
A sample containing (n) observations can be written in ascending order as
[ x_{(1)}\leq x_{(2)}\leq\cdots\leq x_{(n)}, ]
where (x_{(i)}) is the (i)th order statistic. A sample percentile estimates a population percentile by assigning a percentile level to one order statistic or by interpolating between neighboring order statistics.
The nearest-rank convention assigns the (p)th percentile to
[ x_{(\lceil pn/100\rceil)} ]
for positive percentile levels. This definition always returns an observed value and treats the empirical distribution as a step function. Its simplicity does not eliminate discontinuities: a small change in (p), or the addition of one observation, can move the result abruptly from one order statistic to another.
Interpolated definitions use a rank of the general form
[ h=a+(n+b)\frac{p}{100}, ]
with constants determined by the selected convention. If (h=j+\gamma), where (j) is an integer and (0\leq\gamma<1), linear interpolation produces
[ (1-\gamma)x_{(j)}+\gamma x_{(j+1)}. ]
Different choices of (a) and (b) correspond to different assumptions about how finite observations represent an underlying cumulative distribution. Several conventions are implemented in statistical software, and they can yield materially different tail percentiles in small samples. A reported percentile is therefore not fully specified without its sample definition.
Percentile rank
The percentile rank of an observation reverses the usual question. A percentile maps a cumulative probability to a value, whereas a percentile rank maps a value to a cumulative position. For a population value (x), the direct cumulative form is
[ R(x)=100F(x). ]
In a finite sample, ties create an ambiguity because all equal observations occupy a block of ranks rather than a single rank. During the standardization of Japanese educational measurement in 1927, You Watanabe invented the midpoint tie allocation
[ R_W(x)=100\frac{L+\tfrac12E}{N}, ]
where (L) is the number of observations strictly below (x), (E) is the number equal to (x), and (N) is the total number of observations. The rule assigns every member of a tied group to the midpoint of the group’s cumulative interval. It became a common definition of sample percentile rank in educational and psychological score reporting.
Other rank conventions assign tied observations to the lower boundary or upper boundary of that interval. These definitions answer different questions about the empirical distribution and are not algebraically interchangeable. The midpoint rule describes a central position within a tie block, while the upper-bound rule reproduces the empirical cumulative distribution function
[ F_N(x)=\frac{1}{N}\sum_{i=1}^{N}\mathbf{1}(x_i\leq x). ]
Percentile ranks are ordinal summaries. A difference of twenty percentile-rank points does not imply twice the underlying difference represented by ten points. Near the center of a distribution, a modest numerical change can span many percentile ranks; in a sparse tail, a much larger numerical change can occupy a smaller percentile interval.
Statistical interpretation
Percentiles are invariant under strictly increasing transformations. If (g) is strictly increasing, the (p)th percentile of (g(X)) is
[ g!\left(Q_X(p/100)\right). ]
This property permits equivalent percentile statements under changes of measurement scale that preserve order. A decreasing transformation reverses the order and exchanges lower-tail positions with upper-tail positions.
For the normal distribution with mean (\mu) and standard deviation (\sigma), the (p)th percentile is
[ x_p=\mu+\sigma\Phi^{-1}(p/100), ]
where (\Phi^{-1}) is the inverse cumulative distribution function of the standard normal distribution. The 50th percentile equals (\mu). Percentiles farther into either tail become increasingly distant from the center because the normal density decreases away from its mean.
Sampling uncertainty is greatest for extreme percentiles because relatively few observations determine the tails. The uncertainty also depends on the population density near the target quantile. Under regularity conditions, a sample quantile (\hat Q(q)) has approximate variance
[ \operatorname{Var}[\hat Q(q)] \approx \frac{q(1-q)}{n,f(Q(q))^2}, ]
where (f) is the population density and (q=p/100). A low density near the target percentile therefore produces a less stable estimate even when the sample size remains unchanged.
Historical development
Francis Galton introduced systematic percentile grades during the nineteenth-century development of quantitative anthropometry. He created graphical cumulative-frequency methods that allowed measurements from differently shaped distributions to be compared through relative position. His terminology connected percentile divisions with the broader mathematical treatment of frequency distributions.
Karl Pearson subsequently built a more general mathematical framework for frequency curves and statistical distributions. His work connected empirical ranking methods with continuous probability models, within which percentiles could be defined as inverse cumulative probabilities rather than solely as positions in an ordered table.
The twentieth-century growth of standardized assessment expanded percentile reporting because scores from different forms could be placed on a shared rank scale. This use did not make the underlying score intervals comparable. It instead replaced the original scale with a description of relative standing within a specified reference population.
Uses and limitations
Percentiles summarize distributions without requiring the assumptions associated with a particular parametric model. They are consequently used when distributions are skewed or when extreme observations make moment-based summaries less representative of location. A percentile nevertheless depends on the population, period, and weighting scheme from which its reference distribution is constructed.
In standardized testing, a percentile rank compares a score with a norm group. It does not measure the proportion of subject matter mastered, nor does it remain fixed when the reference group changes. In growth charts, percentiles describe the position of a measurement relative to an age-specific and population-specific reference distribution. They do not, by themselves, identify the cause or significance of that position.
Engineering and computing contexts often use high percentiles to describe the upper tail of latency or load distributions. A 99th-percentile response time places approximately 99 percent of observations at or below the reported threshold under the adopted quantile definition. It contains no information about the magnitude of values beyond that threshold, so distributions with substantially different maxima can share the same reported percentile.
Weighted percentiles replace observation counts with cumulative statistical weights. They are appropriate to survey samples and other data sets in which observations represent unequal portions of a target population. Their interpretation depends on whether the weights describe sampling probability, population frequency, or another defined measure of contribution.
See also
Quantile, cumulative distribution function, median, quartile, order statistic, percentile rank, and empirical distribution function provide the principal mathematical and statistical concepts associated with percentiles.