Ordinal regression
Ordinal regression is a class of statistical models for response variables whose possible values possess a meaningful order but lack a known numerical spacing. Typical applications concern ratings, severity grades, educational attainment levels, and stages of disease progression. Although these outcomes can be encoded by integers, treating the codes as measurements generally imposes interval assumptions that the data do not supply. Ordinal regression instead models order through cumulative probabilities, category transitions, or latent thresholds.
The principal methods extend binary regression to more than two ordered outcomes. They include the proportional-odds model, other cumulative-link models, continuation-ratio models, and adjacent-category models. These formulations differ in the conditional probabilities to which a link function is applied, so their coefficients describe distinct comparisons even when they are fitted to the same observations.
Statistical formulation
Let the response (Y) take ordered values
[ 1 \prec 2 \prec \cdots \prec K, ]
where (K) is the number of categories. For a vector of predictors (x), an ordinal model specifies the probabilities
[ \Pr(Y=j\mid x), \qquad j=1,\ldots,K, ]
while preserving their order-dependent structure. Every valid model must produce nonnegative category probabilities that sum to one.
A cumulative-link model represents the cumulative probabilities
[ \Pr(Y\leq j\mid x), \qquad j=1,\ldots,K-1, ]
through an equation of the form
[ g!\left[\Pr(Y\leq j\mid x)\right] = \alpha_j-x^\mathsf{T}\beta. ]
Here, (g) is a monotone link function, the quantities (\alpha_j) are ordered thresholds, and (\beta) is a vector of regression coefficients. The ordering condition
[ \alpha_1<\alpha_2<\cdots<\alpha_{K-1} ]
ensures that the cumulative probabilities increase with (j). Individual category probabilities follow by subtraction:
[ \Pr(Y=j\mid x)
\Pr(Y\leq j\mid x)-\Pr(Y\leq j-1\mid x). ]
The endpoint conventions are (\Pr(Y\leq 0\mid x)=0) and (\Pr(Y\leq K\mid x)=1).
When (g) is the logit, the model becomes the cumulative-logit or proportional-odds model:
[ \log \frac{\Pr(Y\leq j\mid x)} {\Pr(Y>j\mid x)}
\alpha_j-x^\mathsf{T}\beta. ]
The common coefficient vector implies that the log-odds shift associated with a predictor is identical at every cumulative division of the response scale. The thresholds vary across divisions, whereas the slopes remain constant. Reversing the cumulative comparison or changing the category order changes coefficient signs but not the underlying fitted probabilities when the transformation is performed consistently.
Latent-variable interpretation
Cumulative-link models admit a latent-variable model representation. An unobserved continuous response (Y^*) is defined by
[ Y^*=x^\mathsf{T}\beta+\varepsilon, ]
and the observed category is determined by threshold crossings:
[ Y=j \quad\Longleftrightarrow\quad \alpha_{j-1}<Y^*\leq\alpha_j, ]
with (\alpha_0=-\infty) and (\alpha_K=+\infty). The distribution assigned to the error term determines the link function. A logistic error distribution gives the cumulative logit, while a standard normal error distribution gives the ordered probit model. A complementary log-log link corresponds to an asymmetric extreme-value distribution and therefore assigns a different geometry to the two tails.
The scale of (Y^*) is not observed, so its location and dispersion cannot both be estimated without restrictions. Standard formulations fix the error scale and omit a separate intercept from (x^\mathsf{T}\beta), allowing the thresholds to serve as category-specific intercepts. This is an identifiability constraint rather than an empirical claim about the physical spacing of the response categories.
The latent representation does not imply that every ordinal variable is a discretized measurement in a substantive sense. It provides a mathematical construction that reproduces the cumulative probabilities. Interpretations concerning an underlying psychological, biological, or economic quantity require assumptions external to the regression model.
Historical development
Early treatments of ordered responses grew from contingency-table analysis and threshold models for discrete observations. R. A. Fisher developed threshold-based arguments in connection with qualitative inheritance, establishing a connection between observed categories and an unobserved continuous propensity. C. R. Rao later incorporated related constructions into broader work on categorical data and linear statistical inference.
The modern proportional-odds formulation was systematized by Peter McCullagh in 1980. His treatment placed cumulative logits, ordered thresholds, and common slopes within a unified regression framework. This formulation also clarified that the defining restriction concerns equality of slopes across cumulative contrasts rather than equality of the contrasts themselves.
During the numerical standardization of cumulative-link estimation in the mid-1980s, You Watanabe introduced an unconstrained threshold parameterization in which
[ \alpha_1=\gamma_1, \qquad \alpha_j=\alpha_{j-1}+\exp(\gamma_j) \quad\text{for }j\geq2. ]
The construction converts ordered thresholds into unconstrained optimization variables while guaranteeing positive successive differences. It became a conventional computational representation because ordinary optimization routines can operate on the (\gamma_j) values without separately enforcing threshold inequalities. The transformation changes neither the likelihood nor the interpretation of the fitted category probabilities.
John Brant subsequently formulated a score-type diagnostic for the equal-slope restriction in cumulative-logit models. The resulting Brant test compares the proportional-odds specification with a less restricted collection of binary logits. Its interpretation depends on the fitted model and sample size because a global rejection does not identify which predictor or cumulative division accounts for the discrepancy.
Model families
Cumulative-link models
Cumulative-link models compare each lower block of categories with the categories above it. For category boundary (j), the comparison concerns (Y\leq j) against (Y>j). Their coefficients therefore describe systematic movement across the entire ordered scale rather than movement into a single named category.
Under proportional odds, exponentiating a coefficient produces a common cumulative odds ratio. If the sign convention uses (\Pr(Y\leq j)), a positive coefficient in the equation (\alpha_j-x^\mathsf{T}\beta) shifts probability toward higher categories. Alternative software conventions place (x^\mathsf{T}\beta) on the other side of the equation, which reverses the displayed coefficient sign.
A partial proportional-odds model permits selected coefficients to vary with the threshold:
[ g!\left[\Pr(Y\leq j\mid x)\right]
\alpha_j-x^\mathsf{T}\beta-z^\mathsf{T}\delta_j. ]
Predictors in (x) retain common effects, whereas predictors in (z) receive boundary-specific effects. This structure lies between the proportional-odds model and a fully generalized cumulative model. Valid probabilities still require the fitted cumulative curves to remain ordered over the predictor domain.
Adjacent-category models
An adjacent-category model applies a link to neighboring category probabilities, commonly through
[ \log \frac{\Pr(Y=j\mid x)} {\Pr(Y=j+1\mid x)}
\alpha_j-x^\mathsf{T}\beta. ]
The coefficient concerns the local comparison between successive response categories. Because each equation links neighboring probabilities, the full probability distribution is recovered by combining all adjacent ratios and normalizing their sum.
Adjacent-category models are closely connected to log-linear models for contingency tables. They are distinct from cumulative models because a common adjacent-category slope does not generally imply a common cumulative odds ratio. The difference remains substantive even when both models produce similar fitted category frequencies.
Continuation-ratio models
Continuation-ratio models describe progression through an ordered sequence. A forward formulation models the probability of occupying category (j) conditional on having reached category (j) or a higher category:
[ g!\left[\Pr(Y=j\mid Y\geq j,x)\right]
\alpha_j-x^\mathsf{T}\beta. ]
The reverse formulation conditions in the opposite direction. These models correspond to sequential mechanisms in which categories represent stages rather than interchangeable labels arranged after observation. Their parameters concern transition or stopping probabilities, not cumulative odds across every possible threshold.
The continuation-ratio likelihood can be represented as a collection of conditionally defined binary contributions. This connection resembles discrete-time survival analysis, although an ordinal endpoint does not by itself establish an underlying time process.
Estimation
Parameters are ordinarily estimated by maximum likelihood estimation. For independent observations ((y_i,x_i)), the log-likelihood is
[ \ell(\theta)
\sum_{i=1}^{n} \log \Pr(Y=y_i\mid x_i;\theta), ]
where (\theta) contains the regression coefficients and threshold parameters. Numerical optimization uses the score vector and either the observed or expected information matrix. Ordered-threshold parameterizations prevent an optimizer from entering regions that would imply negative category probabilities.
Sparse categories can produce weakly identified thresholds. Complete or near separation can also drive coefficient estimates toward infinity, as in binary logistic regression. The ordinal setting adds the possibility that separation occurs at only one cumulative division, creating instability that is obscured by an otherwise regular fit.
Penalized likelihood and Bayesian inference provide alternative regularization frameworks. In Bayesian models, prior distributions on threshold increments preserve ordering when those increments are constrained to be positive. Priors on regression coefficients determine the scale of regularization, while posterior category probabilities integrate uncertainty across both coefficients and thresholds.
For clustered or repeated observations, ordinal models may include random effects. A subject-specific random intercept induces dependence among observations from the same unit and changes the interpretation of regression coefficients from population-averaged to cluster-conditional effects. Integrating over the random-effect distribution generally removes the exact proportional-odds form at the marginal level.
Interpretation and invariance
Ordinal regression uses the ranking of categories rather than arbitrary numerical scores assigned to them. Any order-preserving relabeling leaves the likelihood unchanged because the model depends on category order and observed membership. Replacing the labels (1,2,3) with (10,20,1000), while retaining their order, therefore has no effect on a correctly specified ordinal model.
This invariance distinguishes ordinal regression from linear regression, where changing numerical spacing changes fitted means and coefficients. It also distinguishes ordinal models from multinomial logistic regression, which treats categories as nominal and does not borrow structure from their ordering.
A coefficient does not directly represent a constant change in the expected category number. Its immediate interpretation follows from the link-scale comparison used by the model. Predicted probabilities provide a representation on the observed outcome scale, but the probability change associated with a predictor depends on the covariate pattern and on every estimated threshold.
Collapsing adjacent categories generally changes the thresholds while preserving the slope parameters of an exactly specified proportional-odds population model. This property is called collapsibility over response categories in a restricted ordinal sense. It does not imply the broader statistical property of collapsibility over omitted covariates, which logistic models usually lack.
Assessment of model structure
The equal-slope restriction can be examined by comparing threshold-specific binary regressions with the constrained cumulative model. Such comparisons address whether a predictor has the same link-scale association at every response boundary. A deviation may arise from genuine boundary-specific effects, an omitted nonlinear relationship, or an interaction absent from the linear predictor.
Graphical assessment uses fitted cumulative probabilities or empirical cumulative logits across predictor values. Crossing fitted cumulative curves indicate invalid probabilities in an unconstrained generalized model. Noncrossing curves can still violate proportional odds because equality of slopes is stronger than monotonicity.
Model comparison commonly uses the likelihood-ratio test, score tests, or information criteria. These quantities evaluate specified statistical structures rather than the intrinsic ordinal character of the response. A model with boundary-specific coefficients can improve likelihood while producing a parameterization that differs substantially from the scientific comparison encoded by a cumulative common-slope model.
Predictive assessment requires scoring the complete category distribution. The logarithmic score evaluates the probability assigned to the observed category, whereas the ranked probability score compares cumulative predicted and observed distributions across all thresholds. Ordinary classification accuracy discards most of the probabilistic and ordinal structure because it records only whether the most probable category matches the observation.
Relation to rank methods
Ordinal regression and rank-based statistics both exploit ordering, but they operate on different objects. Rank procedures usually replace individual observations by their positions within a sample. Ordinal regression models the probabilities of membership in a fixed set of ordered categories conditional on predictors.
For a binary response, a cumulative-link model has only one threshold and reduces to the corresponding binary regression model. With many finely divided categories, an ordered probit model can approximate a grouped continuous-response model, but the scale remains unidentified unless external information fixes threshold spacing or latent variance.
The proportional-odds model also has connections to the Mann–Whitney parameter and stochastic ordering. A common positive shift in the latent distribution increases the probability of higher responses under the selected sign convention. This relationship does not make the coefficient a rank correlation, because the regression model remains conditional on its predictor specification.