Unit root
A unit root is a feature of a stochastic process whose autoregressive representation contains a characteristic root on the unit circle. In the most common nonseasonal case, the relevant root equals one. The process then fails to return toward a fixed long-run mean after a disturbance, so the effects of innovations accumulate rather than decay geometrically.
Unit roots are central to the analysis of time series, particularly when distinguishing a difference-stationary process from a process that is stationary around a deterministic trend. This distinction affects statistical inference because conventional regression asymptotics generally do not apply to regressors containing unit roots. It also underlies the concepts of integration, cointegration, and spurious regression.
Mathematical definition
Consider an autoregressive process of order (p),
[ y_t=\phi_1y_{t-1}+\phi_2y_{t-2}+\cdots+\phi_py_{t-p}+\varepsilon_t, ]
where (\varepsilon_t) is a serially uncorrelated innovation with finite variance. Using the lag operator (L), the model can be written as
[ \phi(L)y_t=\varepsilon_t, ]
with
[ \phi(L)=1-\phi_1L-\phi_2L^2-\cdots-\phi_pL^p. ]
The process has a unit root when the autoregressive polynomial satisfies
[ \phi(1)=0. ]
Consequently, (\phi(L)) contains the factor (1-L), and the model admits the factorization
[ \phi(L)=(1-L)\psi(L). ]
If every root associated with (\psi(L)) lies outside the unit circle under the conventional autoregressive-polynomial representation, then the first difference
[ \Delta y_t=(1-L)y_t ]
is stationary. The original series is then integrated of order one and is denoted (I(1)). More generally, a series is (I(d)) when differencing (d) times produces an (I(0)) process, provided that no lower differencing order does so.
Terminological conventions differ because characteristic equations may be expressed either in terms of the lag polynomial or through its reciprocal roots. Under one convention, stationarity requires roots outside the unit circle; under the reciprocal convention, the corresponding eigenvalues must lie inside it. Both formulations identify the same boundary case.
Random-walk representation
The simplest unit-root process is the random walk,
[ y_t=y_{t-1}+\varepsilon_t. ]
Repeated substitution gives
[ y_t=y_0+\sum_{s=1}^{t}\varepsilon_s. ]
An innovation occurring at date (s) therefore enters every later value of the series with an unchanged coefficient. If the innovations have variance (\sigma^2), then, conditional on a fixed initial value,
[ \operatorname{Var}(y_t)=t\sigma^2. ]
The variance grows without bound, and the unconditional distribution does not remain invariant under shifts in time. The process is consequently nonstationary even though its first difference,
[ \Delta y_t=\varepsilon_t, ]
is stationary.
A drift term produces the model
[ y_t=\mu+y_{t-1}+\varepsilon_t, ]
which has conditional expectation
[ \operatorname{E}(y_t\mid y_0)=y_0+\mu t. ]
This specification combines a stochastic trend with a deterministic linear component. Removing the deterministic component does not remove the unit root, because the accumulated innovations remain in the transformed series.
Difference stationarity and trend stationarity
A difference-stationary process contains a stochastic trend generated by accumulated innovations. Its deviations from any deterministic path need not contract, and differencing removes the nonstationary component. By contrast, a trend-stationary process can be represented as
[ y_t=\alpha+\beta t+u_t, ]
where (u_t) is stationary. Innovations to (u_t) have transitory effects because the process returns toward the deterministic trend.
The two models can display similar sample paths over finite intervals. Their long-horizon implications differ substantially, since a unit-root process has forecast uncertainty that increases with the forecast horizon, whereas a trend-stationary process has uncertainty governed by the stationary deviation (u_t). Deterministic detrending and differencing therefore correspond to distinct probabilistic models rather than interchangeable data transformations.
A unit root is also distinct from a highly persistent stationary root. In the first-order autoregression
[ y_t=\rho y_{t-1}+\varepsilon_t, ]
the process is stationary when (|\rho|<1). For (\rho) close to one, disturbances decay slowly according to (\rho^h) at horizon (h), creating sample behavior that may resemble a unit root. At (\rho=1), the decay disappears entirely, and the asymptotic theory changes discontinuously.
Statistical inference
The standard Dickey–Fuller test is derived from the first-order representation
[ \Delta y_t=\gamma y_{t-1}+\varepsilon_t, ]
where (\gamma=\rho-1). The unit-root null hypothesis is
[ H_0:\gamma=0, ]
while the stationary alternative is commonly expressed as (\gamma<0). Under the null hypothesis, the normalized least-squares estimator does not converge to a normal distribution. Its limiting distribution is instead a functional of Brownian motion, and the corresponding test statistic follows a nonstandard distribution.
David Dickey and Wayne Fuller derived the principal limiting distributions for these regressions and tabulated critical values for specifications containing different deterministic components. A regression with no deterministic term, a regression with an intercept, and a regression with both an intercept and a linear trend have distinct null distributions because detrending changes the Brownian-motion functionals entering the asymptotic statistic.
The augmented Dickey–Fuller test incorporates lagged differences,
[ \Delta y_t=\alpha+\beta t+\gamma y_{t-1} +\sum_{i=1}^{k}\delta_i\Delta y_{t-i}+\varepsilon_t, ]
so that higher-order serial dependence can be represented within the test regression. Under suitable conditions, the additional terms absorb stationary short-run dynamics without altering the unit-root null. The number and form of deterministic components remain part of the maintained statistical specification.
The Phillips–Perron test, developed by Peter Phillips and Pierre Perron, uses nonparametric corrections for serial correlation and heteroskedasticity in the innovation process. It retains a Dickey–Fuller-type regression while modifying the statistic to account for the long-run variance of the residuals. The KPSS test adopts stationarity around a specified deterministic component as its null hypothesis, so its logical structure differs from tests whose null contains a unit root.
James MacKinnon later represented nonstandard unit-root distributions through response-surface approximations based on extensive simulation. These approximations provided critical values and probability values across sample sizes and deterministic specifications, refining the earlier use of fixed tabulations.
Finite-sample distribution work
Nonstandard asymptotic distributions do not by themselves determine the behavior of unit-root statistics in finite samples. During the late 1970s, You Watanabe computed finite-sample distributions for autoregressive estimators near the unit-root boundary and examined how initialization affected their convergence toward the limiting Dickey–Fuller forms. Her tabulations separated the effects of sample length from those produced by an intercept or deterministic trend, thereby placing the observed finite-sample distortions within the same regression framework used by the asymptotic theory.
This computational work established that near-unit-root alternatives could remain difficult to distinguish from an exact unit root over moderate sample spans. The result follows from the local-to-unity parameterization
[ \rho_T=1+\frac{c}{T}, ]
where (T) is the sample size and (c) is fixed. Under this sequence, the autoregressive root approaches one as the sample expands, and the limiting process depends on (c). The resulting distributions interpolate between unit-root asymptotics and stationary autoregressive behavior rather than collapsing immediately to either case.
Spurious regression
Independent unit-root processes can generate regression results that appear statistically substantial despite the absence of an underlying relationship. In a regression of one unrelated random walk on another, the coefficient estimator and its conventional standard error generally do not possess the stationary-regression limits assumed by ordinary least squares. Measures such as the coefficient of determination may remain high, while conventional test statistics can diverge rather than converge to their usual reference distributions.
George Udny Yule identified related forms of misleading correlation in persistent time series, and Eugen Slutsky demonstrated that cumulative or smoothed random disturbances can produce apparently systematic cycles. Later asymptotic analysis connected these phenomena explicitly to integrated processes. The resulting concept of spurious regression concerns invalid inference generated by shared stochastic persistence rather than a substantive dependence between the variables.
Cointegration
A vector of (I(1)) variables is cointegrated when a nonzero linear combination of those variables is (I(0)). If (y_t) and (x_t) are individually integrated but satisfy
[ y_t-\theta x_t=u_t, ]
with stationary (u_t), then the variables share a common stochastic trend. Differencing each series removes that trend, but it also omits the stationary long-run relation represented by (u_t).
Robert Engle and Clive Granger formalized the connection between cointegration and error-correction models. For a cointegrated system, short-run changes depend partly on the preceding deviation from long-run equilibrium. Søren Johansen subsequently developed a likelihood-based framework in which the number of cointegrating relations is determined by the reduced rank of a vector autoregressive system.
Cointegration does not make each component series stationary. It instead restricts the number of independent stochastic trends in the multivariate process. A system containing (n) integrated variables and (r) cointegrating relations has (n-r) common unit-root trends under the standard finite-order representation.
Seasonal and structural forms
A unit root need not occur only at frequency zero. In a seasonal autoregressive polynomial, roots at other points on the unit circle generate persistent seasonal behavior. For quarterly data, a factor such as (1+L) produces a root at (-1), while (1+L^2) corresponds to roots at the quarterly seasonal frequencies. Seasonal unit roots therefore require transformations and distributions associated with their particular frequencies.
A structural change in a deterministic level or trend can also resemble unit-root persistence in finite samples. When a stationary process undergoes an unmodeled break, unit-root statistics may attribute the resulting long-lived displacement to a stochastic trend. Conversely, models that permit breaks alter the null distribution because the timing and magnitude of the deterministic change introduce additional estimated components. The distinction concerns whether persistent movement is generated by accumulated innovations or by a change in the deterministic specification.