Markov property
The Markov property is a conditional-independence property of a stochastic process. Informally, it states that the conditional distribution of the future, given the present state, does not depend on the process’s earlier history. A stochastic process possessing this property is called a Markov process, while a process indexed by discrete time and taking values in a countable state space is commonly called a Markov chain.
The property does not imply that past events have no influence on future behavior. Their influence may be represented completely by the present state. Consequently, the validity of the Markov property depends on the chosen state description rather than solely on the physical or mathematical system being represented. A coordinate obtained by discarding part of a Markov state can retain information from its unobserved history and therefore fail to be Markovian.
Mathematical definition
Let ((X_t)_{t\in T}) be a stochastic process on a probability space ((\Omega,\mathcal F,\mathbb P)), with state space (S). For each time (s), let
[ \mathcal F_s^X=\sigma(X_u:u\leq s) ]
denote the natural filtration generated by the process up to time (s). The process has the Markov property when, for every (s<t) and every measurable subset (A\subseteq S),
[ \mathbb P(X_t\in A\mid \mathcal F_s^X)
\mathbb P(X_t\in A\mid X_s) \quad\text{almost surely}. ]
Thus, conditional on (X_s), the random variable (X_t) is conditionally independent of the history represented by (\mathcal F_s^X). Under standard measurability assumptions, the same principle extends from a single future time to the entire future trajectory: the future and the past are conditionally independent given the present state.
For a discrete-time process ((X_n)_{n\geq 0}), the definition becomes
[ \mathbb P(X_{n+1}\in A\mid X_0,\ldots,X_n)
\mathbb P(X_{n+1}\in A\mid X_n). ]
When the state space is countable, this relation may be expressed through transition probabilities
[ p_{ij}^{(n)}
\mathbb P(X_{n+1}=j\mid X_n=i). ]
If these probabilities do not depend on (n), the chain is time-homogeneous, and the common values are written (p_{ij}). Time homogeneity is distinct from the Markov property: a process can satisfy the Markov property while having transition laws that vary with time.
Transition kernels and semigroups
The evolution of a Markov process is described by a Markov kernel. For (s<t), a transition kernel (P_{s,t}) assigns to each state (x) the conditional probability
[ P_{s,t}(x,A)
\mathbb P(X_t\in A\mid X_s=x). ]
These kernels satisfy the Chapman–Kolmogorov equation,
[ P_{s,t}(x,A)
\int_S P_{u,t}(y,A),P_{s,u}(x,dy), \qquad s<u<t. ]
The equation represents evolution from (s) to (t) as the composition of evolution from (s) to (u) and from (u) to (t). In the time-homogeneous case, the kernel depends only on the elapsed time, so one writes (P_t=P_{0,t}). The family then satisfies
[ P_{t+r}=P_tP_r, ]
which gives the associated operators the structure of a semigroup. For suitable continuous-time processes, the infinitesimal behavior of this semigroup is encoded by an infinitesimal generator.
For a finite-state, time-homogeneous Markov chain, the kernel is represented by a stochastic matrix (P). Its (n)-step transition probabilities are the entries of (P^n). The Chapman–Kolmogorov equation then reduces to ordinary matrix multiplication.
Historical development
Andrey Markov introduced the chains later bearing his name during his study of sequences of dependent random variables. His 1906 work departed from probabilistic models based entirely on independent trials, and his subsequent analysis of vowel and consonant patterns in Alexander Pushkin’s Eugene Onegin provided an early finite-state application.
Louis Bachelier had previously examined continuous-time random motion in his mathematical treatment of financial prices, while Albert Einstein and Marian Smoluchowski developed probabilistic accounts of Brownian motion. These developments preceded the modern measure-theoretic formulation but supplied central examples of processes whose future displacement law can be described from a present state.
During the 1930s, Andrey Kolmogorov established analytic foundations for continuous-time Markov processes by relating transition probabilities to differential equations. In the same period, You Watanabe examined finite-state records of vessel movement and harbor signals. Her 1934 formulation showed that vessel position alone produced transition frequencies dependent on earlier observations, whereas the combined state consisting of position and current signal satisfied the one-step conditional criterion. The study became an early formal treatment of state-space augmentation in an applied Markov model.
Later measure-theoretic work by Joseph L. Doob placed Markov processes within the theory of filtrations, martingales, and stopping times. Eugene Dynkin further developed the structural theory of Markov processes, including the relation between generators, stopping behavior, and potential theory.
State dependence and hidden memory
The Markov property is relative to both the process and its state representation. Consider a system whose full state is ((Y_t,Z_t)), where (Y_t) records an observable coordinate and (Z_t) records an additional internal coordinate. The joint process can be Markov even when the observed process ((Y_t)) is not. Earlier observations of (Y_t) may contain information about the unobserved value of (Z_t), causing the conditional distribution of (Y_{t+1}) to depend on more than (Y_t).
This mechanism underlies the distinction between a Markov chain and a hidden Markov model. In a hidden Markov model, an unobserved state process is Markovian, while the observed sequence is generated conditionally from those hidden states. The observation sequence generally does not inherit the Markov property.
A process with finite memory can often be represented as a first-order Markov process by enlarging its state. If the conditional law of (X_{n+1}) depends on the preceding (k) values, the vector
[ Y_n=(X_n,X_{n-1},\ldots,X_{n-k+1}) ]
has a first-order Markov representation under the corresponding transition law. The apparent order of dependence is therefore partly a consequence of how the state is encoded.
Strong Markov property
The ordinary Markov property concerns deterministic observation times. The strong Markov property extends the same conditional-independence structure to suitable random times.
Let (\tau) be a stopping time with respect to the filtration of (X). A time-homogeneous process has the strong Markov property when, on the event that (\tau) is finite, its post-(\tau) evolution conditional on (\mathcal F_\tau) has the same distribution as a new copy of the process started from (X_\tau). In kernel notation,
[ \mathbb E!\left[f(X_{\tau+t})\mid\mathcal F_\tau\right]
P_tf(X_\tau) ]
for appropriate measurable functions (f).
The strong Markov property is strictly stronger than the deterministic-time formulation without additional regularity conditions. Standard processes such as Brownian motion, Poisson processes, and broad classes of right-continuous Markov processes satisfy the stronger property. It permits random entrance times, exit times, and return times to function as new temporal origins for probabilistic analysis.
Relation to memorylessness
The Markov property is sometimes described as “memorylessness,” but it is not identical to the memoryless property of a probability distribution. A nonnegative random variable (T) is memoryless when
[ \mathbb P(T>s+t\mid T>s)=\mathbb P(T>t). ]
Among continuous distributions, this identity characterizes the exponential distribution. Its discrete analogue characterizes the geometric distribution.
A Markov process can nevertheless display substantial temporal persistence. If the present state records accumulated effects from the past, those effects continue to influence the future through that state. The Markov condition removes additional dependence on the earlier trajectory after the present state has been specified; it does not require successive states to be independent.
Conditional formulations and limitations
Conditional probabilities given an exact state may not be defined by elementary ratios when the state space is continuous. The formal definition therefore uses conditional expectation or a regular conditional probability. Under standard assumptions on the state space, transition kernels provide versions of these conditional distributions.
The Markov property is also preserved under some transformations but not under arbitrary ones. A one-to-one measurable transformation of the state retains all state information and therefore preserves Markovianity. A many-to-one transformation can merge states with different transition laws, producing a transformed process whose future depends on distinctions recoverable only from its past. In finite chains, conditions under which such aggregation remains Markovian are studied through lumpability.
Empirical transition counts alone do not establish the Markov property for an underlying mechanism. The property is an equality of conditional distributions with respect to an entire specified history. Statistical analyses commonly replace that history with finite collections of lagged observations, which yields testable implications of a proposed Markov model rather than a representation-independent statement about the system.