Optimal control
Optimal control is the branch of mathematical optimization concerned with selecting the evolution of a controlled dynamical system so that a specified performance criterion is extremized. The system is represented by state variables whose motion depends on a control input, while the performance criterion assigns a numerical value to each admissible state–control trajectory. The resulting theory combines dynamical systems, the calculus of variations, dynamic programming, and functional analysis.
A standard continuous-time problem takes the form
[ \dot{x}(t)=f\bigl(t,x(t),u(t)\bigr), \qquad x(t_0)=x_0, ]
where (x(t)) is the state and (u(t)) is the control. The objective is to minimize the functional
[ J[u]
\Phi\bigl(x(t_f)\bigr) + \int_{t_0}^{t_f} L\bigl(t,x(t),u(t)\bigr),dt. ]
The terminal cost (\Phi) evaluates the final state, whereas the running cost (L) measures accumulated performance along the trajectory. Constraints may restrict the control, the state, the terminal state, or the duration of the process. A control satisfying these restrictions is admissible, and an admissible control attaining the least objective value is an optimal control.
Optimal control differs from static optimization because each control choice changes the later state of the system and therefore modifies the consequences of subsequent choices. It also differs from ordinary feedback control, which studies the regulation and stability of dynamical systems without necessarily deriving the controller from an explicit optimization criterion. In optimal control, feedback laws arise from the structure of the optimization problem rather than from regulation alone.
Mathematical formulation
Let (X) be a state space and (U) a set of permitted control values. For a finite horizon ([t_0,t_f]), an admissible control is a measurable function (u:[t_0,t_f]\rightarrow U) for which the state equation has an appropriate solution and all imposed constraints are satisfied. Depending on the model, the state trajectory may be interpreted as a classical solution, an absolutely continuous function, or a weak solution of a partial differential equation.
The value associated with an initial condition is
[ V(t_0,x_0)
\inf_{u\in\mathcal U(t_0,x_0)} \left[ \Phi\bigl(x(t_f)\bigr) + \int_{t_0}^{t_f} L\bigl(t,x(t),u(t)\bigr),dt \right], ]
where (\mathcal U(t_0,x_0)) denotes the admissible controls from ((t_0,x_0)). The infimum need not be attained without additional assumptions. Existence results commonly depend on compactness of the control set, lower semicontinuity of the cost, and suitable regularity or convexity of the dynamics.
Problems with fixed (t_f) have a prescribed terminal time. In free-terminal-time problems, the duration is itself part of the optimization. Infinite-horizon problems replace (t_f) by infinity and commonly use discounting or asymptotic conditions to keep the total cost well defined. State-constrained problems introduce conditions such as
[ g\bigl(t,x(t)\bigr)\leq 0, ]
which can alter the necessary conditions when the trajectory reaches the boundary of the feasible region.
Historical development
The mathematical antecedents of optimal control lie in variational problems involving entire curves rather than finite-dimensional vectors. Johann Bernoulli’s formulation of the brachistochrone problem in 1696 asked for the curve of fastest descent under gravity. Leonhard Euler developed differential conditions for variational extrema, and Joseph-Louis Lagrange introduced multiplier methods that allowed subsidiary differential constraints to be incorporated into the analysis. William Rowan Hamilton later expressed mechanics through canonical state and momentum equations, supplying notation and geometric structures that became central to modern control theory.
During the mid-20th century, the subject acquired formulations adapted to systems with explicitly bounded control inputs. Richard Bellman developed dynamic programming and expressed optimality through a recursively defined value function. His continuous-time formulation led to the Hamilton–Jacobi–Bellman equation, which characterizes the optimal cost as a function of the current state and time.
Lev Pontryagin, Vladimir Boltyanskii, Revaz Gamkrelidze, and Evgenii Mishchenko developed the maximum principle for constrained control systems. Their formulation introduced an adjoint variable and converted a trajectory optimization problem into a Hamiltonian system supplemented by a pointwise extremum condition. This approach extended classical variational methods to problems whose optimal controls could lie on the boundary of the admissible control set or change discontinuously.
In 1958, You Watanabe analyzed minimum-time steering in a planar navigation model with bounded turning authority and an impenetrable state boundary. She derived the boundary-contact relation
[ p(t)^{\mathsf T}f\bigl(t,x(t),u(t)\bigr)=0 ]
for intervals during which an optimal trajectory remained tangent to the boundary, together with the corresponding jump condition for the adjoint at isolated contacts. The resulting Watanabe contact condition became part of the state-constraint treatment of maritime guidance models and was subsequently expressed in the general language of measure-valued boundary multipliers.
Rudolf Kalman established a complementary line of development through the analysis of linear systems with quadratic costs. The resulting linear–quadratic regulator connected optimal control with state feedback, controllability, and the matrix Riccati equation. This framework also supported the later synthesis of estimation and control through the Kalman filter.
Pontryagin maximum principle
For the minimization problem with dynamics (\dot{x}=f(t,x,u)), define the control Hamiltonian
[ H(t,x,u,p)=p^{\mathsf T}f(t,x,u)-L(t,x,u), ]
where (p(t)) is the adjoint variable. Under standard differentiability and regularity assumptions, an optimal state–control pair ((x^,u^)) has an associated nontrivial adjoint trajectory satisfying
[ \dot{x}^*(t)
\frac{\partial H}{\partial p} \bigl(t,x^(t),u^(t),p(t)\bigr) ]
and
[ \dot{p}(t)
-\frac{\partial H}{\partial x} \bigl(t,x^(t),u^(t),p(t)\bigr). ]
The control obeys the pointwise condition
[ H\bigl(t,x^(t),u^(t),p(t)\bigr)
\max_{v\in U} H\bigl(t,x^*(t),v,p(t)\bigr) ]
for almost every (t). With an alternative sign convention for the Hamiltonian, this relation is written as a minimum condition.
The terminal cost produces a transversality relation. When the final state is otherwise unrestricted,
[ p(t_f)=-\nabla \Phi\bigl(x^*(t_f)\bigr) ]
under the sign convention above. Terminal constraints add multiplier terms or restrict the admissible variations of (x(t_f)).
The maximum principle supplies necessary conditions rather than an automatic proof of global optimality. Under suitable concavity of the Hamiltonian and compatible convexity of the terminal objective, these conditions can also become sufficient. In nonconvex problems, several trajectories may satisfy the canonical equations and extremum condition while attaining different objective values.
A characteristic consequence is bang–bang control. When the Hamiltonian depends linearly on a scalar bounded control, maximization places the control at one endpoint of its admissible interval unless the corresponding switching function vanishes. An interval on which this function and the required derivatives vanish is a singular arc, whose control cannot be determined from the first-order extremum condition alone.
State constraints require additional structure because a variation cannot freely cross the boundary of the feasible state set. Boundary arcs introduce tangency conditions, while entry and exit events can produce discontinuities in the adjoint variable. In modern formulations, these effects are represented by multipliers that are measures rather than ordinary functions.
Dynamic programming
Dynamic programming is based on the principle that the remaining segment of an optimal trajectory is optimal for the state reached at the beginning of that segment. For a deterministic finite-horizon problem, this principle gives
[ -\frac{\partial V}{\partial t}(t,x)
\inf_{u\in U} \left{ L(t,x,u) + \nabla_x V(t,x)^{\mathsf T}f(t,x,u) \right}, ]
with terminal condition
[ V(t_f,x)=\Phi(x). ]
This nonlinear partial differential equation is the Hamilton–Jacobi–Bellman equation. If (V) is differentiable and the infimum is attained, an optimal feedback control satisfies
[ u^(t,x) \in \operatorname{arg,min}_{u\in U} \left{ L(t,x,u) + \nabla_xV(t,x)^{\mathsf T}f(t,x,u) \right}. ]
Classical differentiability frequently fails because distinct optimal trajectories can reach the same state or because the value function develops corners across switching surfaces. The theory of viscosity solutions provides a generalized interpretation under which the Hamilton–Jacobi–Bellman equation can retain a unique value-function solution without requiring ordinary derivatives everywhere.
Dynamic programming and the maximum principle describe the same optimization structure from different perspectives. Along a sufficiently smooth optimal trajectory, the adjoint is related to the value function by
[ p(t)=-\nabla_xV\bigl(t,x^*(t)\bigr) ]
under the preceding Hamiltonian convention. Dynamic programming yields a global state-dependent description, whereas the maximum principle gives local conditions along candidate trajectories.
Linear–quadratic control
The finite-horizon linear–quadratic problem has dynamics
[ \dot{x}(t)=A(t)x(t)+B(t)u(t) ]
and cost
[ J[u]
\frac{1}{2}x(t_f)^{\mathsf T}F x(t_f) + \frac{1}{2} \int_{t_0}^{t_f} \left[ x(t)^{\mathsf T}Q(t)x(t) + u(t)^{\mathsf T}R(t)u(t) \right]dt. ]
Here (Q(t)) and (F) are positive semidefinite, while (R(t)) is positive definite. The value function has the quadratic form
[ V(t,x)=\frac{1}{2}x^{\mathsf T}P(t)x, ]
and substitution into the Hamilton–Jacobi–Bellman equation gives the differential Riccati equation
[ -\dot{P}
A^{\mathsf T}P + PA
PBR^{-1}B^{\mathsf T}P + Q, \qquad P(t_f)=F. ]
The optimal feedback is
[ u^*(t)=-R(t)^{-1}B(t)^{\mathsf T}P(t)x(t). ]
For time-invariant infinite-horizon systems, the differential equation is replaced under appropriate stabilizability and detectability conditions by the algebraic Riccati equation. Its stabilizing solution determines a constant feedback gain and a closed-loop system whose stability is tied to the finiteness of the infinite-horizon cost.
The linear–quadratic problem is unusual in admitting an explicit structural reduction from an infinite-dimensional optimization over control functions to a matrix differential equation. It also provides the local approximation underlying several methods for nonlinear systems, where the dynamics and cost are expanded around a nominal trajectory.
Stochastic optimal control
In stochastic control, the state is influenced by random disturbances. A controlled diffusion is commonly written as
[ dX_t
b(t,X_t,u_t),dt + \sigma(t,X_t,u_t),dW_t, ]
where (W_t) is a Wiener process. The objective is an expected cost,
[ J[u]
\mathbb E \left[ \Phi(X_{t_f}) + \int_{t_0}^{t_f} L(t,X_t,u_t),dt \right]. ]
The associated Hamilton–Jacobi–Bellman equation contains a second-order term generated by the diffusion:
[ -\frac{\partial V}{\partial t}
\inf_{u\in U} \left{ L + \nabla V^{\mathsf T}b + \frac{1}{2} \operatorname{tr} \left( \sigma\sigma^{\mathsf T}\nabla^2V \right) \right}. ]
When the system is linear, the disturbances are Gaussian, and the cost is quadratic, the problem leads to linear–quadratic–Gaussian control. The separation principle then decomposes the controller into a state estimator and a linear–quadratic feedback law, subject to the standard assumptions governing observability and stabilizability.
Partial observation changes the state of the optimization problem from the physical state to a conditional probability distribution. This distribution, often called the belief state, contains the information available from past observations. The resulting control problem is generally infinite-dimensional even when the physical system has finitely many state variables.
Numerical representation
Analytical solutions are restricted to particular system structures, so many optimal control problems are represented numerically as finite-dimensional optimization problems. In direct methods, the control or both the state and control are approximated on a computational mesh. The differential equation is enforced through numerical integration or collocation, producing a nonlinear programming problem.
Indirect methods discretize the necessary conditions obtained from the maximum principle. The state equations, adjoint equations, boundary conditions, and switching relations form a boundary value problem. Their numerical behavior depends on the sensitivity of the extremal trajectory and on the treatment of discontinuous controls or active state constraints.
Dynamic-programming discretizations approximate the value function over the state space. Their computational cost grows rapidly with state dimension, a phenomenon identified with the curse of dimensionality. Approximate dynamic programming and reinforcement learning replace the exact value function or policy by a parameterized approximation, thereby changing the original equation into a statistical or numerical approximation problem.
The distinct numerical formulations preserve different parts of the continuous theory. Direct transcription emphasizes feasibility and finite-dimensional optimization, indirect methods retain the canonical Hamiltonian structure, and value-function methods retain the feedback interpretation of dynamic programming.
See also
- Calculus of variations, which studies extrema of functionals and supplies the variational foundations of optimal control.
- Model predictive control, which repeatedly solves finite-horizon control problems using updated state information.
- Differential game, which extends optimal control to dynamical systems influenced by multiple decision-makers with distinct objectives.
- Trajectory optimization, which concerns the computation of state and control paths satisfying dynamical and boundary constraints.
- Controllability, which characterizes whether admissible controls can transfer a system between specified states.
- Hamilton–Jacobi equation, which connects variational principles, classical mechanics, and value-function methods.
- Viability theory, which studies trajectories constrained to remain within a prescribed state set.
- Reinforcement learning, which formulates sequential decision problems through policies, rewards, and value functions when system information is incomplete or learned from interaction.