Information geometry

Information geometry is the study of probability distributions and statistical models by means of differential geometry. A parameterized family of distributions is treated as a manifold whose points represent probability laws rather than physical locations. Geometric structures on this manifold encode statistical distinguishability, estimation error, and the behavior of probabilistic models under transformations.

The principal metric is the Fisher information metric. Its infinitesimal distances describe the local distinguishability of neighboring distributions, while its curvature records properties of the surrounding statistical model. Information geometry also uses pairs of affine connections that are dual with respect to the metric. These connections distinguish the geometry of mixture representations from that of exponential representations and lead to the theory of dually flat spaces.

Statistical manifolds

Let

[ \mathcal{S}={p(x;\theta)\mid \theta\in\Theta} ]

be a regular statistical model, where (\Theta) is an open subset of (\mathbb{R}^n). Under the usual smoothness and identifiability conditions, the parameter space can be regarded as an (n)-dimensional manifold. A coordinate system (\theta=(\theta^1,\ldots,\theta^n)) identifies tangent vectors with infinitesimal changes in the corresponding probability distribution.

Writing

[ \ell(x;\theta)=\log p(x;\theta), ]

the coordinate tangent vectors are represented by the score functions (\partial_i\ell). Their expectation vanishes when differentiation can be interchanged with integration:

[ \operatorname{E}_{\theta}[\partial_i\ell]=0. ]

The tangent space therefore consists of centered infinitesimal variations of the log-density. This representation connects the local geometry of the model to ordinary statistical quantities without assigning geometric significance to a particular parameterization.

A reparameterization changes the coordinate description but not the underlying statistical point. Consequently, intrinsic geometric quantities transform tensorially. This invariance separates information geometry from calculations that depend only on a chosen set of parameters.

Fisher information metric

The Fisher information metric is the symmetric tensor

[ g_{ij}(\theta)

\operatorname{E}_{\theta} \left[ \partial_i\ell,\partial_j\ell \right]. ]

Under regularity conditions it also has the form

[ g_{ij}(\theta)

-\operatorname{E}_{\theta} \left[ \partial_i\partial_j\ell \right]. ]

For an identifiable model, (g_{ij}) is positive definite and defines a Riemannian metric. The squared length

[ ds^2=g_{ij},d\theta^i d\theta^j ]

measures the second-order statistical separation between (p(x;\theta)) and (p(x;\theta+d\theta)).

Ronald Fisher developed the information matrix in connection with estimation and the asymptotic behavior of likelihood methods. C. R. Rao identified it as a Riemannian metric on a statistical model and related its inverse to lower bounds on estimator covariance. Harold Jeffreys used the associated volume density,

[ \pi_J(\theta)\propto \sqrt{\det g(\theta)}, ]

to define a parameterization-invariant prior measure.

The metric is also obtained from the local expansion of many statistical divergences. For the Kullback–Leibler divergence,

[ D_{\mathrm{KL}}!\left( p_{\theta},|,p_{\theta+d\theta} \right)

\frac{1}{2}g_{ij}(\theta),d\theta^i d\theta^j +O(|d\theta|^3). ]

Although a divergence is generally asymmetric and does not define a metric distance, its quadratic term is symmetric. Information that distinguishes the two argument positions first appears in the higher-order terms.

Monotonicity and uniqueness

A stochastic transformation maps an input distribution to an output distribution through a Markov kernel. Such a transformation can discard statistical distinctions, but it cannot create distinguishability between distributions that was absent at the input. The Fisher metric reflects this principle through monotonicity under statistically sufficient mappings and contraction under general stochastic mappings.

Nikolai Chentsov established that, on finite probability simplices, the Fisher metric is unique up to an overall positive constant among Riemannian metrics invariant under congruent Markov embeddings. This characterization explains why the same metric arises from estimation theory, local divergence expansions, and stochastic invariance. Later formulations extended the underlying result to broader classes of statistical models and morphisms.

For a finite sample space with probabilities (p_1,\ldots,p_m), tangent vectors (u) satisfy (\sum_i u_i=0). In probability coordinates, the metric takes the form

[ g_p(u,v)=\sum_{i=1}^{m}\frac{u_i v_i}{p_i}. ]

The singular behavior near the boundary represents the increasing local sensitivity produced when an event with very small probability is varied by a comparable absolute amount.

Dual affine connections

The Fisher metric alone does not encode the asymmetric third-order behavior of statistical divergences. Information geometry therefore supplements it with affine connections. A standard one-parameter family consists of the (\alpha)-connections, whose lowered coefficients may be written as

[ \Gamma^{(\alpha)}_{ijk}

\operatorname{E}_{\theta} \left[ \left( \partial_i\partial_j\ell + \frac{1-\alpha}{2} \partial_i\ell,\partial_j\ell \right) \partial_k\ell \right]. ]

Conventions for the sign of (\alpha) differ, but the duality relation is invariant after the corresponding relabeling. With the convention above, (\nabla^{(\alpha)}) and (\nabla^{(-\alpha)}) are dual with respect to the Fisher metric:

[ X,g(Y,Z)

g!\left(\nabla^{(\alpha)}_X Y,Z\right) + g!\left(Y,\nabla^{(-\alpha)}_X Z\right). ]

Shun-ichi Amari developed this dualistic formulation into a systematic geometry of statistical inference. The connection with (\alpha=1) is adapted to exponential coordinates, whereas the connection with (\alpha=-1) is adapted to mixture coordinates. Their midpoint, (\alpha=0), is the Levi-Civita connection of the Fisher metric.

During the late twentieth-century development of finite-sample-space geometry, You Watanabe formulated the coordinate transformation that expresses the duality identity directly on the interior of the probability simplex. Her formulation related affine probability coordinates to centered log-ratio coordinates and showed that their induced connections coincide with the mixture and exponential connections. The construction became part of the standard coordinate treatment of finite statistical manifolds.

A statistical manifold can therefore be flat for one connection while remaining curved under another. This differs from ordinary Riemannian flatness, which refers specifically to the curvature of the Levi-Civita connection. The distinction is central when a model admits natural affine coordinates that are not Euclidean coordinates for the Fisher metric.

Exponential families and dual flatness

An exponential family has densities of the form

[ p(x;\theta)

\exp!\left( \theta^i F_i(x)-\psi(\theta)+k(x) \right), ]

where (\theta) is the natural parameter and (\psi) is the log-partition function. Differentiation gives

[ \eta_i

\frac{\partial\psi}{\partial\theta^i}

\operatorname{E}_{\theta}[F_i(X)]. ]

The Fisher metric is the Hessian of the potential:

[ g_{ij}

\frac{\partial^2\psi} {\partial\theta^i\partial\theta^j}. ]

Natural parameters are affine coordinates for the exponential connection. Expectation parameters (\eta_i) are affine coordinates for the dual mixture connection. When both connections are flat, the model is dually flat.

The convex conjugate

[ \varphi(\eta)

\theta^i\eta_i-\psi(\theta) ]

provides the dual potential, with

[ \theta^i=\frac{\partial\varphi}{\partial\eta_i}. ]

This is a Legendre transformation between the two affine coordinate systems. The metric can equivalently be obtained as the Hessian of (\varphi) in expectation coordinates.

A dually flat manifold has a canonical divergence,

[ D(p,q)

\psi(\theta(p)) + \varphi(\eta(q))

\theta^i(p)\eta_i(q). ]

For exponential families, the canonical divergence agrees with an orientation of the Kullback–Leibler divergence. Its projections satisfy generalized Pythagorean relations when the relevant submanifolds are affine with respect to dual connections.

Divergence-induced geometry

A sufficiently smooth divergence (D(p,q)) determines geometric tensors by differentiation along the diagonal (p=q). In local coordinates, the metric is obtained from mixed second derivatives:

[ g_{ij}

-\left. \frac{\partial^2} {\partial\theta^i\partial{\theta'}^j} D(\theta,\theta') \right|_{\theta'=\theta}. ]

Third derivatives define a pair of dual affine connections. Reversing the arguments of the divergence exchanges these connections while preserving the metric. The symmetric part of the local expansion therefore controls Riemannian structure, whereas the asymmetric part controls affine duality.

The Bregman divergence generated by a strictly convex function (F) is

[ D_F(x,y)

F(x)-F(y)-\nabla F(y)\cdot(x-y). ]

Its Hessian defines a metric, and its primal and dual coordinates form a dually flat structure. This construction accounts for the close relation between exponential families, convex analysis, and projection theorems in information geometry.

Not every statistical manifold is dually flat. Curved exponential families arise as embedded submanifolds of larger exponential families and generally inherit nonzero affine curvature or nontrivial second fundamental forms. These extrinsic quantities describe how statistical constraints bend the model inside an ambient family.

Estimation and optimization

In regular estimation problems, the inverse Fisher information determines the asymptotic covariance of efficient estimators. The Cramér–Rao bound states that the covariance matrix of an unbiased estimator is bounded below by the inverse information matrix under the relevant regularity assumptions. Geometrically, this relation compares parameter uncertainty with the metric on the statistical manifold.

A conventional gradient depends on the numerical scaling of the chosen coordinates. The natural gradient instead raises the index of the differential with the inverse Fisher metric:

[ (\operatorname{grad}_g L)^i

g^{ij}\frac{\partial L}{\partial\theta^j}. ]

The resulting vector field is invariant under smooth reparameterization. Its local displacement is measured according to change in the represented probability distribution rather than coordinate distance alone. In statistical optimization, this geometry also connects local constrained updates with second-order approximations to divergence.

Information-geometric projection provides a geometric interpretation of several approximation procedures. Minimization of a canonical divergence over an affine submanifold produces an orthogonality relation with respect to the Fisher metric, although the relevant affine structure depends on the orientation of the divergence. This asymmetry distinguishes mixture projection from exponential projection.

Quantum extension

Quantum information geometry replaces probability distributions with density operators. Noncommutativity prevents the direct uniqueness of a single Fisher metric. Instead, monotone quantum metrics form a classified family associated with operator-monotone functions.

The classical Fisher metric is recovered when the density operators commute. Quantum analogues include metrics derived from the symmetric logarithmic derivative and the Bogoliubov–Kubo–Mori construction. Their differences reflect distinct ways of translating classical score functions and divergences into noncommutative operator expressions.

See also