---
title: Directional-Fisher Rate Overview
url: https://www.emergentmind.com/topics/directional-fisher-rate
type: topic
---

# Directional-Fisher Rate Overview

Directional-Fisher Rate is the directional projection of the Fisher information rate matrix for a parametrized stochastic process observed through a time series. In the framework of sequence probabilities $P_\theta(X_{1:L})$, it quantifies asymptotic learning efficiency along a specified unit direction $u$ in parameter space, after normalizing Fisher information by observation length. The concept is introduced within a general theory of learning unknown parameters of stochastic processes, including non-Markovian processes of infinite Markov order, where the Fisher information of observed sequence probabilities lower-bounds estimation variance and admits exact closed-form asymptotics through unifilar presentations and belief-state metadynamics [2310.03968].

## 1. Formal setting and definition

Consider a parameter vector $\theta \in \Theta$ and an observed sequence $X_{1:L}=X_1,X_2,\ldots,X_L$ over a countable alphabet $\mathcal{X}$, with sequence probability
\[
P_\theta(X_{1:L}=x_{1:L})
=
P_\theta(x_1)\,P_\theta(x_2\mid x_1)\cdots P_\theta(x_L\mid x_{1:L-1})
\,.
\]

For vector parameters, the Fisher information matrix for sequences of length $L$ is
\[
I_L(\theta)_{m,n}
=
\Big\langle
\big[\partial_{\theta_m}\ln P_\theta(X_{1:L})\big]
\big[\partial_{\theta_n}\ln P_\theta(X_{1:L})\big]
\Big\rangle_{P_\theta(X_{1:L})}
=
-\Big\langle
\partial_{\theta_m}\partial_{\theta_n}\ln P_\theta(X_{1:L})
\Big\rangle_{P_\theta(X_{1:L})}
\,,
\]
under standard regularity conditions, yielding a symmetric positive semidefinite matrix. In the scalar-parameter case,
\[
I_L(\theta)
=
\Big\langle\big[\partial_\theta \ln P_\theta(X_{1:L})\big]^2\Big\rangle
=
-\Big\langle\partial_\theta^2 \ln P_\theta(X_{1:L})\Big\rangle
\,.
\]

The asymptotic per-symbol Fisher information rate is
\[
I_F(\theta)
=
\lim_{L\to\infty}
\frac{1}{L}\,I_L(\theta)
\,.
\]
For stationary processes with finite excess information $\epsilon$, the Fisher information scales as
\[
I_L(\theta)
=
\mathcal{E}_L(\theta)+L\,I_F(\theta)
\quad\Rightarrow\quad
\lim_{L\to\infty}\frac{I_L(\theta)}{L}
=
I_F(\theta)
\,.
\]

The Directional-Fisher Rate is then defined for a unit direction $u$ in parameter space by
\[
I_F^{\mathrm{dir}}(\theta;u)
=
u^\top I_F(\theta)\,u
\,.
\]
It is therefore the quadratic form of the Fisher information rate matrix along a chosen parameter-space direction. In one dimension, the directional quantity coincides with the scalar Fisher information rate [2310.03968].

## 2. Closed-form expression and unifilar presentations

A central result is an exact closed-form expression for the Fisher information rate via any unifilar presentation, valid even for non-Markovian processes of infinite Markov order. Suppose a unifilar hidden-state model with countable recurrent states $\mathcal{S}^+$ and stationary distribution $\pi(s)$ over $\mathcal{S}^+$, and let $P_\theta(x\mid s)$ be the one-step symbol probabilities. Then
\[
I_F(\theta)_{m,n}
=
-\sum_{s\in\mathcal{S}^+}
\pi(s)\sum_{x\in\mathcal{X}}
P_\theta(x\mid s)\,
\partial_{\theta_m}\partial_{\theta_n}\ln P_\theta(x\mid s)
\]
and equivalently
\[
I_F(\theta)_{m,n}
=
\sum_{s\in\mathcal{S}^+}
\pi(s)\sum_{x\in\mathcal{X}}
\frac{
\big[\partial_{\theta_m}P_\theta(x\mid s)\big]\,
\big[\partial_{\theta_n}P_\theta(x\mid s)\big]
}{
P_\theta(x\mid s)
}
\,.
\]

Applying the directional projection yields
\[
I_F^{\mathrm{dir}}(\theta;u)
=
\sum_{s\in\mathcal{S}^+}
\pi(s)
\sum_{x\in\mathcal{X}}
\frac{\big[u\cdot\nabla_\theta P_\theta(x\mid s)\big]^2}{P_\theta(x\mid s)}
\,,
\]
or, equivalently,
\[
I_F^{\mathrm{dir}}(\theta;u)
=
-\sum_{s\in\mathcal{S}^+}
\pi(s)
\sum_{x\in\mathcal{X}}
P_\theta(x\mid s)\,
\big[u\cdot\nabla_\theta\big]^2\ln P_\theta(x\mid s)
\,.
\]

These formulas are valid for any unifilar presentation that faithfully represents the parametrized process in a neighborhood of $\theta$, including non-Markovian processes. The same framework applies when the unifilar presentation is obtained from the mixed-state presentation of a nonunifilar HMM. This suggests that the Directional-Fisher Rate is not restricted to finite-order Markov models, but is instead a property of the predictive state dynamics of the process [2310.03968].

## 3. Belief-state metadynamics and finite-length convergence

The finite-length approach to the asymptotic rate proceeds through estimation states and the mixed-state presentation. Estimation states $s\in\mathcal{S}$ are equivalence classes of histories that induce the same conditional distribution over futures uniformly in a neighborhood of $\theta$. The mixed-state presentation constructs unifilar dynamics over belief distributions $\eta^{(w)}$ over hidden HMM states, updated by Bayes under observed symbols.

The labeled transition operators over estimation states are
\[
W^{(x)}_{s,s'}
=
P_\theta(X_t=x\mid E_t=s)\,\delta_{wx\in s'}
\quad\text{for any }w\in s
\,,
\qquad
W=\sum_{x\in\mathcal{X}}W^{(x)}
\,,
\]
with left stationary eigenvector $\pi$ over recurrent estimation states and right eigenvector of ones.

The myopic Fisher information rate is the finite-length per-step increment
\[
\Delta I_L(\theta)
\equiv
I_L(\theta)-I_{L-1}(\theta)
\,,
\]
with the matrix form
\[
\Delta I_L(\theta)_{m,n}
=
\langle \mathbf{1}_{(\varnothing)}|\,W^{L-1}\,|f_{m,n}\rangle
\,,
\]
where the information vector is $L$-independent and has components
\[
\langle s|f_{m,n}\rangle
=
\sum_{x\in\mathcal{X}}
\frac{
\big[\partial_{\theta_m}P_\theta(x\mid s)\big]\,
\big[\partial_{\theta_n}P_\theta(x\mid s)\big]
}{
P_\theta(x\mid s)
}
\,.
\]

If there is a single attractor, $W_1=|1\rangle\langle\pi|$ and
\[
I_F(\theta)_{m,n}
=
\langle\pi|f_{m,n}\rangle
\,,
\]
so the Fisher information rate is the stationary-average of the information vector. The Directional-Fisher Rate is then the corresponding stationary average after projection along $u$.

The spectral decomposition of powers of $W$ yields the convergence structure of $\Delta I_L(\theta)$ toward $I_F(\theta)$. Contributions with $|\lambda|<1$ yield exponential decay, possibly modulated by polynomials; $\lambda=0$ yields ephemeral terms up to $L\le \nu_0$; and unit-modulus non-1 eigenvalues yield oscillatory modes. The same transient spectrum governs the convergence of the myopic Shannon entropy rate $h_L$ to the Shannon entropy rate $h_\mu$. Accordingly, $\Delta I_L(\theta)\to I_F(\theta)$ and $h_L\to h_\mu$ with exactly the same set of relaxation eigenvalues and timescales drawn from the transient spectrum of the mixed-state presentation [2310.03968].

## 4. Excess information and asymptotic estimation limits

Finite-length deviations from the asymptotic linear scaling are captured by excess information:
\[
\mathcal{E}_L(\theta)
\equiv
I_L(\theta)-L\,I_F(\theta)
=
\sum_{\ell=1}^L
\big[\Delta I_\ell(\theta)-I_F(\theta)\big]
\,.
\]
With $Q=W-W_1$ invertible,
\[
\mathcal{E}_L(\theta)
=
(I-Q)^{-1}(I-Q^L)|f\rangle
-|f\rangle
\,,
\qquad
\boldsymbol{\epsilon}(\theta)
\equiv
\lim_{L\to\infty}\mathcal{E}_L(\theta)
=
(I-Q)^{-1}|f\rangle
-|f\rangle
\,.
\]

This parallels the corresponding formulas for excess entropy. For finite Markov-order $M$ processes, myopic excess information saturates at $M$, while in many cases saturation occurs at the cryptic order $K\le M$. A plausible implication is that finite-length learning transients are controlled not only by Markov order but by the predictive state structure encoded in the cryptic organization of the process.

The estimation-theoretic significance of $I_F(\theta)$ and its directional projection follows from the Cramér–Rao lower bound. For unbiased estimators based on a single length-$L$ sequence,
\[
\mathrm{Var}(\hat\theta)\succeq I_L(\theta)^{-1}
\,,
\qquad
\text{scalar case: }
\mathrm{Var}(\hat\theta)\ge\frac{1}{I_L(\theta)}
\,.
\]
Using
\[
I_L(\theta)=\mathcal{E}_L(\theta)+L\,I_F(\theta)
\,,
\]
the asymptotic bound becomes
\[
\mathrm{Var}(\hat\theta)
\succeq
\frac{1}{L}\,I_F(\theta)^{-1}
\,,
\qquad
\text{scalar: }
\mathrm{Var}(\hat\theta)\ge\frac{1}{L\,I_F(\theta)}
\,.
\]
For $N$ independent sequences of length $L$, $\mathrm{Var}\ge I_L(\theta)^{-1}/N$.

Since the Directional-Fisher Rate is $u^\top I_F(\theta)u$, it specifies the asymptotic learning efficiency in the direction $u$. This suggests that anisotropy of the Fisher information rate matrix induces direction-dependent identifiability and direction-dependent variance scaling [2310.03968].

## 5. Computation for Markov chains, HMMs, and infinite-state presentations

For finite-order Markov chains, the procedure is explicit. One takes the unifilar states to be the last $M$ symbols, or a minimal equivalent, computes the stationary distribution $\pi$ over recurrent unifilar states, evaluates $P_\theta(x\mid s)$ and their derivatives, forms the information vector components
\[
\langle s|f_{m,n}\rangle
=
\sum_{x\in\mathcal{X}}
\frac{
\big[\partial_{\theta_m}P_\theta(x\mid s)\big]\,
\big[\partial_{\theta_n}P_\theta(x\mid s)\big]
}{
P_\theta(x\mid s)
}
\,,
\]
and then averages over $\pi$:
\[
I_F(\theta)_{m,n}=\sum_s \pi(s)\langle s|f_{m,n}\rangle
\,,
\qquad
I_F^{\mathrm{dir}}(\theta;u)=u^\top I_F(\theta)u
\,.
\]
For finite Markov order, the transient $W$ has only zero eigenvalues with index $\nu_0=M$, yielding exact convergence for $L\ge M+1$. Often saturation occurs at cryptic order $K\le M$.

For HMMs, including infinite Markov order, one begins from an HMM $\mathcal{M}=(\mathcal{X},(T^{(x)}),\mu)$, constructs mixed states
\[
\eta_\theta^{(w)}
=
\frac{\mu\,T^{(w)}}{\mu\,T^{(w)}\mathbf{1}}
\,,
\]
groups them into estimation states $s$, and computes
\[
P_\theta(x\mid s)
=
\eta_\theta^{(w)}\,T^{(x)}\mathbf{1}
\quad\text{for any }w\in s
\,.
\]
The metadynamics matrices $W^{(x)}$ are then built, the stationary distribution over recurrent estimation states is obtained, and the same stationary averaging formula yields $I_F(\theta)$ and $I_F^{\mathrm{dir}}(\theta;u)$.

For binary alphabets, the information vector simplifies to
\[
\langle s|f_{m,n}\rangle
=
\frac{
\big[\partial_{\theta_m}P_\theta(1\mid s)\big]\,
\big[\partial_{\theta_n}P_\theta(1\mid s)\big]
}{
P_\theta(0\mid s)\,P_\theta(1\mid s)
}
\,.
\]

For processes with no finite unifilar HMM, such as Parentheses Matching, the mixed-state presentation is constructed directly over the possibly infinite set of mixed states generated by the HMM, and one proceeds on the recurrent component. For nonergodic mixtures with disconnected components, one computes $I_F(\theta)$ within each block and combines with mixing weights; directional projection again takes the form $u^\top I_Fu$ [2310.03968].

## 6. Representative processes and explicit forms

The framework admits exact expressions for a range of qualitatively distinct stochastic processes.

For the IID biased coin with scalar parameter $p$,
\[
I_F(p)=\frac{1}{p(1-p)}
\,,
\qquad
I_F^{\mathrm{dir}}(p;u)=I_F(p)
\quad\text{(1D)}.
\]
The myopic excess information is $\epsilon=0$, and $\Delta I_L=I_F$ for all $L$.

For the General $\mathscr{M}$–$K$ Golden Mean process with scalar parameter $p$,
\[
\langle\pi|
=
\frac{1}{1+(\mathscr{M}+K-1)p}\,[1,\; p,\;\dots,\;p]
\,,
\qquad
\vec f
=
\bigl[\tfrac{1}{p(1-p)},\,0,\,\dots,\,0\bigr]^\top
\,,
\]
which yields
\[
I_F(p)
=
\frac{1}{p(1-p)\,[1+(\mathscr{M}+K-1)p]}
\,.
\]
Convergence satisfies $\Delta I_L=f$ for $L>K$, showing cryptic-order saturation.

For the Even Process, which has infinite Markov order,
\[
I_F(p)
=
\frac{1}{p(1+p)(1-p)}
\,.
\]
The convergence is oscillatory with period $2$:
\[
\Delta I_L
=
\begin{cases}
I_F - \dfrac{p^{L/2}}{p(1+p)^2}, & L \text{ even},\\[6pt]
I_F + \dfrac{p^{(L-1)/2}}{p(1+p)^2}, & L \text{ odd}.
\end{cases}
\]
Its excess information is
\[
\mathcal{E}_L
=
\frac{1}{p(1+p)^2}\,\big[1 - p^{L/2}\,\delta_{L/2\in\mathbb{Z}}\big]
\quad\Rightarrow\quad
\boldsymbol{\epsilon}
=
\frac{1}{p(1+p)^2}
\,.
\]

For the Teddy Bear process with two parameters $(p,q)$,
\[
I_F(p,q)
=
\gamma
\begin{bmatrix}
\dfrac{1-q}{p} & 1\\[6pt]
1 & \dfrac{1-p}{q}
\end{bmatrix}
\,,
\qquad
\gamma
=
\frac{1}{(1-p-q)(1+2p+4q)}
\,,
\]
and for a unit direction $u=(u_p,u_q)$,
\[
I_F^{\mathrm{dir}}(p,q;u)
=
\gamma\left[
\frac{1-q}{p}\,u_p^2
+2\,u_pu_q
+\frac{1-p}{q}\,u_q^2
\right]
\,.
\]
Its convergence shows ephemeral cutoff by cryptic order $K=2$, after which a period-3 decaying mode dominates.

For overparametrized, non-identifiable models, Fisher information matrices are singular, but the Fisher information rate can still be computed. An example based on the Even process with $g=g(p,q,r)$ yields
\[
I_F
=
\gamma\,(\nabla_\theta g)\,(\nabla_\theta g)^\top
\quad\text{with}\quad
\gamma=\frac{1}{g(1+g)(1-g)}
\,,
\]
which is rank-$1$, and
\[
I_F^{\mathrm{dir}}(\theta;u)
=
\gamma\,\big[u\cdot\nabla_\theta g\big]^2
\,.
\]
In such cases, the Drazin inverse or Moore–Penrose pseudoinverse can be used to assess variance bounds orthogonal to invariant subspaces [2310.03968].

## 7. Assumptions, scope, and conceptual significance

The theory assumes a countable alphabet $\mathcal{X}$ and discrete-time observations, with continuous-time obtained via an appropriate limit. Standard regularity conditions are required for Fisher information, including exchange of derivative and expectation. Stationarity is assumed when asserting limits and stationary distributions over recurrent estimation states. Unifilar presentations must be faithful in a neighborhood of $\theta$, and the mixed-state presentation provides a canonical unifilar presentation even for nonunifilar HMMs. The spectrum of $W$ is assumed to comprise isolated eigenvalues. When using $Q=W-W_1$ invertibility, deterministic periodicities on the recurrent component are excluded; otherwise unit-modulus modes must be included explicitly in spectral sums [2310.03968].

Within these assumptions, the Directional-Fisher Rate furnishes an exact asymptotic measure of inferential sensitivity along selected parameter-space directions. Its role is simultaneously geometric, through the quadratic form $u^\top I_Fu$; dynamical, through the mixed-state metadynamics and transient spectrum; and statistical, through the asymptotic Cramér–Rao scaling. The central structural result is that both learning and synchronization are governed by the same transient spectral data: the myopic Fisher-information rate and the myopic entropy rate relax with identical eigenvalues and timescales. This establishes a direct connection between predictive-state relaxation and asymptotic parameter-learning efficiency in both Markovian and non-Markovian stochastic processes [2310.03968].

Source: https://www.emergentmind.com/topics/directional-fisher-rate