---
title: Projected Expected Fisher Operator
url: https://www.emergentmind.com/topics/projected-expected-fisher-operator
type: topic
---

# Projected Expected Fisher Operator

Searching arXiv for the cited works and topic variants to ground the article in the relevant literature.
The **Projected Expected Fisher Operator** is not a single universally fixed object but a family of closely related constructions in which Fisher information is first formed as an expected squared score, Hessian, or covariance-like operator and then restricted, projected, or conditionally averaged onto the statistically meaningful directions of a model. Across infinite-dimensional inverse problems, noncommutative prediction, nonlinear semiparametric inference, stochastic optimization, continual learning, quantum sensing, and covariance estimation, the common role of the construction is to remove non-identifiable, ill-posed, or operationally irrelevant directions while preserving the geometry that controls estimation error, diffusion, or metrological sensitivity. In the most direct Hilbert-space form this yields operators such as $\bar I_{\mathrm{proj}} = P\,\mathbb E[J(\theta)^\ast \Gamma^{-1} J(\theta)]\,P$, whereas in operator-algebraic settings it appears as a conditional expectation $E_t(J^\ast J)\in N_t$, and in projected Fisher geometry for SGD as $F^{\star}(\theta)=\Pi(\theta) I(\theta)\Pi(\theta)^\top$ [1203.5397], [2601.14355], [2603.02417].

## 1. Conceptual form and scope

The underlying pattern is a two-stage construction. First, one forms an **expected Fisher object**. Depending on context, this may be the classical expected Fisher information $I(\theta)=\mathbb E[s_\theta s_\theta^\top]$, the inverse-problem operator $J(\theta)^\ast \Gamma^{-1} J(\theta)$, the operator-valued square of a conjugate variable $E_t(J^\ast J)$, or a Haar-averaged quantum Fisher-information operator. Second, one introduces an explicit **projection mechanism**. In the cited literature this mechanism may be an orthogonal projection $P$ onto a closed parameter subspace, a projection $\Pi(\theta)$ onto an identifiable tangent space, a conditional expectation onto an information algebra, or a symmetry projector $\Pi_i$ onto an invariant quantum sector [1203.5397], [2603.02417], [2512.13294].

This common structure serves several distinct purposes. In infinite-dimensional Gaussian models it is needed because the inverse covariance is only meaningful on the Cameron–Martin space. In nonlinear inverse and regression models it isolates the identified domain and codomain Hilbert spaces on which the Fisher operator becomes invertible. In noncommutative prediction it converts the squared score into an $N_t$-valued quantity adapted to the available information. In optimization it removes degenerate directions and exposes the matrix-valued noise geometry of SGD or the Fisher-orthogonal complement required for continual learning. In quantum sensing it restricts the Fisher operator to invariant or postselected sectors where the relevant orbit has enhanced metrological compatibility [1203.5397], [2601.13254], [2601.14355], [2601.12816].

A recurring misconception is that such projections are merely numerical truncations. The cited works instead show that projection often has a structural role: it can be required for well-posedness, for invertibility, for conditional measurability, or for the very definition of the relevant score-based operator. This suggests that the adjective “projected” is not incidental but identifies the operative statistical geometry.

## 2. Infinite-dimensional Gaussian inverse problems

In the Hilbert-space inverse-problem framework of Nordebo et al., the data model is
\[
y = F(\theta) + \varepsilon,
\]
with $F:X\to Y$ Fréchet differentiable, $X$ and $Y$ separable Hilbert spaces, $\varepsilon$ zero-mean Gaussian on $Y$, covariance $\Gamma:Y\to Y$ positive, self-adjoint, and trace class, and Jacobian $J(\theta):X\to Y$ Hilbert–Schmidt. The natural data space for Fisher analysis is not $Y$ itself but the Cameron–Martin space
\[
H_\Gamma = \operatorname{ran}(\Gamma^{1/2}) \subset Y,
\]
with inner product
\[
\langle v,u\rangle_{H_\Gamma}=\langle \Gamma^{-1/2}v,\Gamma^{-1/2}u\rangle_Y.
\]
The reason is that in infinite dimensions $\operatorname{ran}(\Gamma)\neq Y$, so $\Gamma^{-1}$ is not defined on all of $Y$; it is defined on $\operatorname{ran}(\Gamma)$ and extended as a pseudo-inverse acting trivially on $N(\Gamma)$. Consequently, the range condition $\operatorname{ran}(J(\theta))\subset H_\Gamma$ is required in order that $\Gamma^{-1}J(\theta)$ be meaningful [1203.5397].

Under that condition, the Fisher operator is
\[
I(\theta)=J(\theta)^\ast \Gamma^{-1} J(\theta)
\]
in the complex case, with the paper’s real-case convention inserting a factor $2$. If $S\subset X$ is a closed subspace and $P:X\to S$ is the orthogonal projection, the projected Fisher operator is
\[
I_{\mathrm{proj}}(\theta)=P\,J(\theta)^\ast \Gamma^{-1}J(\theta)\,P.
\]
For Gaussian likelihood with $\Gamma$ independent of $\theta$, the expected Fisher equals the conditional Fisher because the Hessian of the log-likelihood is $-J(\theta)^\ast\Gamma^{-1}J(\theta)$ and does not depend on $y$. With a prior $\pi$ on $\theta$, the prior-averaged operator is
\[
\bar I=\mathbb E_\pi[J(\theta)^\ast \Gamma^{-1}J(\theta)],
\qquad
\bar I_{\mathrm{proj}}=P\bar I P=\mathbb E_\pi[PJ(\theta)^\ast \Gamma^{-1}J(\theta)P].
\]
This is the most direct version of a projected expected Fisher operator in infinite-dimensional Gaussian inference [1203.5397].

Trace-class criteria are governed by the spectral interaction between the singular values $\sigma_i$ of $J$ and the eigenvalues $\lambda_j$ of $\Gamma$. A sufficient condition for $\operatorname{ran}(J)\subset \operatorname{ran}(\Gamma)$ is
\[
\sum_i \sigma_i^2 \sum_j \frac{1}{\lambda_j^2} |\langle \phi_j,u_i\rangle|^2 < \infty,
\]
and trace class of $I(\theta)$ follows if
\[
\operatorname{tr}(I(\theta))
=
\sum_i \sigma_i^2 \sum_j \frac{1}{\lambda_j} |\langle \phi_j,u_i\rangle|^2
<\infty.
\]
In the diagonal case $u_i=\phi_i$, this reduces to $\sum_i \sigma_i(\theta)^2/\lambda_i<\infty$. The paper’s electromagnetic inverse-source example shows that the asymptotic decay rates of $\sigma^2_{\tau lm}$ and $\lambda_{\tau lm}$ can force opposite behaviors for Fisher and Cramér–Rao operators: with external spherically isotropic noise, $\sigma^2_{\tau lm}/\lambda_{\tau lm}\to\infty$, so the infinite-dimensional Fisher operator diverges, while the pseudo-inverse remains trace class and the finite-dimensional CRB converges; with added internal white noise, the Fisher becomes trace class but the CRB diverges as the truncation level $L\to\infty$ [1203.5397].

The paper explicitly states that in this setting projection onto parameter subspaces is not merely computationally expedient but mathematically necessary. The projected expected Fisher operator is therefore both a statistical quantity and a regularity device.

## 3. Identified Hilbert spaces and nonlinear statistical models

In nonlinear regression models of the form
\[
Y_i = G(\theta)(X_i)+\varepsilon_i,
\]
with $G:\Theta\to L^2_\lambda$ mapping a Borel subset of a separable Hilbert space $(V,\langle\cdot,\cdot\rangle_V)$ into a nonlinear submanifold of $L^2_\lambda$, the expected Fisher operator is determined by the $L^2$-linearization $I_{\theta_0}:V_0\to L^2_\lambda$ at a base point $\theta_0$ and the Fisher information of the noise,
\[
I_\varepsilon = 4\int_{\mathbb R^p} (\nabla \sqrt{q_\varepsilon})(y)(\nabla \sqrt{q_\varepsilon})(y)^T\,dy.
\]
The model is differentiable in quadratic mean, and the score operator $A_{\theta_0}$ satisfies
\[
A_{\theta_0}^\ast A_{\theta_0}=I_{\theta_0}^\ast I_\varepsilon I_{\theta_0}.
\]
This yields the expected Fisher operator
\[
I(\theta_0)\equiv I_{\theta_0}^\ast I_\varepsilon I_{\theta_0}:H\to S,
\]
where $H$ is the completion of $V_0$ under the LAN norm
\[
\|h\|_H=\|I_\varepsilon^{1/2}I_{\theta_0}[h]\|_{L^2_\lambda},
\]
and $S$ is the identified dual space defined through the $V$-pairing. Under injectivity of $I_{\theta_0}$, Theorem 3.4 shows that $I(\theta_0)$ is an isometric homeomorphism from $H$ onto $S$ [2601.13254].

In this framework, the projected expected Fisher operator is the restriction of the expected Fisher to identified directions,
\[
I_{\mathrm{proj}}(\theta_0):=I_{\theta_0}^\ast I_\varepsilon I_{\theta_0}:H\to S.
\]
If one introduces an ambient pivot space $\mathcal H=\overline V_0$, with orthogonal projectors $P_H$ and $P_S$ onto the closures of $H$ and $S$, one may write
\[
I_{\mathrm{proj}}(\theta_0)=P_S\,I_{\mathrm{full}}(\theta_0)\,P_H.
\]
The associated Moore–Penrose pseudoinverse is
\[
I_{\mathrm{full}}(\theta_0)^+ = I_{\mathrm{proj}}(\theta_0)^{-1} P_S.
\]
The central point is that invertibility is automatic once the score is injective, provided the operator is viewed between the identified spaces $H$ and $S$ rather than on an ambient Hilbert space [2601.13254].

This identified-space perspective is operational rather than merely formal. In the reaction–diffusion example on $[0,T]\times \mathbb T^d$, the two-sided estimate
\[
\|h\|_{H^1(\mathbb T^d)^\ast}\lesssim \|I_{\theta_0}[h]\|_{L^2([0,T]\times \mathbb T^d)}\lesssim \|h\|_{H^1(\mathbb T^d)^\ast}
\]
implies $H=H^1(\mathbb T^d)^\ast$ and $S=H^1(\mathbb T^d)$. In incompressible $2$D Navier–Stokes, one analogously obtains $H=(H^1)^\ast$ and $S=H^1$ on divergence-free fields. These spaces feed directly into efficient Gaussian limits, influence-function equations of the form
\[
\phi=I_{\mathrm{proj}}(\theta_0)^{-1}\dot F^\ast[\cdot],
\]
and semiparametric lower bounds [2601.13254].

A plausible implication is that the projected expected Fisher operator in nonlinear inverse models should be understood less as a truncation of a larger ambient operator than as the canonical Fisher operator once identification has been completed.

## 4. Conditional operator-valued Fisher under information filtrations

In the noncommutative framework of a von Neumann algebra $M$ with an increasing family of abelian subalgebras $(N_t)_{t\in[0,T]}$ and normal $\varphi_\rho$-preserving conditional expectations $E_t:M\to N_t$, the role of projection is played by conditioning onto the information algebra. The Hilbert space $L^2(M,\varphi_\rho)$ carries inner product
\[
\langle X,Y\rangle_{\varphi_\rho}=\varphi_\rho(X^\ast Y),
\]
and $E_t$ is the $L^2(M,\varphi_\rho)$-orthogonal projection onto $L^2(N_t,\varphi_\rho)$. For a self-adjoint variable $X$ relative to a conditioning algebra $D=N_t$ and a background algebra $B\supset D$, the paper defines the operator-valued conjugate variable
\[
J_D(X:B)=\partial_{X:B}^\ast(1\otimes 1),
\]
where $\partial_{X:B}$ is the $D$-bimodular free difference quotient. The operator-valued Fisher information is then
\[
I_D(X:B)=E_D(J_D(X:B)^\ast J_D(X:B))\in D_+.
\]
The paper explicitly identifies the phrase “Projected Expected Fisher Operator” with the $N_t$-valued quantity
\[
\mathcal J_t(X:B)=E_t(J_{N_t}(X:B)^\ast J_{N_t}(X:B))\in N_t^+,
\]
and in the multivariate case with the sum over coordinates [2601.14355].

Here “expected” means application of the conditional expectation $E_t$ to the squared score, and “projected” means that the result lies in $N_t$. This is not only semantic. Because $E_t$ is the optimal $L^2$ predictor, $\mathcal J_t(X:B)$ is directly linked to conditional prediction limits. If $T\in N_t$ is any unbiased $N_t$-measurable estimator or predictor with $E_t(T)=E_t(X)$ and $I_{N_t}(X:B)$ is invertible in $N_t$, the operator-valued Cramér–Rao inequality gives
\[
E_t\big((T-X)^2\big)\succeq I_{N_t}(X:B)^{-1}.
\]
For the optimal predictor $T^\ast=E_t(X)$ this becomes
\[
E_t\big((X-E_t(X))^2\big)\succeq I_{N_t}(X:B)^{-1}.
\]
In the commutative reduction $M=L^\infty(\Omega,\mathcal F)$, $N_t=L^\infty(\Omega,\mathcal F_t)$, and $E_t$ classical conditional expectation, this reduces to conditional expected Fisher information and the classical conditional Cramér–Rao bound [2601.14355].

The framework also exposes a significant limitation. In compound Poisson lattice-jump models, conjugate variables typically do not exist, so the operator-valued Fisher information is infinite and the CR route may degenerate. The paper then computes the exact minimal conditional mean-square error directly from jump intensities:
\[
E_t\big((R_{t,h}-E_tR_{t,h})^2\big)=h\,(\Delta x)^2\sum_\alpha \alpha^2\gamma_\alpha.
\]
This sharp error floor replaces Fisher-based bounds in pure-jump regimes. The broader lesson is that projected expected Fisher operators remain meaningful only when the relevant score object exists in the chosen operator geometry.

## 5. Projected Fisher geometry in stochastic optimization and continual learning

In stochastic optimization, the projected expected Fisher operator appears as the intrinsic covariance geometry of gradient noise. For per-sample gradients $\psi(\theta;z)=\nabla_\theta \ell(\theta;z)$ and population risk $L(\theta)=\mathbb E[\ell(\theta;Z)]$, the well-specified likelihood case uses the expected Fisher
\[
I(\theta)=\mathbb E[\psi(\theta;Z)\psi(\theta;Z)^\top],
\]
while general $M$-estimation uses
\[
B(\theta)=\mathbb E[\psi(\theta;Z)\psi(\theta;Z)^\top].
\]
If $T_\theta$ is the identifiable tangent space and $\Pi(\theta)$ the Euclidean-orthogonal projector onto it, then the projected expected Fisher is
\[
F^\star(\theta)=\Pi(\theta) I(\theta)\Pi(\theta)^\top,
\]
with Godambe generalization
\[
G^\star(\theta)=\Pi(\theta) B(\theta)\Pi(\theta)^\top.
\]
Under exchangeable sampling and mini-batching, the mini-batch covariance is, to leading order, $(1/b)F^\star(\theta^\star)$ in the correctly specified likelihood case and $(1/b)G^\star(\theta^\star)$ in general $M$-estimation. This identification fixes the diffusion approximation
\[
d\theta_s=-\nabla L(\theta_s)\,ds+\sqrt{\tau}\,C^\star(\theta_s)\,dW_s,
\qquad \tau=\eta/b,
\]
with $C^\star(\theta)C^\star(\theta)^\top=G^\star(\theta)$. Near a stationary point $\theta^\star$, the OU linearization has stationary covariance $\Sigma$ satisfying
\[
H^\star \Sigma + \Sigma H^{\star\top} = \tau S(\theta^\star),
\]
where $S(\theta^\star)=F^\star(\theta^\star)$ in likelihood models and $S(\theta^\star)=G^\star(\theta^\star)$ in general losses. The same geometry drives oracle-complexity results expressed through an intrinsic effective dimension and a Fisher/Godambe condition number rather than ambient dimension [2603.02417].

An important negative result accompanies this formulation. Scalar-temperature surrogates match only total noise power and cannot reproduce the directional anisotropy, off-diagonal stationary covariances, or rotation-sensitive structure encoded by $F^\star$ or $G^\star$. The projected expected Fisher operator is therefore the matrix-valued noise geometry of SGD, not merely a rescaling parameter [2603.02417].

A distinct optimization use arises in continual learning, where the expected Fisher serves as the local Riemannian metric and projection is imposed to prevent interference with past tasks. With
\[
F(\theta)=\mathbb E_{(x,y)\sim D}\big[\nabla_\theta \log p(y|x,\theta)\nabla_\theta \log p(y|x,\theta)^\top\big],
\]
the Fisher-induced projector onto $\operatorname{span}(V)$ is
\[
P_F = V(V^\top F V)^{-1}V^\top F,
\qquad
\Pi_F^\perp = I-P_F.
\]
FOPNG uses an old-task Fisher $F_{\mathrm{old}}$ for orthogonality and a new-task Fisher $F_{\mathrm{new}}$ for step control. With
\[
M=F_{\mathrm{new}}^{-1/2}F_{\mathrm{old}}G,
\qquad
\Pi^\perp_{F_{\mathrm{new}};F_{\mathrm{old}},G}
=
I-M(M^\top M)^{-1}M^\top,
\]
the paper defines the projected expected Fisher operator acting on the incoming gradient $g$ by
\[
\mathcal P_{EF}(g)
=
F_{\mathrm{new}}^{-1/2}
\Pi^\perp_{F_{\mathrm{new}};F_{\mathrm{old}},G}
F_{\mathrm{new}}^{-1/2}g.
\]
The normalized trust-region step is then proportional to $\mathcal P_{EF}(g)$. This operator whitens by $F_{\mathrm{new}}^{-1/2}$, projects in Euclidean coordinates against old-task directions measured by $F_{\mathrm{old}}$, unwhitens, and thereby yields a reparameterization-invariant descent direction in the Fisher metric. In practice the diagonal Fisher is the default implementation, together with damping and low-rank gradient storage [2601.12816].

These optimization formulations broaden the meaning of the term. The projected expected Fisher operator is no longer only a statistical information operator; it also becomes the generator of anisotropic diffusion and the projector defining admissible descent directions.

## 6. Quantum metrology and covariance-operator realizations

In quantum sensing, the projected expected Fisher operator is formulated through invariant-subspace or postselection projections and Haar moment operators. For a phase generator $G$, an invariant block $\mathcal H_i$ with projector $\Pi_i$, and Haar moment operators
\[
\mathcal M_{\mu_H}^{(k)}(O)=\mathbb E_{U\sim \mu_H}[U^{\otimes k} O \,\overline U^{\otimes k}],
\]
the two-copy expected Fisher-information operator is
\[
\widehat{\mathcal F}_{\mathrm{proj,exp}}^{(2)}(\Pi_i,G;\mu_H)
=
4\Big(
\mathcal M_{\mu_H}^{(1)}((\Pi_i G \Pi_i)^2)\otimes I
-
\mathcal M_{\mu_H}^{(2)}((\Pi_i G \Pi_i)^{\otimes 2})
\Big).
\]
Its expectation on $|\psi_0\rangle\otimes |\psi_0\rangle$ yields the Haar-averaged QFI, with leading behavior
\[
\mathbb E_{U\sim \mu_H}[\mathcal F]
\approx
\frac{4\,\operatorname{Tr}[(\Pi_i G \Pi_i)^2]}{d_i}.
\]
The cited work shows shot-noise scaling $\mathcal O(n)$ for full $SU(2^n)$ averaging, but Heisenberg scaling $\mathcal O(n^2)$ in the permutation-symmetric subspace, and likewise for a projected-ensemble protocol in which outcome-conditioned effective generators $G_{\mathrm{eff}(z_e)}$ are averaged after local measurement and feed-forward. In this setting the projection is what restores metrologically compatible orbits [2512.13294].

A more classical operator realization appears in the Gaussian model with Wishart-randomized precision. On the space $S^n$ of symmetric matrices, with
\[
P(u)(s)=u s u,
\qquad
(u\otimes u)(s)=u\,\operatorname{Tr}(u s),
\]
the expected Fisher operator of the marginal model is
\[
I_p(\sigma)
=
\frac{(2p+1)P(\sigma^{-1})-(\sigma^{-1}\otimes \sigma^{-1})}{2(2p+3)}.
\]
Its inverse is
\[
I_p(\sigma)^{-1}
=
\frac{2(2p+3)}{2p+1}P(\sigma)
+
\frac{2(2p+3)}{(2p+1)(2p+1-n)}(\sigma\otimes \sigma).
\]
For a subspace $S\subset S^n$ with orthogonal projector $\Pi$, the projected expected Fisher operator is
\[
I_p(\sigma)_{\mathrm{proj},S}=\Pi I_p(\sigma)\Pi.
\]
On the weighted-traceless subspace $T_{\sigma^{-1}}=\{s:\operatorname{Tr}(\sigma^{-1}s)=0\}$, the rank-one term vanishes and only the $P(\sigma^{-1})$ part survives; on trace-containing subspaces the $\sigma^{-1}\otimes \sigma^{-1}$ term remains as a single distinguished direction. This decomposition makes explicit how projection alters conditioning by suppressing or preserving trace sensitivity [2211.14137].

Taken together, these realizations show that the projected expected Fisher operator is a genuinely operator-theoretic notion. It may live on Hilbert spaces, identified Sobolev scales, von Neumann subalgebras, tangent spaces, two-copy quantum spaces, or symmetric-matrix cones, but in each case it isolates the effective information-bearing sector while retaining the geometry that determines bounds, equilibria, or sensitivities.

Source: https://www.emergentmind.com/topics/projected-expected-fisher-operator