---
title: Residual Operator Learning in Neural PDEs
url: https://www.emergentmind.com/topics/residual-operator-learning
type: topic
---

# Residual Operator Learning in Neural PDEs

Searching arXiv for the cited papers and closely related work on residual operator learning.
arXiv search query: residual operator learning neural operator PDE residual correction

Residual operator learning denotes a family of formulations in which the learned map is not the full target operator in one shot, but a correction, defect, closure, or refinement relative to a baseline object such as an auxiliary trajectory, a coarse surrogate, a linearized solve, a symbolic prefix, or a bag-derived estimator. In the recent literature, this pattern appears in neural PDE solvers, probabilistic operator surrogates, learned preconditioners, weakly supervised uncertainty estimation, robot manipulation, teleoperation, and residual policy refinement. The common reconstruction pattern is additive or compositional: a baseline prediction is first produced or retrieved, and a learned residual operator supplies the unresolved part needed to obtain the final output [2406.09795][2306.12047][2603.18527][2512.12749][2606.05248].

## 1. Conceptual scope and terminology

The expression “residual operator” has been used in several technically distinct but structurally related senses. In PDE solving, DeltaPhi reformulates direct operator learning
\[
\mathcal{G}:\mathcal{A}\to\mathcal{U}
\]
as residual operator learning
\[
\mathcal{G}^{\Delta}:\mathcal{A}^2\to \Delta\mathcal{U},
\]
so that the model predicts \(u_i-u_j\) between a target solution and a similar auxiliary solution rather than \(u_i\) directly [2406.09795]. In the residual-based corrector literature, the residual operator is the linearized variational correction
\[
\mathcal F^C(m,\tilde u)=\tilde u-\delta_u R(m,\tilde u)^{-1}R(m,\tilde u),
\]
applied after a neural-operator surrogate has produced \(\tilde u\) [2306.12047]. In learned iterative solvers, the learned object is a residual-to-update map embedded inside Born-series, Newton–Kantorovich, or preconditioned iterations rather than an end-to-end solution operator [2603.18527][2511.19980].

Outside PDEs, the same phrase is used for structured correction operators anchored to a pre-existing decision map. In weakly supervised multi-instance uncertainty estimation, the residual operator is
\[
R(\mathbf{x})=g_{\phi}(f_{\psi}(\mathbf{x})) + r_{\pi}(f_{\psi}(\mathbf{x})),
\]
a residual-corrected instance estimator derived from a bag-level symmetric-function classifier [2405.04405]. In inverse manipulation, a symbolic planner first restores what it can, and a residual policy learns the continuous control needed to satisfy unresolved inverse predicates, yielding
\[
\pi_{\text{inv}}=\pi_{\text{sym}}\circ \pi_{\text{res}}.
\]
[2606.05248]

A useful unifying description is that residual operator learning shifts the learning target from the full response manifold to a smaller discrepancy space. In some papers that discrepancy is deterministic and additive; in others it is a solver increment, a metric-weighted defect, a probabilistic correction flow, or a constrained residual control law. This suggests that the field is less a single architecture than a modeling principle about what should be learned.

## 2. Core mathematical patterns

A first recurring pattern is baseline-plus-correction reconstruction. DeltaPhi trains on residual labels
\[
\mathcal{T}^{res}=\{(a_i,a_{k_i}),u_i-u_{k_i}\}_{i=1}^N,
\]
with objective
\[
\ell=\sum_{i=1}^{N}\mathcal{L}\big(\mathcal{G}_{\theta}^{\Delta}(a_i,a_{k_i}),u_i-u_{k_i}\big),
\]
and reconstructs the final prediction as
\[
\hat{u}_i = Q(v_l)+u_{k_i}.
\]
The auxiliary sample is selected by a retrieval function based on similarity and top-\(K\) random sampling during training [2406.09795]. Closely related additive reconstructions appear in residual video diffusion for PDEs,
\[
\hat{\mathbf x}_0=\hat{\mathcal R}_{unnormed}+\mathbf x^{prior},
\]
after S-DeepONet supplies a coarse prior [2507.06133], in low-to-high fidelity probabilistic transport,
\[
\mathcal{G}[a]=\mathcal{G}_{\text{LF}}[a]+\mathcal{C}[a],
\]
where \(\mathcal C\) is learned probabilistically by flow matching in function space [2512.12749], and in field-space closure models for PIC simulation,
\[
\hat{\mathbf u}=\hat{\mathbf u}_{up}+\delta\hat{\mathbf u}_r,
\]
with an additional latent residual closure on the source representation [2606.17733].

A second pattern is residual-to-update learning inside an iterative solver. The Neural Preconditioned Born Series replaces the scalar CBS relaxation by a neural operator:
\[
\bm u^{m+1}=\bm u^m+\mathcal{M}_\theta\!\left(G_\eta(\bm f-A\bm u^m)\right),
\]
so the network learns a correction operator in preconditioned residual coordinates rather than the solution field directly [2603.18527]. CHONKNORIS adopts the same principle in a Newton–Kantorovich setting by learning Cholesky factors of a regularized inverse Hessian-like operator and updating
\[
v_{n+1}=v_n-\alpha_n\, \mathcal R(u,v_n,\lambda_n)\mathcal R(u,v_n,\lambda_n)^T
\left(\frac{\delta \mathcal F}{\delta v}(u,v_n)\right)^*\mathcal F(u,v_n).
\]
[2511.19980]

A third pattern is residual weighting or residual geometry. The preconditioner-learning literature argues that matching residuals in the Euclidean norm can be mismatched to the actual operator geometry, and instead defines
\[
\|\bm r\|_{R_\eta}:=\|G_\eta \bm r\|_2,\qquad R_\eta=G_\eta^\ast G_\eta,
\]
so that residual learning is carried out in the metric induced by the shifted-Laplacian/Born reference operator [2603.18527]. The adaptive-training literature makes a complementary point: residual-based weighting rules can be derived variationally, with exponential weights targeting \(L^\infty\)-like objectives and linear residual weights recovering a variance/\(L^2\)-type regime [2509.14198].

These patterns are mathematically distinct, but all replace “learn the whole operator” by “learn the unresolved part of an operator action.”

## 3. PDE-centered residual operator learning

The most fully developed use of the term is in neural PDE solving. DeltaPhi is explicitly motivated by limited-data, low-resolution, and biased-data regimes in which direct neural operators may overfit the training resolution or training distribution. It keeps the neural-operator backbone
\[
\mathcal{G}_{\theta}=Q\circ \sigma(W_l+\mathcal K_l)\circ \cdots \circ \sigma(W_1+\mathcal K_1)\circ P
\]
unchanged, but changes the inputs and labels to residual form by concatenating \(a_i\oplus a_{k_i}\) and predicting \(u_i-u_{k_i}\). The paper also introduces customized auxiliary inputs such as partial auxiliary history, the auxiliary output \(u_{k_i}(x)\), and similarity scores \(Score_{i,k_i}\), and reports that these improve optimization [2406.09795].

Residual formulations in PDE surrogates also appear as hierarchical refinement. ResFNO learns the map from cure cycle \(T_a(t)\) to temperature history \(T_c(t)\), but replaces direct fitting of the raw time-domain residual with a Fourier residual mapping based on a low-mode reconstruction of the cure cycle. The purpose is to reduce the singular error near non-differentiable turning points of piecewise cure cycles and to preserve time-resolution independence of the operator parameterization [2111.10262]. The two-stage S-DeepONet plus video-diffusion framework similarly decomposes the spatio-temporal solution into a coarse, physics-consistent prior and a residual containing high-frequency structures. On lid-driven cavity flow, the reported mean relative \(L_2\) error drops from \(4.58\%\) for S-DeepONet to \(0.829\%\) for prior-conditioned residual diffusion, while on dogbone plasticity it drops from \(4.43\%\) to \(2.94\%\) [2507.06133].

Residual closure is also used for multi-field kinetic plasma simulation. LRC-FNO introduces two residual levels: a latent closure refiner corrects information lost by source compression, and a residual-closure FNO restores the high-resolution field correction missed by a coarse field solver. In the 2D scrape-off layer benchmark, the paper reports relative \(L_2\) errors of \(0.0447\) for the self-consistent potential and \(0.0251\) for the magnetic vector potential in single-step prediction, while also using the network as an initialization for iterative corrections in closed-loop PIC integration [2606.17733].

Residual operator learning can also be coupled to geometry normalization. NDNO maps varying 3D component geometries to a common reference domain by a diffeomorphic neural network and then learns the operator from residual stress fields to deformation fields on that reference domain. The learned deformation is finally mapped back to the original geometry. This converts variable-domain operator learning into reference-domain residual-stress-to-deformation learning with smoothness, invertibility, and Sinkhorn similarity constraints on the geometry map [2509.12237].

A probabilistic variant is given by residual-augmented flow matching in function space. There the model learns a stochastic correction transport from a low-fidelity PDE approximation to the high-fidelity solution manifold rather than learning the full conditional solution distribution from scratch. The operator backbone is FiLM-conditioned and resolution invariant, and the learned vector field is decomposed into a linear operator plus a nonlinear operator for stability and expressiveness [2512.12749].

## 4. Correctors, variational correctness, and operator-theoretic analysis

A major branch of the literature treats residual operator learning not as a surrogate for the solution operator itself, but as a post-processing or solver mechanism with explicit error-control semantics. The residual-based corrector operator for nonlinear variational boundary-value problems linearizes the PDE residual around a neural-operator prediction and solves
\[
\delta_u R(m,\tilde u)(e^C)=-R(m,\tilde u),
\]
after which \(u^C=\tilde u+e^C\). Under boundedness and invertibility assumptions, the corrected error is quadratic in the original prediction error. In numerical experiments on a nonlinear reaction–diffusion model, the paper reports almost two orders of increase in accuracy, and in topology optimization the error of surrogate-based optimizers can be as high as \(80\%\) without correction but below \(7\%\) with correction [2306.12047].

The variationally correct operator-learning literature makes a different but related point: standard PDE-residual losses are often not equivalent to the true solution error because of non-compliant norms and ad hoc boundary penalties. RBNO therefore constructs first-order system least-squares objectives whose residual values are provably equivalent to the solution error in PDE-induced norms, and constrains the network output to a conforming reduced basis. In this setting the residual loss serves as a reliable, computable a posteriori error estimator [2512.21319].

Learned preconditioners provide a third interpretation. The Born-series-inspired metric paper argues that the relevant residual geometry is not Euclidean for indefinite operators such as high-frequency Helmholtz. Using the identity
\[
I-G_\eta V_\eta = G_\eta A = L_\eta^{-1}A,
\]
it defines the metric \(R_\eta=G_\eta^\ast G_\eta\) and the loss
\[
\mathcal L_{\mathrm{bs}^{R_\eta}}(\theta)
=
\E_{\bm r\sim\mathcal D}
\frac{\|A\mathcal M_\theta(G_\eta \bm r)-\bm r\|_{R_\eta}}{\|\bm r\|_{R_\eta}},
\]
which matches the geometry used at inference time [2603.18527]. The adaptive variational framework generalizes this logic by linking residual transformations, sampling distributions, and target norms; exponential weighting corresponds to a softmax tilt and targets uniform-error reduction, while quadratic potentials recover linear residual weighting and a variance/\(L^2\)-type regime [2509.14198].

At a more abstract level, operator defect theory treats iterative approximation schemes as products of contractions \(T_n=A_nA_{n-1}\cdots A_1\) with defect operator \(D_A=(I-A^\ast A)^{1/2}\). The exact telescoping identity
\[
\|x\|^2=\|T_Nx\|^2+\sum_{n=1}^N\|D_{A_n}T_{n-1}x\|^2
\]
interprets residuals as operator-theoretically defined defect energies rather than heuristic errors. This framework is then applied to Kaczmarz-type methods, RKHS interpolation, kernel compression, and greedy KPCA [2601.18080].

## 5. Extensions beyond PDE surrogates

Residual operator learning has also been adopted in settings where the “baseline object” is not a PDE approximation. In multi-instance uncertainty estimation, MIREL derives an instance estimator \(T=g\circ f\) from the permutation-invariant bag classifier
\[
S(X)=g\big(\sum_{k=1}^K f(\mathbf x_k)\big),
\]
then corrects it by a residual branch
\[
R(\mathbf x)=T(\mathbf x)+r_\pi(f_\psi(\mathbf x)).
\]
The output is mapped to a Dirichlet evidential predictor by \(\bm\alpha_{\text{ins}}=\mathcal A(R(\mathbf x))+1\). On MNIST-bags, the ablation reports bag-level \(\overline{UE}=86.47\) and instance-level \(\overline{UE}=78.63\) for the residual formulation; on CAMELYON16, bag-level \(\overline{UE}=72.06\) and instance-level \(\overline{UE}=72.85\) are reported [2405.04405].

In inverse manipulation, the forward skill is first abstracted as a STRIPS-like operator \(o=\langle Pre,Add,Del\rangle\), and the inverse target is defined by
\[
\mathcal I(o)=Pre\cup Del\cup \neg Add.
\]
A BFS-based symbolic planner executes scripted primitives, yielding a handoff state \(s_h\); satisfied predicates become fences, unresolved predicates become the active residual objective, and a Soft Actor-Critic policy learns the continuous correction. On ManiSkill3 PushCube, the symbolic prefix alone yields \(16.6\pm 6.6\) mm mean distance and \(10\%\) success @ 1 cm, whereas symbolic + RL yields \(1.4\pm 3.2\) mm and \(90\%\) success [2606.05248].

In dexterous teleoperation, ResPilot begins with an optimization-based retargeter, HKVM, then learns fingerwise Gaussian-process residuals in an angle-aware 2D vector representation. The final command is
\[
\mathbf q_d[f]\gets \mathbf a\!\left(\pmb{\xi}^{f^\ast}+\mathbf V(\mathbf q_o^\ast[f])\right),
\]
so the GP corrects rather than replaces the baseline hand-mapping policy. The method uses 24 calibration poses, about 4.5 minutes total calibration and GP fitting time, GP inference of about 9 ms across all fingers, and reports joint workspace \(0.2399\ \mathrm{rad}^4\) versus \(0.2028\) for HKVM and fingertip workspace \(1511.05\ \mathrm{cm}^3\) versus \(1090.75\) [2409.09140].

Residual refinement is also used in long-horizon control. KORR keeps the standard executed action
\[
\mathbf a_{\text{exe},t}=\mathbf a_{\text{base},t}+\mathbf a_{\text{res},t},
\]
but conditions \(\mathbf a_{\text{res},t}\) on a Koopman-predicted next latent state
\[
\mathbf z^{\text{base}}_{t+1}=\mathbf A g_\theta(\mathbf x_t)+\mathbf B\,\mathbf a_{\text{base},t}.
\]
On FurnitureBench tasks, the reported gains are largest under stronger perturbations, for example One_Leg High with disturbance improves from \(3.20\) for ResiP to \(8.30\) for KORR, and Lamp Med with disturbance improves from \(29.98\) to \(39.16\) [2509.12562].

A different use appears in evolutionary multitasking, where MFEA-RL uses a VDSR model to generate a high-dimensional residual representation of an individual, a ResNet-based mechanism for dynamic skill-factor assignment, and a random mapping mechanism to project crossover back into the original decision space. The benchmarks are CEC2017-MTSO, WCCI2020-MTSO, and a Sensor Coverage Problem [2503.21347].

## 6. Empirical tendencies, misconceptions, and limitations

A recurrent empirical finding is that residual formulations are especially effective when the baseline captures large-scale structure and the learned residual is smaller, more regular, or statistically more diverse than the full target. DeltaPhi explicitly attributes its gains to auxiliary trajectories acting as a nonparametric physical prior and to reshaping the residual label distribution through retrieval range \(K\); the paper reports consistent improvements over direct learning on Darcy Flow, Navier–Stokes, and five irregular-domain tasks, with gains often larger when training data are scarce and under zero-shot resolution generalization [2406.09795]. The S-DeepONet-plus-diffusion paper makes the same point in generative form, arguing that diffusion can focus on sharpening high-frequency structures once the global low-frequency scaffold is supplied by the prior [2507.06133].

A common misconception is to equate residual operator learning with ordinary architectural skip connections. The literature is more specific. In several papers the residual object is an operator on discrepancy space, a variational corrector, or a solver increment with its own objective and inference semantics, not merely a residual block in a neural network [2306.12047][2603.18527][2511.19980]. Another recurrent clarification is that a small residual loss is not automatically physically meaningful; the variationally correct literature emphasizes that residual norms must be compatible with PDE-induced norms and boundary conditions if they are to function as error estimators [2512.21319].

The literature also records clear limitations. DeltaPhi states that benefits may be limited when input–output correlation is weak, as in chaotic systems, and that residual learning can be harder to optimize than direct learning, requiring sufficient backbone expressivity; for long-horizon time-series PDEs, gains may shrink as the relation between initial input and future output weakens [2406.09795]. The probabilistic residual-flow framework assumes access to a useful low-fidelity model [2512.12749]. KORR reports that non-linear dynamics predictors can hurt performance relative to Koopman linearity in future-state-guided residual refinement [2509.12562]. ResPilot notes that contact constraints are task-dependent and may hinder some dynamic rotations or pivoting tasks [2409.09140].

These results suggest a general, but not universal, principle. Residual operator learning is most effective when a baseline operator, prior, or planner already captures a substantial invariant component of the target behavior, leaving a correction problem that is simpler than the original one. When that decomposition is poorly aligned with the underlying dynamics, the residual formulation can lose its advantage.

Source: https://www.emergentmind.com/topics/residual-operator-learning