---
title: Damped Residual ODEs
url: https://www.emergentmind.com/topics/damped-residual-odes
type: topic
---

# Damped Residual ODEs

Damped residual ODEs are a class of continuous-depth neural network models in which explicit linear damping terms are introduced to the governing ordinary differential equations associated with residual and non-residual network architectures. These damped ODE formulations provide a unified mathematical framework for interpreting, analyzing, and interpolating between deep residual networks (ResNets), convolutional neural networks (CNNs), and their continuous analogs—Neural ODEs (NODEs)—by modulating the damping effect via tunable parameters. Rigorous analysis of such models has led to new results in interpolation, universal approximation, control-theoretic characterization, stability, and robustness gains under stochastic and adversarial perturbations [2110.08761], [2006.05749].

## 1. Damped Residual ODE Models

Central to the theory is the modification of the standard ResNet ODE
\[
\frac{d\,x(t)}{dt}=f\bigl(x(t),t;\theta\bigr)
\]
to include a linear damping term parameterized by a nonnegative interpolation coefficient $\lambda$:
\[
\frac{d\,x(t)}{dt} = -\lambda\,x(t) + \rho(\lambda)\,f(x(t),t;\theta),\qquad x(0)=x_0.
\]
Here, $\rho(\lambda)$ is a nonnegative scalar function with asymptotics $\rho(\lambda)\to 1$ as $\lambda\to 0^+$ and $\rho(\lambda)\sim \lambda$ as $\lambda\to+\infty$. The discrete-step update corresponding to a uniform discretization with step size $\Delta t=1$ is
\[
x_{n+1} = e^{-\lambda}x_n + \frac{1-e^{-\lambda}}{\lambda}\,\rho(\lambda)f_n(x_n).
\]
This model smoothly interpolates between residual and non-residual architectures:
- $\lambda\to 0$: recovers the classic ResNet block $x_{n+1}=x_n+f_n(x_n)$,
- $\lambda\gg 1$: recovers a plain CNN layer $x_{n+1}=f_n(x_n)$ (up to scaling).

More sophisticated “damped” NODEs arise in the analysis of Momentum ResNets and memory-augmented models. For example, the momentum-augmented NODE has
\[
\ddot{x}(t)+\dot{x}(t) = w(t)\,\sigma(\langle a(t),x(t)\rangle+b(t)),
\]
with $p=\dot{x}$ defining a first-order system with explicit velocity damping [2110.08761].

## 2. Approximation, Interpolation, and Controllability Results

Damped residual ODEs exhibit strong theoretical properties with respect to interpolation and universal approximation. These models are formulated as controlled ODEs $\dot{Z}=F(Z;u(t))$, with trainable, time-dependent controls $u(t)$.

### Main Theorems for Momentum ResNets [2110.08761]:
- **Simultaneous controllability/interpolation:** For any finite set of initial conditions $(x_i,p_i)$ and targets $y_i$, there exist bounded controls $(w,a,b)$ such that the solution $x(T;x_i,p_i)=y_i$ for all $i$.
- **Universal $L^2$-approximation:** For any $f\in L^2(\Omega;\mathbb{R}^d)$ and $\varepsilon>0$, there exists $T_{\min}>0$ such that for all $T\geq T_{\min}$, there are controls $(w,a,b)$ yielding $\|f-x(T;\cdot,0)\|_{L^2(\Omega)}<\varepsilon$.

### Memory-Augmented NODEs [2110.08761]:
- **Simultaneous state–memory control:** For any finite collection $\{(x_i,p_i)\},\{(y_i,\varphi_i)\}$, there exist controls such that $(x(T),p(T))=(y_i,\varphi_i)$ for each $i$.
- **Simultaneous tracking controllability:** For suitably large memory dimension $d_p\ge 2d$, and continuous–BV target maps $M(\cdot,t)$, trajectories $x_i(t)$ can be steered to track any set of continuous curves $y_i(\cdot)$ within uniform error $\varepsilon$ over $[0,T]$.
- **Universal tracking approximation:** The flow $x(t;\cdot)$ can uniformly approximate $M(\cdot,t)$ in $L^2(\Omega)$ to any precision.

## 3. Stability and Control-Theoretic Insights

Damping imparts favorable stability properties to the dynamical system underlying the neural architecture. For time-invariant $f$, equilibria $x^*$ with $f(x^*)=0$ are locally exponentially stable if
\[
\lambda > L\,\rho(\lambda)
\]
where $L$ is the local Lipschitz constant of $f$.
Linearization yields a Jacobian $J_\lambda(x^*)=\rho(\lambda)\,\partial_x f(x^*)-\lambda I$, shifting all eigenvalues by $-\lambda$ in the complex plane. This leftward spectral shift ensures that for sufficiently large $\lambda$, all perturbations are suppressed and the equilibrium is stable.

These results generalize to the memory-augmented and momentum-damped NODE settings, where the simultaneous-control framework enables intricate schemes for mesh interpolation, contraction flows, marker-based approximation, and piecewise-linear trajectory tracking [2110.08761].

## 4. Simultaneous-Control Viewpoint and Proof Strategies

Both the momentum and memory-augmented NODEs are analyzed through the lens of simultaneous controllability: employing a single control law $u(t)$ to steer multiple initial states simultaneously to prescribed terminal targets. The “toolbox” for constructing precise flows involves:
- **Inactive region dynamics:** In regions where the ReLU is inactive, the second-order ODE reduces to free damping: $p\mapsto p\,e^{-t}$, $x\mapsto x + (1-e^{-t})p$.
- **Active region dynamics:** Damped oscillators or real-eigenvalue contraction flows drive specific coordinates.
- **Piecewise control:** In the interpolation proof, time is partitioned and appropriate controls ensure marker-points are moved or compressed as needed before applying contraction to achieve global accuracy.
- **Sequential tracking:** Memory variables are alternately configured to induce exact piecewise-linear drifts for time-dependent trajectory tracking.

## 5. Empirical Evaluation and Simulation Studies

Experiments and simulations reinforce the theoretical guarantees and showcase the practical advantages of damped residual ODEs.

### Binary Disk Classification [2110.08761]:
- **Setup:** $N=100$ points sampled from $[-1,1]^2$, with the target a disk indicator.
- **Results:** Momentum ResNet (second-order with damping) achieves $L^1$ error $\approx 0.06$, capturing the nontrivial topology, whereas standard NODEs (first-order) yield $L^1\approx 0.11$ and cannot produce closed boundaries.

### Function Tracking [2110.08761]:
- **Task:** Tracking $M(x)(t)=\sin(x\,t)$ over $x\in(0,1)$, $N=5$ samples, using memory NODE ($d_p=2$).
- **Results:** Simultaneous tracking error $\lesssim 0.063$; global tracking error $\lesssim 0.11$.

### Robustness to Corruption and Attack [2006.05749]:
- **Noise robustness:** On CIFAR-10, In-ResNet-110 attains $72.7\%$ accuracy under four noise types (vs. $53.7\%$ for ResNet-110).
- **Adversarial robustness:** Under FGSM ($\epsilon=2/255$), In-ResNet-110 scores $55.2\%$ (vs. $41.5\%$ for ResNet-110); under PGD, $31.7\%$ (vs. $5.6\%$).
- **Loss landscape:** Damped models exhibit flatter, more stable loss landscapes under attack.

## 6. Relationships to Classical Architectures and Unified View

Damped residual ODEs subsume both ResNets and plain CNNs as limiting cases by continuous variation of the damping parameter $\lambda$:
- $\lambda\approx 0$: classical residual network behavior.
- $\lambda\gg 1$: non-residual, feedforward CNN regime.

The parameter $\lambda$ can be learned per layer, yielding architectures that adaptively interpolate between skip-connection dominance and full feature transform, while achieving smoothly tunable robustness and stability [2006.05749].

| Model/Regime           | $\lambda$ Value   | Discrete Update                | Architecture Recovered |
|------------------------|-------------------|-------------------------------|-----------------------|
| ResNet                 | $\lambda\to 0$    | $x_{n+1}=x_n+f_n(x_n)$        | Residual network      |
| Damped (General)       | $\lambda>0$       | $x_{n+1}=e^{-\lambda}x_n+\ldots$ | Interpolated         |
| CNN (Plain)            | $\lambda\to +\infty$ | $x_{n+1}=f_n(x_n)$            | Non-residual CNN      |

A plausible implication is that damped ODE frameworks present an overarching control-theoretic formalism for depth-continuous machine learning, encompassing, extending, and theoretically bridging the properties of both residual and non-residual deep network architectures.

## 7. References

- Ruiz-Balet, A.; Affili, E.; Zuazua, E. "Interpolation and approximation via Momentum ResNets and Neural ODEs" [2110.08761]
- Chen, R.T.Q., et al. "Neural ordinary differential equations," NeurIPS 2018.
- Sander, T.; Haber, E.; Rumpf, M. "Lagrangian deep learning," ICLR 2021.
- Ruiz-Balet, A.; Zuazua, E. "Simultaneous control of First-Order Neural ODEs," J. Math. Control 2021.
- "Interpolation between Residual and Non-Residual Networks" [2006.05749]

Source: https://www.emergentmind.com/topics/damped-residual-odes