---
title: 'HADES-NN: Neural Estimation for Discontinuous Inputs'
url: https://www.emergentmind.com/topics/hades-nn
type: topic
---

# HADES-NN: Neural Estimation for Discontinuous Inputs

Searching arXiv for the HADES-NN paper and closely related context.
HADES-NN, short for **Harmonic Approximation of Discontinuous External Signals using Neural Networks**, is a parameter-estimation method for non-autonomous differential-equation models whose external driving signals are discontinuous or abruptly varying. It is formulated for systems of the form
$$
\frac{d}{dt}\vec{y}(t)=F\big(\vec{y}(t),\vec{p},S(t)\big),
$$
where $\vec{y}(t)$ is the state, $\vec{p}$ is the parameter vector, and $S(t)$ is an exogenous signal. The central idea is to replace a discontinuous input $S(t)$ by a smooth neural approximation $\tilde S_n(t)$ and to alternate between signal reconstruction and parameter fitting, thereby turning a non-smooth inverse problem into a sequence of smooth ones. The method was introduced in "Neural Network-Based Parameter Estimation for Non-Autonomous Differential Equations with Discontinuous Signals" [2507.06267].

## 1. Problem class and motivation

HADES-NN is designed for inverse problems in which model dynamics are externally forced and the forcing is not smooth. In the reported formulation, observations are available at times $\{t_j^o\}_{j=1}^N$, and the ideal target is
$$
\vec{p}_*=\arg\min_{\vec{p}\in\mathbb{R}^{D_p}}
\left(
\frac{1}{N}\sum_{j=1}^N\sum_{i=1}^{D_y}
\big[y_{\mathrm{obs},i}(t_j^o)-y_i(t_j^o;\vec{p},S(t_j^o))\big]^2
\right)^{1/2}.
$$
When $S(t)$ is smooth, this optimization is likewise smooth in $\vec p$ and standard methods such as Levenberg–Marquardt, L-BFGS, or related local solvers are effective. The distinctive regime addressed by HADES-NN is the one in which $S(t)$ is step-like, switching, binary, or otherwise discontinuous. In that setting, the induced loss landscape can exhibit kinks, many local minima, and non-differentiable structure, particularly when trajectory features move relative to the signal discontinuities [2507.06267].

The method is motivated by applications in which abrupt exogenous signals are intrinsic rather than pathological. The reported examples include a Lotka–Volterra predator–prey model driven by a binary Markov input, circadian clock dynamics regulated by external light measured via wearable devices, and a yeast mating-response network driven by pheromone exposure. In each case, the external signal is not a smooth nuisance term but a central component of the model class.

## 2. Two-stage iterative construction

HADES-NN operates in two iterated stages. In the first stage, the discontinuous signal is approximated by a neural network. Given signal observations $\{(t_i^s,S(t_i^s))\}_{i=1}^M$, the network $S_{\mathrm{NN}}(t;\theta_n)$ defines
$$
\tilde S_n(t)=S_{\mathrm{NN}}(t;\theta_n),
$$
and is trained by minimizing the empirical $L^2$ loss
$$
\mathcal{L}_S(\theta_n)=
\frac{1}{M}\sum_{i=1}^M\big|S(t_i^s)-\tilde S_n(t_i^s)\big|^2.
$$
In the second stage, $\tilde S_n(t)$ is treated as the external input to a smooth non-autonomous ODE,
$$
\frac{d}{dt}\vec y_n(t)=F\big(\vec y_n(t),\vec p,\tilde S_n(t)\big),
$$
and parameters are estimated through
$$
\vec{p}_n=\arg\min_{\vec{p}\in\mathbb{R}^{D_p}}
\left(
\frac{1}{N}\sum_{j=1}^N\sum_{i=1}^{D_y}
\left[
y_{\mathrm{obs},i}(t_j^o)-y_{n,i}\big(t_j^o;\vec p,\tilde S_n(t_j^o)\big)
\right]^2
\right)^{1/2}.
$$
The paper uses the Levenberg–Marquardt algorithm for this second stage [2507.06267].

The iteration is warm-started in both blocks. For $n>1$, the signal network is initialized from the previous iteration’s parameters, and the Levenberg–Marquardt step is initialized from $\vec p_{n-1}$. Iteration stops when
$$
\left(\sum_{k=1}^{D_p}(p_{n,k}-p_{n-1,k})^2\right)^{1/2}\le \varepsilon.
$$
This architecture makes HADES-NN an outer-loop framework around a conventional solver rather than an end-to-end neural replacement of the governing equations.

The “harmonic” component of the method refers to the signal-approximation network. Its first hidden layer contains 16 cosine units,
$$
\phi_l(t)=\cos(\omega_l t+\phi_l),\qquad l=1,\dots,16,
$$
followed by five fully connected hidden layers with 16 ELU units each and a linear output layer. In the reported interpretation, this first layer acts as a finite harmonic or Fourier-type basis for representing non-smooth temporal structure.

## 3. Mathematical structure and convergence theory

The theoretical analysis of HADES-NN rests on three elements: regularity of the vector field, approximation of the external signal, and convergence of the smooth optimizer. The reported assumptions include identifiability of the target parameter $\vec p_*$ and Lipschitz continuity of $F$ in both $\vec y$ and $S$. Concretely, there exist constants $C_1$ and $C_2$ such that
$$
\|F(\vec y_1(t),\vec p,S(t))-F(\vec y_2(t),\vec p,S(t))\|_2
\le C_1\|\vec y_1(t)-\vec y_2(t)\|_2,
$$
and
$$
\|F(\vec y(t),\vec p,S_1(t))-F(\vec y(t),\vec p,S_2(t))\|_2
\le C_2|S_1(t)-S_2(t)|.
$$

A key lemma establishes continuity of the ODE solution with respect to the external signal in $L^2$: if two systems share the same initial condition and parameter vector but use signals $S_1$ and $S_2$, then
$$
\|\vec y_1-\vec y_2\|_{L_2([0,T])}
\le C(T)\,\|S_1-S_2\|_{L_2([0,T])}.
$$
This supplies the bridge between the neural approximation of the forcing and the underlying parameter-estimation problem [2507.06267].

The reported convergence theorem states that if $\tilde S_n\to S$ in $L^2$, then the sequence of global minimizers $\vec p_n$ of the smoothed problems converges to the true minimizer $\vec p_*$ of the original discontinuous-input problem. The argument combines the continuity lemma with the uniqueness of the minimizer. The theoretical discussion also invokes sensitivity equations for derivatives $z_{ij}=\partial y_i/\partial p_j$, supporting the use of smooth local solvers once the forcing has been replaced by $\tilde S_n(t)$.

This suggests that HADES-NN is best understood as a regularization strategy on the exogenous signal rather than on the parameter space itself. The smoothing is not imposed directly on the loss; it is introduced through the forcing term that generates the trajectories.

## 4. Reported empirical applications

The first application is a Lotka–Volterra predator–prey system driven by a binary Markov input $S(t)\in\{0,1\}$ with transition matrix
$$
T=
\begin{bmatrix}
0.95 & 0.05\\
0.05 & 0.95
\end{bmatrix}.
$$
In the smooth-input case, standard optimizers recover the parameters accurately from sparse sampling. In the discontinuous-input case, the reported loss contours become jagged and classical optimizers fail badly. HADES-NN improves as the outer iteration proceeds: at early iterations the signal approximation is poor and the parameter estimates are inaccurate, whereas by later iterations both $\tilde S_n(t)$ and the inferred parameters approach the truth. For sampling densities $n_s=20$ and $n_s=80$, the reported comparison indicates that HADES-NN and NeuralODE converge toward the true parameters, with HADES-NN giving the lowest Mean Absolute Percentage Error [2507.06267].

The second application concerns the Forger–Jewett–Kronauer circadian pacemaker model driven by real light exposure measured by Actiwatch2. The inferred parameter vector is
$$
\vec p=(\tau_c,\gamma,G,k),
$$
where $\tau_c$ is the intrinsic circadian period, $\gamma$ the oscillator stiffness, $G$ a gain factor, and $k$ a relative-effect parameter. Using 80 samples of a normalized core-body-temperature proxy, the reported benchmark compares HADES-NN with NeuralODE, LM, L-BFGS, SLSQP, Nelder–Mead, and Differential Evolution across 100 runs with wide initializations. The reported outcome is that only HADES-NN consistently recovers accurate and tight parameter estimates near the true values, whereas NeuralODE, LM, and the other optimizers fail or exhibit large scatter.

The third application is a yeast mating-response network with 7 ODEs and 23 unknown parameters, using GFP trajectories under multiple pheromone protocols. Relative to a previously used evolutionary algorithm, HADES-NN produces markedly narrower parameter distributions while maintaining good agreement between simulated and observed GFP trajectories. The reported interpretation is not merely improved point estimation but increased parameter precision in a biologically structured inverse problem.

## 5. Methodological position and common interpretations

A common misconception is to classify HADES-NN as an ordinary NeuralODE method. In the reported construction, that description is not accurate. The second stage still solves the governing ODE numerically and performs parameter optimization with Levenberg–Marquardt; the neural network is used to approximate the discontinuous external signal, thereby restoring smooth dependence of the residuals on the parameters. HADES-NN is therefore a hybrid inverse-problem framework rather than a pure learned dynamical system.

Its methodological advantages, as reported, are tied specifically to discontinuous forcing. It does not claim universal superiority for smooth-input problems. Instead, it targets the regime in which abrupt switching corrupts the geometry of the optimization landscape. The reported limitations are correspondingly conventional for inverse problems: computational cost due to repeated ODE solves, sensitivity to neural-network capacity and outer-loop tolerance, the practical risk of local minima despite the convergence theorem being stated for global minimizers, and the inability to resolve structural non-identifiability when the underlying model is unidentifiable [2507.06267].

The paper also treats HADES-NN as distinct from smoothing the observed state trajectories. The smoothing acts on $S(t)$, not on $\vec y_{\mathrm{obs}}(t)$, preserving the mechanistic role of the forcing in the differential equation. This design choice is central to the claim that the method extends the range of model systems that can be fitted to measurements involving abrupt exogenous signals.

## 6. Implementation profile, extensions, and scope

The reported implementation requires three ingredients: time series for the observed state variables, sampled values of the external signal, and a mechanistic non-autonomous ODE model. In practice, the procedure trains the signal network on the signal samples, freezes the resulting $\tilde S_n(t)$ for the current outer iteration, and then solves the parameter-estimation problem with standard numerical integration and Levenberg–Marquardt. The practical guidance given in the paper emphasizes warm starts, diagnostic comparison of $\tilde S_n(t)$ with the measured signal, and residual inspection after fitting.

The computational burden is dominated by the repeated ODE solves in the second stage rather than by the signal network itself, whose architecture is small. This suggests that HADES-NN is particularly attractive when the parameter dimension is moderate and the forward model remains cheaper than brute-force derivative-free exploration of a non-smooth loss surface.

Several extensions are explicitly proposed. These include replacing the current harmonic-feature network with alternative periodic architectures such as SIREN-style models or random Fourier features, incorporating the approach into Physics-Informed Neural Networks, jointly optimizing the signal approximation and model parameters through a unified objective, and extending the underlying idea to stochastic differential equations or partial differential equations with discontinuous forcing [2507.06267]. A plausible implication is that the method belongs to a broader class of neural-assisted regularization schemes for mechanistic inference, but in its published form it remains specifically an iterative procedure for non-autonomous ODE parameter estimation under discontinuous external inputs.

Source: https://www.emergentmind.com/topics/hades-nn