---
title: Rational Neural Networks
url: https://www.emergentmind.com/topics/rational-neural-networks
type: topic
---

# Rational Neural Networks

Searching arXiv for recent papers on rational neural networks and closely related formulations.
Rational neural networks are neural-network models in which rational functions—ratios of polynomials—enter either as trainable activation functions, as the explicit parameterization of layers and filters, or as the algebraic objects used to characterize representable maps. In the literature, the term covers feedforward networks with activations of the form \(R(x)=P(x)/Q(x)\), controller and symbolic-regression architectures whose units are directly given by polynomial numerators and denominators, graph networks with rational spectral filters, rational Gaussian wavelet layers, and rational-weight ReLU networks studied through exact depth lower bounds [2004.01902], [2307.06287], [1808.10073], [2502.06283]. This suggests that rational neural networks are best viewed as a family of algebraically structured models rather than a single canonical architecture.

## 1. Terminological scope and basic forms

A common definition makes the nonlinearity itself rational. In this formulation, each activation is a learnable function
\[
\operatorname{R}(x)=\frac{\operatorname{P}(x)}{\operatorname{Q}(x)}=\frac{\sum_{j=0}^{m} a_{j} x^{j}}{1 + \sum_{k=1}^{n} b_{k} x^{k}},
\]
or, equivalently, a low-degree rational function
\[
F(x)=\frac{\sum_{i=0}^{r_P} a_i x^i}{\sum_{j=0}^{r_Q} b_j x^j},
\]
with trainable numerator and denominator coefficients [2102.09407], [2004.01902]. In this line of work, the attraction of rationals is their ability to let each neuron adapt its nonlinear shape during optimization, rather than fixing the activation to ReLU, sigmoid, or tanh.

A second formulation makes the entire layer rational. In the control-oriented parameterization, one writes
\[
x_{i}^{k+1}=\frac{p_i(x^k)}{q_i(x^k)},
\qquad
\pi_i(u)=\frac{p_i(x^\ell)}{q_i(x^\ell)},
\]
so that the trainable objects are polynomial coefficients inside numerators and denominators rather than affine weights followed by a separate activation [2307.06287]. In symbolic-regression-oriented models, the output itself is directly constrained to be rational, for example
\[
\hat y=\frac{E_N C}{E_D d},
\]
with polynomial basis terms explicitly enumerated in the input layer and sparsity imposed on the coefficient vectors [2109.08813].

A third usage is architectural rather than neuron-wise. Rational Gaussian wavelet models define a mother wavelet by
\[
\psi^{\boldsymbol\eta}(t)=C(\boldsymbol\eta)\,P^{\boldsymbol\eta}(t)\,v^{\boldsymbol\eta}(t)\,e^{-t^2/2},
\]
where the zeros of the polynomial factor and the poles of the rational factor become trainable parameters that shape the feature extractor [2502.01282]. Spectral graph models similarly replace polynomial graph filters by rational ones of the form \(P(\Lambda)/Q(\Lambda)\) [1808.10073].

A fourth usage concerns weights rather than activations. In the theory of rational ReLU networks, the activation remains ReLU, but weights are restricted to rational forms such as decimal fractions or, more generally, \(N\)-ary fractions [2502.06283]. This version is central when exact representability and depth lower bounds are the main object of study.

Across these variants, pole avoidance is a recurring technical constraint. Safe rational functions with absolute values in the denominator, positivity-enforcing parameterizations, and softplus-based positive denominators are repeatedly introduced to prevent real poles and training instability [2102.09407], [2004.01902], [2602.04006].

## 2. Approximation-theoretic foundations

The modern theory begins from the observation that rational functions and ReLU networks approximate each other surprisingly efficiently. For any ReLU network, there exists a rational function of degree \(O(\mathrm{polylog}(1/\epsilon))\) that is \(\epsilon\)-close, and for suitable rational functions there exists a ReLU network of size \(O(\mathrm{polylog}(1/\epsilon))\) that is \(\epsilon\)-close [1706.03301]. The key classical ingredient is Newman’s approximation of \(|x|\), from which
\[
(x)_+ = \frac{x+|x|}{2}
\]
yields efficient rational approximants to ReLU. By contrast, polynomials require degree \(\Omega(\mathrm{poly}(1/\epsilon))\) to approximate even a single ReLU [1706.03301].

This bidirectional relationship was sharpened by work that treated rational activations as a native network primitive rather than as an external surrogate. Composing \(k\) rational functions of type \((3,2)\) yields a rational function of type \((3^k,3^k-1)\) while using only \(7k\) parameters, so the represented degree grows exponentially with depth although the number of trainable coefficients grows only linearly [2004.01902]. The same paper proves that smooth Sobolev functions can be approximated by rational networks with size
\[
\mathcal{O}\!\left(\epsilon^{-d/n}\log(\log(1/\epsilon))\right)
\]
and depth
\[
\mathcal{O}\!\left(\log(\log(1/\epsilon))\right),
\]
which is presented as an exponential depth saving relative to comparable ReLU constructions [2004.01902]. Empirically, the same study reported mean squared error \(1.2\times10^{-7}\) for a rational network on a Korteweg–de Vries approximation task, compared with \(1.9\times10^{-4}\) for ReLU [2004.01902].

A later approximation-theoretic program extends the comparison from ReLU to a broad family of modern fixed activations. It shows that any network built from standard fixed activations can be uniformly approximated on compact domains by a rational-activation network with only \(\mathrm{poly}(\log\log(1/\varepsilon))\) overhead in size, while the converse provably requires \(\Omega(\log(1/\varepsilon))\) parameters in the worst case [2602.12390]. The separation is stated not only for scalar activations such as GELU, SiLU, Mish, ELU, SELU, CELU, Softplus, Sigmoid, Tanh, Softmin, Softmax, and LogSoftmax, but also for gated activations and transformer-style nonlinearities [2602.12390].

Rational approximation has also been pushed into derivative-sensitive regimes. For \(ReQU(x)=\max(x,0)^2\), there exist rational functions of type \((n+1,n-1)\) with
\[
\|ReQU-\tilde R_n\|_{\mathcal C([-1,1])}=\mathcal O(e^{-\sqrt n})
\]
and
\[
\|ReQU'-\tilde R_n'\|_{\mathcal C([-1,1])}=\mathcal O(n^{-(K-2)/2})
\]
for every \(K\ge 3\) [2508.19672]. Using these approximants, suitable Hölder-smooth functions \(f:[0,1]^d\to\mathbb R^p\) can be approximated in \(\mathcal C^1\) by rational neural networks of width of order \(N\), constant depth, and maximal rational degree of order \(N^\epsilon\), with
\[
\|f-\mathcal R_f^{\lfloor\beta\rfloor,N}\|_{\mathcal C^1([0,1]^d)}\le cN^{-(\beta-1)}.
\]
The same framework yields \(\mathcal C^1\)-approximation results for the \(\mathrm{EQL}^{\div}\) and ParFam architectures used in symbolic regression and physical law learning [2508.19672].

## 3. Exact representability and depth lower bounds

A distinct branch of the literature studies exact representation rather than approximation, and here the test function
\[
F_n(x_1,\ldots,x_n)=\max\{0,x_1,\ldots,x_n\}
\]
is central because it is the support function of the standard simplex \(\Delta_n=\operatorname{conv}(0,e_1,\dots,e_n)\) [2502.06283]. The underlying conjecture, due to Hertrich, Basu, Di Summa, and Skutella, predicts that any ReLU network exactly representing \(F_n\) should need at least \(\lceil \log_2(n+1)\rceil\) hidden layers.

For rationally representable weights of arithmetic relevance, the strongest general theorem currently available states that if \(p\) is a prime not dividing \(N\), then every ReLU network whose weights are \(N\)-ary fractions needs at least
\[
\left\lceil \log_p(n+1)\right\rceil
\]
hidden layers to exactly represent \(F_n\) [2502.06283]. For decimal fractions, choosing \(p=3\) yields the concrete lower bound
\[
\left\lceil \log_3(n+1)\right\rceil.
\]
The same paper also proves a denominator-sensitive asymptotic lower bound: there exists a constant \(C>0\) such that for all \(n,N\ge 3\), every ReLU network with \(N\)-ary fraction weights that exactly represents \(F_n\) has depth at least
\[
C\cdot \frac{\ln n}{\ln\ln N}.
\]
These are presented as the first non-constant lower bounds on the depth of practically relevant rational-weight ReLU networks [2502.06283].

The proof combines a polyhedral characterization of positively homogeneous ReLU representability with a modular volume obstruction. Representability is expressed through the sum-union closure \(SU^k(P_0(R^n))\) of point polytopes, together with the equivalence
\[
h_P\in ReLU_n^R(k)\quad\Longleftrightarrow\quad P+A=B \text{ for some }A,B\in SU^k(P_0(R^n)).
\]
A denominator-clearing lemma then reduces rational weights to integer weights:
\[
M^{k+1}f \in ReLU_n^{\mathbb Z}(k)
\]
whenever \(f\) is representable with \(k\) hidden layers and all rational weights have common denominator \(M\) [2502.06283]. The contradiction is obtained by showing that certain normalized face volumes must be divisible by a prime \(p\), whereas the simplex has normalized volume \(1\).

These results are explicitly described as a partial confirmation of the original conjecture. They establish that depth must grow with \(n\) for exact max computation in rational settings, but they do not settle the \(\lceil\log_2(n+1)\rceil\) bound for arbitrary real weights [2502.06283].

## 4. Architecture families and empirical domains

In graph learning, RationalNet replaces polynomial spectral filters by rational ones,
\[
\mathcal g * x = \mathcal P(\mathcal L)\mathcal Q(\mathcal L)^{-1}x,
\]
in order to better approximate jump discontinuities and non-smooth high-pass behavior [1808.10073]. The motivation is classical Gibbs-type failure of Chebyshev approximants at discontinuities. The paper gives a rational convergence rate
\[
\sup_{x\in[-c,c]} |f_{1,2} - R_n(x)| \leq C e^{-\sqrt{n}},
\]
contrasted with a polynomial rate
\[
\sup_{x\in[-c,c]} |f_{1,2} - P_n(x)| \le \frac{C\beta}{n},
\]
and uses a relaxed Remez algorithm to initialize the rational coefficients [1808.10073]. On a 1000-node synthetic graph, RationalNet achieved spectral MSE around \(5\times 10^{-6}\) on \(|x|\), and Remez initialization improved spectral MSE by 56.26% for \(|x|\) and 81.39% for \(\operatorname{sign}(x)\) [1808.10073].

Wavelet-based rational neural networks form another important family. Rational Gaussian wavelets define an adaptive mother wavelet
\[
\psi^{\boldsymbol\eta}(t)=C(\boldsymbol\eta)\,P^{\boldsymbol\eta}(t)\,v^{\boldsymbol\eta}(t)\,e^{-t^2/2},
\]
where the zeros of \(P^{\boldsymbol\eta}\) and the poles of \(v^{\boldsymbol\eta}\) directly control wavelet morphology [2502.01282]. A variable-projection layer
\[
\mathcal C_{\Psi(\boldsymbol\eta)}(\boldsymbol f)=\Psi(\boldsymbol\eta)^+\boldsymbol f
\]
then produces interpretable wavelet coefficients, while the reconstruction map
\[
\mathcal P_{\Psi(\boldsymbol\eta)}(\boldsymbol f)=\Psi(\boldsymbol\eta)\big(\Psi(\boldsymbol\eta)^+\boldsymbol f\big)
\]
supports an additional regularizer [2502.01282]. On ventricular ectopic beat detection from MIT-BIH ECG data, the reported setup used \(m=10\) wavelet coefficients, \(p=3\) zeros, and \(n=4\) poles, and achieved 98.51% total accuracy [2502.01282].

The same rational Gaussian wavelet idea was later transferred from variable projection to discrete convolution for nonperiodic acoustic segments in UAV recognition [2605.26310]. The RGW convolution layer maps a sampled signal to multiple scale-dependent responses, after which top-\(Q\) pooling keeps the dominant coefficients. In indoor swarm detection, indoor drone classification, and noisy outdoor detection, the reported mean accuracies were 92.52%, 99.20%, and 90.55%, respectively; in the outdoor scenario, the proposed model outperformed CNN, with CNN at 89.60% [2605.26310].

In scientific model discovery, rational function neural networks are used as symbolic-form estimators rather than purely predictive models. RafNN constrains the output to a rational expression with explicitly enumerated numerator and denominator basis terms, uses sparsity penalties
\[
E(\theta)=\left(y-\hat y\right)^2+\lambda_1|C|+\lambda_2|d|,
\]
and applies thresholding to prune small coefficients [2109.08813]. On synthetic rock-physics data with 1% random noise, the method reconstructed Gassmann’s equation; after 35 independent training runs, the active term counts converged to \(N_c=4\) and \(N_d=4\), and the reported recovered coefficients differed from the target values by below 0.03%, with best-case differences of 0.01% and 0.004% [2109.08813].

A recent architectural synthesis is the Rational-ANOVA Network, which combines Padé-style rational units with a functional-ANOVA interaction topology,
\[
f(x)\approx \sum_{i=1}^d r_i(x_i) + \sum_{(i,j)\in S} r_{ij}(x_i,x_j),
\]
and enforces a strictly positive denominator
\[
d_i(x)=1+\mathrm{softplus}\!\left(\sum_{b=1}^{n}\beta_b x^b\right)+\epsilon
\]
to avoid poles [2602.04006]. Under matched budgets, the paper reports roughly 59.05% on CIFAR-10 at around 1.0M parameters, compared with 56.95% for MLP and 56.45% for KAN, and a ViT-Tiny variant whose top-1 accuracy rises from 72.3% to 74.2% when the FFN is replaced by RAN [2602.04006].

## 5. Optimization and training regimes

One influential theme treats rationality as adaptive activation plasticity. In deep reinforcement learning, rational activations of order \((m,n)\) are proposed as trainable replacements for static nonlinearities, and a central theorem states that a rational function embeds a residual connection if and only if \(m>n\) [2102.09407]. The same work introduces a regularized parameter-sharing version, called joint-rational, in which one rational activation is shared across layers. In Atari experiments, replacing fixed activations by rational ones led to consistent improvements for DQN, with joint-rational often performing especially well on stationary games and making simple DQN competitive with DDQN and Rainbow [2102.09407].

Another training line fixes the rational activation coefficients and changes the loss. For the one-degree activation
\[
R(x)=\frac{a_0+a_1x}{b_0+b_1x},
\qquad b_0+b_1x>0,
\]
with coefficients chosen as the best rational \((1,1)\) approximation to ReLU on \([-1,1]\), network training under the uniform loss
\[
L({\bf W},b)=\max_i \left| y^i-R({\bf W}x^i+b)\right|
\]
becomes a quasiconvex generalized rational uniform approximation problem [2111.02602]. This permits bisection and differential-correction methods instead of least-squares optimization. On TwoLeadECG, the reported test accuracies were 87.71% for bisection, 70.2% for the MATLAB toolbox, and 55.31% for differential correction; on an imbalanced SonyAIBORobotSurface1 split with class 1 underrepresented, the corresponding values were 73.54%, 52.2%, and 65.06% [2111.02602].

In control, rational parameterization is used to make neural feedback loops compatible with Sum of Squares programming. Standard rational activations such as
\[
\mathrm{Rtanh}(x)=\frac{4x}{x^2+4},
\qquad
\mathrm{Rsig}(x)=\frac{(x+4)^2}{2(x^2+16)}
\]
satisfy exact polynomial equalities after clearing denominators, which is advantageous for Positivstellensatz-based stability certificates [2307.06287]. The same work proposes a more general rational neural network structure that is convex in the network parameters and a refined architecture with state-dependent denominators to avoid numerical recovery issues in SOS synthesis. The resulting method recovers stabilizing rational neural network controllers for unstable and nonlinear plants with saturation, noise, and parametric uncertainty [2307.06287].

## 6. Algebraic, logical, and geometric theories

The rational perspective has also produced a substantial formal theory. In continuous-time systems, a broad class of recurrent neural networks can be embedded into rational or polynomial systems under mild assumptions on the activation function [1903.05609]. The standing hypothesis is that the activation is analytic and satisfies a differential-algebraic closure property, equivalently a nontrivial polynomial relation among \(\sigma,\sigma^{(1)},\ldots,\sigma^{(N)}\). This includes \(\tanh\) and the sigmoid \(S(x)=1/(1+e^{-x})\), since \(y'=1-y^2\) and \(y'=y(1-y)\), respectively [1903.05609]. The embedding transfers questions of realizability, minimality, reachability, and observability from RNNs to the established realization theory of rational systems.

In graph logic, however, rationality does not simply enlarge expressivity. The logic of rational graph neural networks studies message-passing GNNs whose combination networks use rational activations \(R(X_1,\dots,X_m)=P/Q\) with no real pole [2310.13139]. The main negative result is that some depth-3 \(GC2\) queries cannot be expressed by any rational GNN, even though ReLU GNNs are known to capture the full logical fragment \(GC2\) uniformly [2310.13139]. To delimit the positive fragment, the paper defines \(RGC2\) and proves that rational GNNs can express every query in that fragment uniformly over all graphs [2310.13139].

A complementary logic-theoretic result shows that rational-weight ReLU networks admit fuzzy-logic characterizations. Using a scaling map
\[
scale_k(x)=\frac{k+x}{2k},
\]
the paper proves that rational-weight feedforward ReLU networks have the same expressive power, with respect to scaling, as Rational Pavelka Logic \(RPL\), fragments of \(\mathit{L\Pi}\frac12\), and a generalized polynomial ring over \(\mathbb Q\) in countably many variables with \(ReLU\) permitted [2605.03064]. The translation is constructive in both directions and identifies proto-neurons with degree-\(\le 1\) terms in the generalized polynomial structure [2605.03064].

Algebraic geometry supplies another viewpoint. For RationalNets with activation \(\sigma(x)=1/x\), the output of a fixed architecture can be written as a tuple of homogeneous polynomial fractions with common denominator, and the set of all such outputs is called the neuromanifold; its Zariski closure is the neurovariety [2509.11088]. For one-hidden-layer architectures \(d=(2,m,k)\), the paper shows that the Zariski closure is filling but the actual neuromanifold is not; for deep binary RationalNets, it classifies when the Zariski closure fills the ambient space and proves that this happens only for architectures of the form \(d=(2,\dots,2,1)\) [2509.11088]. Membership algorithms are also given for deciding whether a prescribed rational function belongs to the neuromanifold [2509.11088].

A related algebraic-combinatorial analysis links rational neural networks to VC-theory. Using the Erzeugungsgrad of Boolean algebras of constructible sets and degree notions for constructible families, one obtains bounds in which VC-dimension and Krull dimension are linearly related up to logarithmic factors [2504.11345]. These results are then applied to parameterized families of neural networks with rational activation function, yielding bounds on growth functions, VC-dimension, and densities of correct test sequences [2504.11345].

## 7. Limitations and open problems

Despite their approximation-theoretic strength, rational neural networks are not uniformly dominant across all formulations. A direct comparison between classical rational approximation and neural networks with rational activations reports that direct rational approximation is consistently more accurate than neural-network-based approximation when both are given the same number of decision variables, including on nonsmooth and non-Lipschitz targets [2303.04436]. This suggests that, in some regimes, the compositional neural parameterization is a restriction rather than an advantage.

Training stability remains a persistent concern. Unconstrained rational functions can develop poles and unstable gradients, which is why many implementations adopt absolute values in denominators, positivity constraints, or softplus-based denominator parameterizations [2004.01902], [2102.09407], [2602.04006]. Recent work on adaptive rational activations also argues that normalization layers can interfere with adaptive rationals by introducing non-identifiability and stochasticity that interact badly with coefficient sensitivity [2602.12390].

Expressivity is also architecture-dependent. Rational activations do not make GNNs maximally expressive in the uniform logical sense, since some \(GC2\) queries remain unattainable [2310.13139]. In exact-representation theory, current depth lower bounds for rational-weight ReLU networks stop short of the conjectured \(\lceil\log_2(n+1)\rceil\) barrier for arbitrary real weights, so the general real-weight case remains open [2502.06283]. In \(\mathcal C^1\)-approximation, the extension from \(ReQU\) to \(RePU_p\) is conjectured for \(p\ge 3\) but not proved [2508.19672].

The cumulative picture is therefore differentiated rather than uniform. Rational neural networks offer unusually strong tools for approximation of non-smooth structure, exact algebraic modeling, symbolic regression, and control-compatible synthesis; yet they also introduce specific numerical and structural issues, and in several subfields their expressive envelope is now known to differ sharply from that of ReLU-based models [2602.12390], [2310.13139].

Source: https://www.emergentmind.com/topics/rational-neural-networks