Papers
Topics
Authors
Recent
Search
2000 character limit reached

Hyperbolic Embedding Reservoir (HypER)

Updated 9 July 2026
  • Hyperbolic Embedding Reservoir (HypER) is an ESN architecture that embeds neurons in a unit Poincaré ball to naturally mirror chaotic stretch-and-fold dynamics.
  • It uses hyperbolic-uniform sampling and an exponential decay kernel based on hyperbolic distance to enforce curvature-aware connectivity.
  • The design retains standard ESN features while introducing geometric inductive bias, leading to improved performance on chaotic and real-world forecasting tasks.

Searching arXiv for the cited HypER paper and closely related hyperbolic-geometry work to ground the article. Hyperbolic Embedding Reservoir (HypER) is an Echo State Network (ESN) architecture in which the reservoir’s latent geometry and recurrent wiring are explicitly matched to the stretch-and-fold structure and exponential divergence of chaotic dynamics. In HypER, reservoir neurons are embedded in the unit Poincaré ball with curvature fixed at K=1K=-1, and recurrent couplings decay exponentially with hyperbolic geodesic distance. The design preserves standard ESN mechanisms—sparsity, leaky integration, spectral-radius normalization, and training restricted to a Tikhonov-regularized readout—while introducing a curvature-aware inductive bias intended to align local reservoir expansion and contraction with Lyapunov growth and the stable and unstable directions of chaotic flows (Singh et al., 25 Aug 2025).

1. Concept and motivation

HypER was introduced to address a specific mismatch between conventional reservoir computing and chaotic dynamics. Chaotic flows exhibit local stretch-and-fold behavior: infinitesimal perturbations separate exponentially along unstable directions and are subsequently folded back toward a fractal attractor. In the formulation used for HypER, if two initial conditions differ by δx0\|\delta x_0\|, then their separation grows as δxtδx0eλmaxt\|\delta x_t\| \approx \|\delta x_0\| e^{\lambda_{\max} t} when the largest Lyapunov exponent satisfies λmax>0\lambda_{\max} > 0 (Singh et al., 25 Aug 2025).

Classical ESNs typically operate in Euclidean latent spaces with polynomial volume and metric growth. According to the HypER formulation, such reservoirs can reproduce short-term structure but often struggle to sustain forecasts beyond approximately $5$–$8$ Lyapunov times because their latent geometry does not encode an inductive bias for exponential expansion and contraction (Singh et al., 25 Aug 2025). HypER addresses this by using negative curvature, so that exponential metric growth is built directly into the reservoir geometry.

The central intuition is geometric. Hyperbolic spaces have exponential expansion of volume and geodesic distance, so small coordinate changes near the boundary of the Poincaré ball can induce large geodesic increments. HypER exploits this by assigning many reservoir units to near-boundary regions under hyperbolic-uniform sampling, thereby representing strongly expanding directions, while more central units capture contracting modes. This suggests that the reservoir is not merely a random nonlinear dynamical system, but a geometry-shaped surrogate whose local gain structure is intended to mirror Lyapunov behavior (Singh et al., 25 Aug 2025).

2. Geometric foundation in the Poincaré ball

HypER uses the dd-dimensional unit Poincaré ball

Bd={xRd:x<1},B^d = \{ x \in \mathbb{R}^d : \|x\| < 1 \},

with constant negative sectional curvature fixed at K=1K=-1 (Singh et al., 25 Aug 2025). The conformal Riemannian metric is

ds2=4dx2(1x2)2.ds^2 = \frac{4 \|dx\|^2}{(1 - \|x\|^2)^2}.

The hyperbolic geodesic distance between δx0\|\delta x_0\|0 is defined as

δx0\|\delta x_0\|1

with

δx0\|\delta x_0\|2

Geodesics are straight lines through the origin or circular arcs orthogonal to the boundary, and angles are preserved because the metric is conformal (Singh et al., 25 Aug 2025).

In HypER, curvature is not itself tuned. The construction fixes δx0\|\delta x_0\|3 and uses the unit ball, while the effective expansion and contraction are controlled primarily by the hyperbolic kernel width δx0\|\delta x_0\|4 and the sparsity parameter δx0\|\delta x_0\|5 (Singh et al., 25 Aug 2025). This distinguishes HypER from formulations that treat curvature as a trainable parameter. A plausible implication is that the method emphasizes architectural simplicity and ESN-style fixed-reservoir design over manifold learning.

The use of hyperbolic geometry in HypER is conceptually consistent with broader machine-learning work exploiting exponential volume growth. Hyperbolic embeddings have been used to encode latent structure in networks through the δx0\|\delta x_0\|6 framework, where radial position represents popularity or degree and angular position represents similarity (García-Pérez et al., 2019). More generally, hyperbolic representation learning has been motivated by the ability of negatively curved spaces to embed tree-like and hierarchical structure with low distortion, although this comes with well-documented numerical constraints in both the Poincaré and Lorentz models (Mishne et al., 2022). HypER differs in using hyperbolic geometry not for static hierarchy modeling, but for recurrent dynamics and chaotic forecasting (Singh et al., 25 Aug 2025).

3. Reservoir construction and state dynamics

HypER defines each reservoir neuron δx0\|\delta x_0\|7 by a coordinate δx0\|\delta x_0\|8, written as δx0\|\delta x_0\|9, where δxtδx0eλmaxt\|\delta x_t\| \approx \|\delta x_0\| e^{\lambda_{\max} t}0 and δxtδx0eλmaxt\|\delta x_t\| \approx \|\delta x_0\| e^{\lambda_{\max} t}1 (Singh et al., 25 Aug 2025). Two sampling schemes are specified.

The Euclidean-isotropic baseline samples δxtδx0eλmaxt\|\delta x_t\| \approx \|\delta x_0\| e^{\lambda_{\max} t}2 uniformly in Euclidean volume within a ball of radius δxtδx0eλmaxt\|\delta x_t\| \approx \|\delta x_0\| e^{\lambda_{\max} t}3, using

δxtδx0eλmaxt\|\delta x_t\| \approx \|\delta x_0\| e^{\lambda_{\max} t}4

with δxtδx0eλmaxt\|\delta x_t\| \approx \|\delta x_0\| e^{\lambda_{\max} t}5 uniform on δxtδx0eλmaxt\|\delta x_t\| \approx \|\delta x_0\| e^{\lambda_{\max} t}6 (Singh et al., 25 Aug 2025). This baseline ignores the native hyperbolic volume measure.

HypER uses hyperbolic-uniform sampling. The hyperbolic radius δxtδx0eλmaxt\|\delta x_t\| \approx \|\delta x_0\| e^{\lambda_{\max} t}7 is drawn with density

δxtδx0eλmaxt\|\delta x_t\| \approx \|\delta x_0\| e^{\lambda_{\max} t}8

and related to the Euclidean radius by

δxtδx0eλmaxt\|\delta x_t\| \approx \|\delta x_0\| e^{\lambda_{\max} t}9

Using the normalized cumulative distribution

λmax>0\lambda_{\max} > 00

with λmax>0\lambda_{\max} > 01, the radius is sampled as

λmax>0\lambda_{\max} > 02

followed by λmax>0\lambda_{\max} > 03 and uniform λmax>0\lambda_{\max} > 04 on λmax>0\lambda_{\max} > 05 (Singh et al., 25 Aug 2025). This sampling places more nodes near the boundary, matching the exponential hyperbolic volume element.

Pairwise hyperbolic distances λmax>0\lambda_{\max} > 06 determine the recurrent weight matrix through the kernel

λmax>0\lambda_{\max} > 07

where λmax>0\lambda_{\max} > 08 is the kernel width (Singh et al., 25 Aug 2025). Small hyperbolic distances yield stronger coupling, whereas near-boundary nodes, which are far apart in hyperbolic distance, have weaker direct coupling.

Sparsity is enforced by retaining only the top-λmax>0\lambda_{\max} > 09 entries in each row, producing a sparse matrix $5$0. This typically yields a directed, asymmetric graph; symmetric variants are possible but are not used in the reported work (Singh et al., 25 Aug 2025). The sparse matrix is then scaled to a target spectral radius $5$1:

$5$2

where $5$3 denotes the spectral radius. The purpose is to preserve the geometry-induced structure while keeping the internal dynamics near the edge of stability and maintaining the echo-state property (ESP) (Singh et al., 25 Aug 2025).

The state update follows the standard leaky ESN form

$5$4

where $5$5 is the leak rate, $5$6 is a random input-weight matrix, and $5$7 is a $5$8-Lipschitz pointwise activation such as $5$9 (Singh et al., 25 Aug 2025). Mixed node-wise activations, specifically tanh–sine–linear combinations, are also explored and are reported to help most when combined with hyperbolic connectivity.

The readout uses a second-order polynomial feature map

$8$0

with prediction

$8$1

Only $8$2 is trained, preserving the characteristic fixed-reservoir workflow of ESNs (Singh et al., 25 Aug 2025).

4. Training protocol and theoretical analysis

HypER trains only the readout by Tikhonov-regularized least squares. After washout $8$3, the method collects feature-target pairs $8$4 for $8$5 and solves

$8$6

with closed-form solution

$8$7

where $8$8 stacks row targets, $8$9 stacks column features, and dd0 is the identity matrix (Singh et al., 25 Aug 2025). Cholesky or SVD is used for numerical stability.

Two operating modes are distinguished. During training and validation, teacher forcing (open loop) feeds the true input dd1 at every step, ensuring stability and single-step error learning. During testing, autoregressive forecasting (closed loop) feeds dd2 back as input, so long-horizon fidelity depends on whether the reservoir dynamics suppress or amplify relevant perturbations in a task-appropriate manner (Singh et al., 25 Aug 2025).

The paper develops a formal analysis of curvature-controlled divergence. For the row-sparse hyperbolic-kernel matrix dd3 with distinct points dd4 and minimum pairwise hyperbolic distance

dd5

Lemma 1 gives

dd6

As dd7, the ratio dd8 increases, which the paper interprets as lifting the spectral floor (Singh et al., 25 Aug 2025).

For a dd9 activation Bd={xRd:x<1},B^d = \{ x \in \mathbb{R}^d : \|x\| < 1 \},0 with Bd={xRd:x<1},B^d = \{ x \in \mathbb{R}^d : \|x\| < 1 \},1 on the reachable domain, the one-step Jacobian is

Bd={xRd:x<1},B^d = \{ x \in \mathbb{R}^d : \|x\| < 1 \},2

where

Bd={xRd:x<1},B^d = \{ x \in \mathbb{R}^d : \|x\| < 1 \},3

Lemma 2 provides

Bd={xRd:x<1},B^d = \{ x \in \mathbb{R}^d : \|x\| < 1 \},4

and

Bd={xRd:x<1},B^d = \{ x \in \mathbb{R}^d : \|x\| < 1 \},5

Since spectral normalization fixes Bd={xRd:x<1},B^d = \{ x \in \mathbb{R}^d : \|x\| < 1 \},6, the upper gain remains controlled, while the lower gain depends on the hyperbolic-kernel structure (Singh et al., 25 Aug 2025).

The main state-divergence result defines

Bd={xRd:x<1},B^d = \{ x \in \mathbb{R}^d : \|x\| < 1 \},7

If two input streams are identical up to time Bd={xRd:x<1},B^d = \{ x \in \mathbb{R}^d : \|x\| < 1 \},8, differ at Bd={xRd:x<1},B^d = \{ x \in \mathbb{R}^d : \|x\| < 1 \},9, and coincide thereafter, then for the resulting state difference K=1K=-10, if K=1K=-11,

K=1K=-12

with K=1K=-13 (Singh et al., 25 Aug 2025). The paper interprets this as a lower bound on reservoir state divergence that mirrors Lyapunov growth while preserving ESP via spectral normalization.

This analysis establishes the conceptual distinction from standard ESNs and graph-structured reservoirs. Standard ESNs and engineered topologies such as small-world or cycle reservoirs alter connectivity patterns in flat geometry, whereas HypER uses negative curvature and hyperbolic-distance decay to shape the local expansion–contraction spectrum itself (Singh et al., 25 Aug 2025).

5. Empirical evaluation

HypER is evaluated on synthetic chaotic systems and real-world time series. The synthetic benchmarks are Lorenz–63 with parameters K=1K=-14 and largest Lyapunov exponent approximately K=1K=-15, the Rössler system with K=1K=-16 and largest Lyapunov exponent approximately K=1K=-17, and the hyperchaotic Chen–Ueta system with K=1K=-18 (Singh et al., 25 Aug 2025). Trajectories are generated with LSODA at K=1K=-19 for ds2=4dx2(1x2)2.ds^2 = \frac{4 \|dx\|^2}{(1 - \|x\|^2)^2}.0 samples; the first ds2=4dx2(1x2)2.ds^2 = \frac{4 \|dx\|^2}{(1 - \|x\|^2)^2}.1 are discarded as washout, the next ds2=4dx2(1x2)2.ds^2 = \frac{4 \|dx\|^2}{(1 - \|x\|^2)^2}.2 are used for training, and the final ds2=4dx2(1x2)2.ds^2 = \frac{4 \|dx\|^2}{(1 - \|x\|^2)^2}.3 for testing.

The real-world datasets are MIT–BIH Arrhythmia (ECG), Santa Fe Dataset B (chaotic laser), and the Sunspot monthly index (SILSO v2.0). Signals are normalized to ds2=4dx2(1x2)2.ds^2 = \frac{4 \|dx\|^2}{(1 - \|x\|^2)^2}.4 and represented through a ds2=4dx2(1x2)2.ds^2 = \frac{4 \|dx\|^2}{(1 - \|x\|^2)^2}.5D delay embedding using standard ds2=4dx2(1x2)2.ds^2 = \frac{4 \|dx\|^2}{(1 - \|x\|^2)^2}.6 splits (Singh et al., 25 Aug 2025).

Baselines include Euclidean ESN, SCR, CRJ, SW-ESN, MCI-ESN, and DeepESN. Reservoir sizes are ds2=4dx2(1x2)2.ds^2 = \frac{4 \|dx\|^2}{(1 - \|x\|^2)^2}.7 units except for DeepESN, which uses ds2=4dx2(1x2)2.ds^2 = \frac{4 \|dx\|^2}{(1 - \|x\|^2)^2}.8. Global hyperparameters are tuned with ds2=4dx2(1x2)2.ds^2 = \frac{4 \|dx\|^2}{(1 - \|x\|^2)^2}.9, target δx0\|\delta x_0\|00, leak δx0\|\delta x_0\|01, and ridge regularization δx0\|\delta x_0\|02. Results are averaged over δx0\|\delta x_0\|03 runs (Singh et al., 25 Aug 2025).

The reported metrics are normalized root mean squared error (NRMSE), Valid Prediction Time (VPT), Attractor Deviation (ADev), and Welch power spectral density (PSD) (Singh et al., 25 Aug 2025).

Benchmark Result reported for HypER Comparison reported
Lorenz VPT δx0\|\delta x_0\|04 best baseline δx0\|\delta x_0\|05
Rössler VPT δx0\|\delta x_0\|06 best baseline δx0\|\delta x_0\|07
Chen–Ueta VPT δx0\|\delta x_0\|08 best baseline δx0\|\delta x_0\|09
Lorenz ADev δx0\|\delta x_0\|10 best baseline δx0\|\delta x_0\|11
Rössler ADev δx0\|\delta x_0\|12 best baseline δx0\|\delta x_0\|13

On Lorenz–63, HypER attains NRMSE δx0\|\delta x_0\|14 at δx0\|\delta x_0\|15 steps and δx0\|\delta x_0\|16 at δx0\|\delta x_0\|17 steps, outperforming MCI-ESN (δx0\|\delta x_0\|18), SW-ESN (δx0\|\delta x_0\|19), CRJ (δx0\|\delta x_0\|20), Euclidean ESN (δx0\|\delta x_0\|21), and DeepESN (δx0\|\delta x_0\|22) at the δx0\|\delta x_0\|23-step horizon (Singh et al., 25 Aug 2025).

On Rössler, HypER maintains an order-of-magnitude lower error than Euclidean baselines at every horizon up to δx0\|\delta x_0\|24 steps; the paper gives the example of δx0\|\delta x_0\|25 for HypER versus δx0\|\delta x_0\|26 for MCI-ESN at δx0\|\delta x_0\|27 steps (Singh et al., 25 Aug 2025). On Chen–Ueta, HypER reports δx0\|\delta x_0\|28 at δx0\|\delta x_0\|29 steps, lower than SW-ESN at δx0\|\delta x_0\|30 and lower than the other listed baselines (Singh et al., 25 Aug 2025).

PSD analysis is used to assess frequency-domain fidelity. The paper reports that HypER preserves the Lorenz power spectrum, whereas flat ESNs inject spurious high-frequency noise (Singh et al., 25 Aug 2025). This suggests that the benefit is not limited to pointwise trajectory error, but extends to attractor-level dynamical statistics.

On real-world datasets, the gains are strongest for signals treated as chaotic or nonstationary. On MIT–BIH, HypER reports NRMSE δx0\|\delta x_0\|31 at δx0\|\delta x_0\|32 steps compared with the best baseline MCI-ESN at δx0\|\delta x_0\|33, corresponding to an improvement of approximately δx0\|\delta x_0\|34–δx0\|\delta x_0\|35 across horizons (Singh et al., 25 Aug 2025). On Santa Fe, HypER gives δx0\|\delta x_0\|36 at δx0\|\delta x_0\|37 steps versus the best baseline δx0\|\delta x_0\|38. On sunspots, HypER is competitive but not superior, with δx0\|\delta x_0\|39 at δx0\|\delta x_0\|40 steps versus MCI-ESN at δx0\|\delta x_0\|41 (Singh et al., 25 Aug 2025). The paper interprets this as consistent with theory: when δx0\|\delta x_0\|42, raising δx0\|\delta x_0\|43 is unnecessary.

6. Ablations, implementation, and relation to adjacent work

The ablation results identify geometry, sampling, kernel scale, and latent dimension as the main determinants of performance. For Lorenz at the δx0\|\delta x_0\|44-step horizon, moving from Euclidean to hyperbolic geometry with Euclidean-volume sampling reduces NRMSE by approximately δx0\|\delta x_0\|45, and switching to hyperbolic-uniform sampling yields an additional approximately δx0\|\delta x_0\|46 reduction, together giving the best VPT and ADev (Singh et al., 25 Aug 2025). Increasing the latent dimension from the common Poincaré disk setting δx0\|\delta x_0\|47 to δx0\|\delta x_0\|48 or δx0\|\delta x_0\|49 weakens gains, which the paper attributes to curvature dilution.

The kernel width δx0\|\delta x_0\|50 has a clear optimum around δx0\|\delta x_0\|51. If δx0\|\delta x_0\|52 is too small, the graph becomes under-connected and VPT decreases; if δx0\|\delta x_0\|53 is too large, curvature effects are diluted and δx0\|\delta x_0\|54 (Singh et al., 25 Aug 2025). Performance plateaus for sparsity δx0\|\delta x_0\|55, with denser graphs slightly improving ADev, but NRMSE gains attributed mainly to geometry rather than overall connectivity scale.

Mixed activations improve all reservoir types, but the improvement is strongest for HypER. For Lorenz, the reported mixed-activation comparison gives HypER NRMSE δx0\|\delta x_0\|56 versus the best graph baseline at δx0\|\delta x_0\|57, and VPT δx0\|\delta x_0\|58 versus δx0\|\delta x_0\|59 (Singh et al., 25 Aug 2025).

A typical implementation uses δx0\|\delta x_0\|60 units, δx0\|\delta x_0\|61, δx0\|\delta x_0\|62, δx0\|\delta x_0\|63, δx0\|\delta x_0\|64, δx0\|\delta x_0\|65, δx0\|\delta x_0\|66, and δx0\|\delta x_0\|67 (Singh et al., 25 Aug 2025). Dense adjacency construction is δx0\|\delta x_0\|68; row-wise sparsification reduces subsequent structure to δx0\|\delta x_0\|69. Readout training is dominated by solving linear systems of size δx0\|\delta x_0\|70 (Singh et al., 25 Aug 2025). No manifold optimization is used because the reservoir is fixed after sampling and normalization.

HypER’s use of the Poincaré ball also places it near a broader numerical discussion in hyperbolic ML. Work on the numerical stability of hyperbolic representation learning has shown that hyperbolic models can suffer catastrophic NaNs and floating-point pathologies near the Poincaré boundary or at large Lorentz time coordinates, with different trade-offs between representational capacity and optimization behavior (Mishne et al., 2022). HypER avoids many of these issues because it does not train manifold coordinates directly; neuron positions are sampled once, and only the Euclidean readout is optimized (Singh et al., 25 Aug 2025). This suggests that fixed-geometry reservoir computing may be an attractive regime for exploiting hyperbolic structure without incurring the full optimization burden of end-to-end hyperbolic learning.

The method is also distinct from static hyperbolic embedding frameworks such as Mercator, which infer latent δx0\|\delta x_0\|71 coordinates of network nodes under the Popularityδx0\|\delta x_0\|72Similarity model (García-Pérez et al., 2019), and from hyperbolic dense retrieval systems such as HypRAG, where hyperbolic geometry is used to encode semantic hierarchy and preserve specificity through norm- or radius-based separation (Madhu et al., 8 Feb 2026). The shared principle is the use of exponential metric growth as inductive bias, but HypER applies it to dynamical forecasting rather than network embedding or retrieval.

7. Limitations and prospective directions

HypER fixes curvature at δx0\|\delta x_0\|73 and uses hyperbolic-uniform sampling with a maximum radius parameter δx0\|\delta x_0\|74. Although effective in the reported experiments, performance still depends on δx0\|\delta x_0\|75, δx0\|\delta x_0\|76, δx0\|\delta x_0\|77, δx0\|\delta x_0\|78, δx0\|\delta x_0\|79, and input scaling (Singh et al., 25 Aug 2025). Careful tuning remains necessary, so the geometric prior does not eliminate hyperparameter sensitivity.

The method’s computational construction is naively δx0\|\delta x_0\|80 in the number of reservoir units, which limits straightforward scaling to very large reservoirs. The paper explicitly identifies scalable approximations such as hyperbolic δx0\|\delta x_0\|81-nearest neighbors and fast kernel evaluation as useful directions (Singh et al., 25 Aug 2025). It also notes that gains diminish as latent dimension increases, leaving open the question of how to balance representational capacity against curvature concentration.

Several extensions are proposed: adaptive or learned node placement, task-specific curvature or variable metrics, hybrid knowledge-based plus reservoir schemes, online adaptation, output feedback, and multi-layer hyperbolic reservoirs (Singh et al., 25 Aug 2025). A systematic relation between Lyapunov exponents and geometric control parameters such as curvature and δx0\|\delta x_0\|82 is also identified as an open problem. This suggests a possible future program in which reservoir geometry is calibrated directly to the Lyapunov spectrum of the target system rather than chosen through heuristic tuning.

A common misconception would be to view HypER as simply an ESN with a different graph topology. The paper argues for a stronger claim: the essential modification is not merely graph sparsity or nonrandom wiring, but the introduction of a negative-curvature metric into the latent state space, so that the reservoir’s local expansion–contraction properties can track the geometry of chaotic divergence (Singh et al., 25 Aug 2025). Whether this principle generalizes beyond the studied chaotic and real-world benchmarks remains an empirical question, but within the reported scope, HypER presents a geometrically explicit formulation of reservoir computing in which hyperbolic structure functions as a dynamical inductive bias rather than a static embedding device.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Hyperbolic Embedding Reservoir (HypER).