Papers
Topics
Authors
Recent
Search
2000 character limit reached

Heracles in Astronomy, ML, and Robotics

Updated 12 July 2026
  • Heracles is a chemically distinct, metal-poor inner-Galaxy proto-component identified via orbital dynamics and Plummer density modeling in Galactic archaeology.
  • Heracles denotes a hybrid SSM–Transformer model that integrates global state-space modules and local convolution, achieving competitive high-resolution image and time-series analysis.
  • Heracles also refers to a state-conditioned diffusion middleware in humanoid robotics that bridges precise tracking with generative recovery, improving robustness under large deviations.

Searching arXiv for relevant papers on “Heracles” to ground the article and disambiguate the multiple uses of the term. Found multiple distinct research usages of “Heracles,” including Milky Way archaeology, machine learning, and humanoid control. I’ll rely on the specific papers supplied in the data block and cite them precisely. Heracles is a name used for several unrelated research objects in contemporary literature. In Galactic archaeology, it denotes a chemically defined inner-Galaxy proto-Galactic component identified in APOGEE–Gaia analyses and modeled as a significant fragment of the proto-Milky Way (Horta et al., 2024). In machine learning, it denotes a hybrid SSM–Transformer architecture for high-resolution image and time-series analysis built from a Hartley-kernel global SSM, a localized convolutional SSM, and attention-based token interaction (Patro et al., 2024). In humanoid robotics, it denotes a state-conditioned diffusion middleware that operates between reference motions and low-level physics trackers, preserving identity-like tracking near nominal states and synthesizing recovery trajectories under large deviations (Tao et al., 29 Mar 2026). The shared name therefore functions as a disciplinary homonym rather than a single concept.

1. Disciplinary scope and nomenclature

In the Milky Way literature, “Heracles” appears as a stellar structure or proto-Galactic fragment inferred from chemistry, spatial concentration, and orbital properties. Horta & Schiavon identify populations in the inner Galaxy largely associated with the Heracles structure through a purely chemical dissection of APOGEE–Gaia stellar populations, and argue that the proto-Milky Way is at least comprised of two significant fragments: the main in situ progenitor and the Heracles structure (Horta et al., 2024).

In machine learning, “Heracles” names a model architecture introduced for vision and time-series analysis. It is explicitly described as “a Hybrid SSM–Transformer Model for High-Resolution Image and Time-Series Analysis,” with a three-component block comprising a global SSM, a local SSM, and an attention module (Patro et al., 2024).

In humanoid control, “Heracles” names a control framework subtitled “Bridging Precise Tracking and Generative Synthesis for General Humanoid Control.” Its core contribution is a state-conditioned diffusion middleware that adapts between precise tracking and generative recovery without explicit mode switching (Tao et al., 29 Mar 2026).

This multiplicity of meanings is important because the astronomical, machine-learning, and robotics usages are technically unrelated, despite sharing the same label.

2. Heracles in Galactic archaeology: identification and chemistry

Heracles was initially characterized in APOGEE–Gaia work as a metal-poor, low-energy, high-eccentricity substructure confined to the inner few kpc of the Milky Way halo. In APOGEE DR17 + Gaia EDR3, 300 Heracles candidate stars were selected solely by orbital parameters plus a chemical cut in the [Mg/Mn]–[Al/Fe] plane. The full criteria were e>0.6e>0.6, 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}, a two-segment selection in the [Al/Fe] vs. [Mg/Mn] plane, and [Fe/H]>1.7[\mathrm{Fe/H}]>-1.7; the resulting sample lay at RGC4R_{\rm GC}\lesssim 4 kpc (Horta et al., 2022).

A later APOGEE–Gaia analysis used a purely chemical dissection. In that framework, stars of the “unevolved” locus in the [Mg/Mn]–[Al/Fe] plane satisfy [Mg/Mn]>0.15[\mathrm{Mg}/\mathrm{Mn}]>0.15 or [Al/Fe]<0.2[\mathrm{Al}/\mathrm{Fe}]<-0.2, with boundary

[Mg/Mn]=2[Al/Fe]+0.6,[\mathrm{Mg}/\mathrm{Mn}] = 2\,[\mathrm{Al}/\mathrm{Fe}] + 0.6,

and Heracles is then isolated by additional cuts in the [Mg/Fe]–[Fe/H] plane, selecting the high-[Mg/Fe], moderate-[Fe/H] locus that is spatially confined to the inner Galaxy (Horta et al., 2024).

Chemically, Heracles is consistently described as metal poor and α\alpha-enhanced. One APOGEE-based characterization gives a metallicity distribution spanning roughly 1.7[Fe/H]1.0-1.7 \lesssim [\mathrm{Fe/H}] \lesssim -1.0 with a peak around [Fe/H]1.3[\mathrm{Fe/H}] \approx -1.3, median 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}0 dex, and dispersion 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}1 dex. For Mg, the sequence is approximately 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}2 dex over 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}3, with no statistically significant knee detected up to 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}4. Fe-peak and neutron-capture tracers were reported as 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}5 dex at 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}6 rising to 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}7 dex at 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}8, 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}9 flat at [Fe/H]>1.7[\mathrm{Fe/H}]>-1.70, and [Fe/H]>1.7[\mathrm{Fe/H}]>-1.71 at [Fe/H]>1.7[\mathrm{Fe/H}]>-1.72 (Horta et al., 2022).

The chemically dissected proto-Galaxy analysis gives a closely related but not identical description: Heracles stars span roughly [Fe/H]>1.7[\mathrm{Fe/H}]>-1.73, peak around [Fe/H]>1.7[\mathrm{Fe/H}]>-1.74, and are distinguished at fixed [Fe/H]>1.7[\mathrm{Fe/H}]>-1.75 by relatively high [Fe/H]>1.7[\mathrm{Fe/H}]>-1.76-enhancement, [Fe/H]>1.7[\mathrm{Fe/H}]>-1.77–[Fe/H]>1.7[\mathrm{Fe/H}]>-1.78, compared with lower-[Mg/Fe] debris such as Gaia-Enceladus/Sausage (Horta et al., 2024).

3. Spatial structure, density modeling, and mass of the Galactic Heracles

The most explicit density model for Heracles is a triaxial, oblate Plummer profile. For Heracles-dominated chemical cells, the three-dimensional density is written as

[Fe/H]>1.7[\mathrm{Fe/H}]>-1.79

or, in mass-normalized form,

RGC4R_{\rm GC}\lesssim 40

For the core Heracles cell, the best-fit parameters are RGC4R_{\rm GC}\lesssim 41, RGC4R_{\rm GC}\lesssim 42, and RGC4R_{\rm GC}\lesssim 43 (Horta et al., 2024).

These values are consistent with the broader proto-Galaxy fit, which finds the chemically defined populations to be well represented by a Plummer model with a scale radius of RGC4R_{\rm GC}\lesssim 44 kpc and an oblate ellipsoid with flattening parameters RGC4R_{\rm GC}\lesssim 45 and RGC4R_{\rm GC}\lesssim 46. The interpretation offered there is that the Milky Way plausibly hosts a low-mass, metal-poor, bulge component (Horta et al., 2024).

Mass estimates depend on the adopted chemical cell definition. Integrating the density of Cell 2 within RGC4R_{\rm GC}\lesssim 47 yields

RGC4R_{\rm GC}\lesssim 48

Summing over all three inner-Galaxy chemical cells gives a combined mass of RGC4R_{\rm GC}\lesssim 49, while modeling all three together gives [Mg/Mn]>0.15[\mathrm{Mg}/\mathrm{Mn}]>0.150 (Horta et al., 2024).

Morphologically, the best-fit [Mg/Mn]>0.15[\mathrm{Mg}/\mathrm{Mn}]>0.151 indicates a distinctly oblate ellipsoid, and [Mg/Mn]>0.15[\mathrm{Mg}/\mathrm{Mn}]>0.152 indicates only mild triaxiality. Heracles occupies the inner [Mg/Mn]>0.15[\mathrm{Mg}/\mathrm{Mn}]>0.153–[Mg/Mn]>0.15[\mathrm{Mg}/\mathrm{Mn}]>0.154 kpc apocenter halo, overlaps the classical bulge/bar region, and remains chemically distinct from the bar’s younger, higher-[Fe/H] populations. Its centrally concentrated Plummer core with [Mg/Mn]>0.15[\mathrm{Mg}/\mathrm{Mn}]>0.155 kpc was described as resembling a low-mass, metal-poor bulge component, a “poor old heart,” embedded within the boxy/peanut bar (Horta et al., 2024).

An earlier APOGEE-based interpretation, using orbital decay and chemical arguments, estimated a progenitor stellar mass [Mg/Mn]>0.15[\mathrm{Mg}/\mathrm{Mn}]>0.156, with accretion at high redshift [Mg/Mn]>0.15[\mathrm{Mg}/\mathrm{Mn}]>0.157, and interpreted Heracles as a building block whose debris now dominate the very inner halo (Horta et al., 2022). This suggests that published mass and origin estimates depend sensitively on the adopted selection and modeling strategy.

4. Competing interpretations: accreted fragment, proto-Galactic component, or Aurora analogue

The principal controversy surrounding Galactic Heracles concerns whether it is best understood as an ex situ merger remnant, a distinct proto-Galactic fragment, or the chemical twin of an in situ population called Aurora.

One APOGEE-based study states that Heracles differs chemically from in situ populations such as Aurora and its inner halo counterparts in a statistically significant way. Using a 13-element [Mg/Mn]>0.15[\mathrm{Mg}/\mathrm{Mn}]>0.158 comparison,

[Mg/Mn]>0.15[\mathrm{Mg}/\mathrm{Mn}]>0.159

with [Al/Fe]<0.2[\mathrm{Al}/\mathrm{Fe}]<-0.20 for the inner high-[Al/Fe]<0.2[\mathrm{Al}/\mathrm{Fe}]<-0.21 comparison and [Al/Fe]<0.2[\mathrm{Al}/\mathrm{Fe}]<-0.22 for Heracles versus Aurora, it reports at [Al/Fe]<0.2[\mathrm{Al}/\mathrm{Fe}]<-0.23: [Al/Fe]<0.2[\mathrm{Al}/\mathrm{Fe}]<-0.24 and

[Al/Fe]<0.2[\mathrm{Al}/\mathrm{Fe}]<-0.25

The same work argues that in every major element, especially O, Mg, and Si, Heracles is [Al/Fe]<0.2[\mathrm{Al}/\mathrm{Fe}]<-0.26–[Al/Fe]<0.2[\mathrm{Al}/\mathrm{Fe}]<-0.27 dex lower than the in situ high-[Al/Fe]<0.2[\mathrm{Al}/\mathrm{Fe}]<-0.28 sequence at the same [Al/Fe]<0.2[\mathrm{Al}/\mathrm{Fe}]<-0.29, and interprets this as evidence that the star formation rate was lower in Heracles than in the early Milky Way (Horta et al., 2022).

A different analysis reaches the opposite conclusion. In an unsupervised decomposition of the local stellar halo, Myeong et al. report that “Aurora is entirely consistent with the chemical properties of the so-called Heracles merger.” In APOGEE, Aurora has [Mg/Mn]=2[Al/Fe]+0.6,[\mathrm{Mg}/\mathrm{Mn}] = 2\,[\mathrm{Al}/\mathrm{Fe}] + 0.6,0, [Mg/Mn]=2[Al/Fe]+0.6,[\mathrm{Mg}/\mathrm{Mn}] = 2\,[\mathrm{Al}/\mathrm{Fe}] + 0.6,1, and [Mg/Mn]=2[Al/Fe]+0.6,[\mathrm{Mg}/\mathrm{Mn}] = 2\,[\mathrm{Al}/\mathrm{Fe}] + 0.6,2; in GALAH it has [Mg/Mn]=2[Al/Fe]+0.6,[\mathrm{Mg}/\mathrm{Mn}] = 2\,[\mathrm{Al}/\mathrm{Fe}] + 0.6,3, [Mg/Mn]=2[Al/Fe]+0.6,[\mathrm{Mg}/\mathrm{Mn}] = 2\,[\mathrm{Al}/\mathrm{Fe}] + 0.6,4, [Mg/Mn]=2[Al/Fe]+0.6,[\mathrm{Mg}/\mathrm{Mn}] = 2\,[\mathrm{Al}/\mathrm{Fe}] + 0.6,5, together with elevated Ba, Y, and Eu. That study proposes that the Heracles signature is “the fossil relic of the earliest, bursty phase of the Milky Way’s own disk growth (‘in situ’), rather than a disrupted satellite” (Myeong et al., 2022).

The later APOGEE–Gaia proto-Galaxy modeling of Horta & Schiavon again treats Heracles as one of at least two significant proto-Galactic fragments, alongside the main in situ progenitor. In that interpretation, Heracles is “comparably massive” to the main progenitor, formed more slowly, and contributes significantly to the metal-poor bulge/inner halo (Horta et al., 2024).

Taken together, the literature does not offer a single settled interpretation. The stable points of agreement are that Heracles is chemically old, metal poor, centrally concentrated, and associated with the ancient inner Galaxy; the disputed point is whether those properties are best attributed to an accreted building block, a separate proto-Galactic fragment, or the earliest in situ Milky Way starburst.

5. Heracles as a hybrid SSM–Transformer model for image and time-series analysis

In machine learning, Heracles is a model architecture designed to address limitations attributed both to vision transformers and to earlier vision-oriented state space models. The architecture is defined as a three-component block: a global SSM based on a real-valued Hartley kernel, a localized convolutional SSM for fine spatial detail, and an attention-based token interaction module in deeper layers. Starting from the continuous-time state-space model

[Mg/Mn]=2[Al/Fe]+0.6,[\mathrm{Mg}/\mathrm{Mn}] = 2\,[\mathrm{Al}/\mathrm{Fe}] + 0.6,6

the model uses zero-order-hold discretization,

[Mg/Mn]=2[Al/Fe]+0.6,[\mathrm{Mg}/\mathrm{Mn}] = 2\,[\mathrm{Al}/\mathrm{Fe}] + 0.6,7

and the convolutional kernel view

[Mg/Mn]=2[Al/Fe]+0.6,[\mathrm{Mg}/\mathrm{Mn}] = 2\,[\mathrm{Al}/\mathrm{Fe}] + 0.6,8

Its global operator replaces a complex FFT by the real Hartley transform [Mg/Mn]=2[Al/Fe]+0.6,[\mathrm{Mg}/\mathrm{Mn}] = 2\,[\mathrm{Al}/\mathrm{Fe}] + 0.6,9, defining

α\alpha0

where α\alpha1 is a learned Hartley-domain filter. The local stream is a discrete convolution

α\alpha2

and deeper layers append multi-head self-attention with

α\alpha3

The layer flow is: input tokens α\alpha4 of shape α\alpha5; split into global and local streams; sum streams, apply LayerNorm and a two-layer MLP; use only SSM+MLP in the first α\alpha6 blocks and append multi-head attention in the later α\alpha7 stages (Patro et al., 2024).

For ImageNet-1K, the training recipe is AdamW with α\alpha8, α\alpha9, 1.7[Fe/H]1.0-1.7 \lesssim [\mathrm{Fe/H}] \lesssim -1.00, a 10-epoch linear warm-up plus 310-epoch cosine decay, batch size 128 on 1.7[Fe/H]1.0-1.7 \lesssim [\mathrm{Fe/H}] \lesssim -1.01V100 GPUs, 1.7[Fe/H]1.0-1.7 \lesssim [\mathrm{Fe/H}] \lesssim -1.02 input resolution, and RandAug, CutMix, MixToken, and Token Labeling (Patro et al., 2024).

Variant Params / FLOPs ImageNet top-1
Heracles-C-Small 21.7 M / 4.1 G 84.5%
Heracles-C-Base 32.5 M / 6.5 G 85.2%
Heracles-C-Large 54.1 M / 13.4 G 85.9%
Heracles-C-Huge 156.7 M / 39.3 G 86.4%

The reported ImageNet comparisons state that Heracles-C-Small at 1.7[Fe/H]1.0-1.7 \lesssim [\mathrm{Fe/H}] \lesssim -1.03 G achieves 1.7[Fe/H]1.0-1.7 \lesssim [\mathrm{Fe/H}] \lesssim -1.04 versus BiFormer-S 1.7[Fe/H]1.0-1.7 \lesssim [\mathrm{Fe/H}] \lesssim -1.05 and iFormer-S 1.7[Fe/H]1.0-1.7 \lesssim [\mathrm{Fe/H}] \lesssim -1.06; Heracles-C-Base at 1.7[Fe/H]1.0-1.7 \lesssim [\mathrm{Fe/H}] \lesssim -1.07 G achieves 1.7[Fe/H]1.0-1.7 \lesssim [\mathrm{Fe/H}] \lesssim -1.08 versus Wave-ViT-B 1.7[Fe/H]1.0-1.7 \lesssim [\mathrm{Fe/H}] \lesssim -1.09 and MaxViT-S [Fe/H]1.3[\mathrm{Fe/H}] \approx -1.30; Heracles-C-Large at [Fe/H]1.3[\mathrm{Fe/H}] \approx -1.31 G achieves [Fe/H]1.3[\mathrm{Fe/H}] \approx -1.32 versus VOLO-D3 [Fe/H]1.3[\mathrm{Fe/H}] \approx -1.33 and Wave-ViT-L [Fe/H]1.3[\mathrm{Fe/H}] \approx -1.34; and Heracles-C-Huge at [Fe/H]1.3[\mathrm{Fe/H}] \approx -1.35 G achieves [Fe/H]1.3[\mathrm{Fe/H}] \approx -1.36 versus LiT-22B [Fe/H]1.3[\mathrm{Fe/H}] \approx -1.37. The paper states that Heracles-C-small achieves state-of-the-art performance on ImageNet with [Fe/H]1.3[\mathrm{Fe/H}] \approx -1.38 top-1 accuracy, and that Heracles-C-Large and Heracles-C-Huge further improve accuracy to [Fe/H]1.3[\mathrm{Fe/H}] \approx -1.39 and 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}00, respectively (Patro et al., 2024).

The same work reports transfer-learning results on CIFAR-10, CIFAR-100, Flowers-102, and Cars-196: 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}01, 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}02, 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}03, and 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}04, respectively, comparing favorably to DeiT-B. In MS-COCO instance segmentation with Mask R-CNN 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}05 schedule on val2017, Heracles-C-S reaches 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}06, 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}07, 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}08, 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}09, 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}10, and 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}11. In time-series forecasting on ETTm1/2, ETTh1/2, Electricity, and Weather, using MSE and MAE at horizons 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}12, it achieves best or second-best on almost all tables; examples given are ETTm1 at 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}13, MSE 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}14 versus best 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}15, and Electricity at 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}16, MSE 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}17 versus best 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}18 (Patro et al., 2024).

Ablation results isolate the architectural contribution of the parallel design. Global SSM only gives 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}19, local SSM only gives 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}20, a series Hartley2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}21Conv arrangement gives 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}22, and the parallel Hartley2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}23Conv arrangement gives 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}24, the best result in Table A1. The paper attributes this to joint global/local streams injecting both translation equivariance and long-range inductive biases absent in pure attention. It also reports that early spectral blocks reduce token dimension before quadratic attention, and gives an A100 latency of 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}25 ms for Heracles-C-S versus 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}26 ms for GFNet-H-S at 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}27 (Patro et al., 2024).

6. Heracles as state-conditioned diffusion middleware for humanoid control

In humanoid robotics, Heracles is a hierarchical controller built from three modules: high-level reference motions, a state-conditioned diffusion middleware, and a low-level physics tracker. The reference signal is a time-indexed trajectory 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}28. The middleware is a lower-frequency 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}29 Hz planner 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}30 that observes the robot’s physical state 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}31 and the commanded reference 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}32, and synthesizes a short-horizon keyframe trajectory 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}33. The low-level policy 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}34 runs at 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}35 Hz on a densified version 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}36 of 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}37 and the proprioceptive state, issuing joint-position commands executed via PD control at 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}38 Hz. The receding-horizon loop is: observe 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}39; generate 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}40; densify 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}41; execute 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}42 for the next 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}43 steps (Tao et al., 29 Mar 2026).

The middleware uses a continuous flow-matching diffusion model in residual space. With static baseline

2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}44

the residual is

2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}45

If 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}46, then 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}47, preserving identity-like behavior. Flow matching is defined with 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}48 as normalized ground-truth residual, 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}49, interpolation path

2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}50

conditioning vector

2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}51

and velocity-matching loss

2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}52

Inference begins from

2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}53

with directional warm-start residual

2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}54

and integrates

2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}55

using a small number of Euler steps, for example 5 steps from 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}56 (Tao et al., 29 Mar 2026).

The small-deviation regime is treated explicitly as an identity-map regime. Because the model is parameterized around the current state, and because the first token is “pinned” via inpainting,

2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}57

the learned field satisfies 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}58 for 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}59, causing

2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}60

For large deviations, the same system transitions into generative synthesis: as 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}61 grows, the residual departs from zero, the directional warm start seeds a coarse straight-line plan, and the learned field edits it into a human-like recovery in very few steps, akin to SD-Edit. Because the pipeline runs in closed loop at 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}62 Hz, it can continuously replan and blend back toward the original reference once stability is regained (Tao et al., 29 Mar 2026).

Quantitatively, on 101 unseen motion sequences spanning locomotion, dance, martial arts, daily activities, acrobatics, and discontinuous clips, the reported Completion Rate is 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}63 for Heracles versus 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}64 for the best non-generative tracker and 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}65 for an MLP baseline; root orientation error is 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}66 rad, a 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}67 change versus tracker alone, and root height error is 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}68 m, a 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}69 change. On the fall-and-recovery subset, Completion Rate is 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}70 versus 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}71 for the MLP and 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}72 for VQ-VAE; joint-position error is 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}73 rad versus 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}74 rad for the MLP, and orientation error is 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}75 rad versus 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}76 rad. With a fixed iFSQ tracker, removing directional warm start reduces Completion Rate by 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}77, removing noisy-state augmentation reduces it by 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}78, and removing kinematics-aware loss weights reduces it by 2.6<E<2.0×105km2s2-2.6 < E < -2.0 \times 10^5\,\mathrm{km}^2\,\mathrm{s}^{-2}79. Qualitatively, under severe pushes and out-of-distribution martial-arts sequences, Heracles is reported to remain anthropomorphic, exhibiting compensatory stepping, arm countermotions, and torso realignment, while pure trackers collapse or execute jerky corrections; on Unitree G1, it executes omnidirectional lie-to-stand recoveries from supine, lateral, and prone configurations without task-specific engineering (Tao et al., 29 Mar 2026).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Heracles.