- The paper presents a physics-informed reinforcement learning framework integrating a PINN-based inverse dynamics estimator for precise, sensorless GRF estimation.
- It achieves significant noise and impact reduction through an impact-aware reward shaping strategy that adapts to diverse footwear and surface conditions.
- Experimental results show a 6–7x reduction in RMSE and an R² improvement from 0.39 to 0.99, demonstrating a strong sim2real transfer for humanoid locomotion.
Introduction
The deployment of humanoid robots in human-centric environments imposes constraints far beyond basic locomotion, particularly with respect to acoustic noise and hardware longevity during interaction with the ground. The ground reaction force (GRF) generated during walking is a primary etiological factor for both structural vibration and impulsive acoustic events. Robust, low-noise locomotion remains a challenge, especially when considering contact-dynamics variation introduced by footwear diversity. Existing approaches primarily rely on either kinematic proxies or noisy force sensor feedback, both of which exhibit limitations regarding generalization and long-term system durability. "QuietWalk" (2604.23702) proposes a tightly integrated, physics-informed reinforcement learning (RL) framework aimed at mitigating these deficiencies, introducing a physically-constrained neural estimator for GRFs and an RL loop that leverages sensorless, GRF-aware reward shaping to achieve acoustically and mechanically robust bipedal locomotion under diverse footwear and surface conditions.

Figure 1: Humanoid robot validation in simulation and on hardware under barefoot, skate shoes, athletic sneakers, and high heels.
Accurate and physically-consistent GRF estimation is essential for closed-loop locomotion control with impact- and noise-minimizing objectives. The proposed approach utilizes a physics-informed neural network (PINN), architected to embed the inverse-dynamics structure of the robot, leveraging proprioceptive signals ([q,q˙​,q¨​]) spanning six timesteps as input. The network integrates a Kolmogorov–Arnold Network (KAN) layer for nonlinear feature parameterization and explicitly predicts dynamic model constituents: inertia, Coriolis/centrifugal, and potential terms, yielding a residual dynamics equation for GRF estimation. This strictly constrains the prediction manifold to physically plausible behaviors and mitigates overfitting to dataset statistics.

Figure 2: Structure of PINN-Based robot inverse dynamics.
Ablative studies reveal that omitting the inverse-dynamics residual in the learning objective leads to a 6–7x increase in vertical GRF root mean squared error (RMSE) and a severe collapse in R2; the inclusion of dynamics-consistency constraints improves R2 from 0.39 to 0.99 on held-out hardware data, dramatically narrowing the sim2real gap and providing a robust alternative to direct force-sensor feedback. The smoothness and swing-phase masking regularization further enhance temporal stability and classification between stance and swing.

Figure 3: Ablation study of the GRF predictor under different loss configurations (C1–C5).
Transferring locomotion controllers across varying footwear, including high compliance (barefoot, sneakers) and extreme stiffness (high heels), demands explicit modeling of foot-end geometry and inertial parameters. Following a unified pipeline, multiple footwear assets are constructed and merged via mesh union, ensuring precise leg-length compensation and inertia recalibration. This design maximally preserves the repeatability and integrity of style- and contact-induced distributional shifts across both simulation domains and hardware deployments.

Figure 4: Overview of the proposed physics-informed reinforcement learning framework with cross-footwear deployment.
The RL loop is augmented with a deployable, frozen GRF predictor as a reward critic, directly penalizing predicted per-foot normal impact forces squared rather than using indirect kinematic proxies. The impact-aware reward is staged via a curriculum, ramping the impact-penalty coefficient in concert with increased surface and footwear randomization. The policy operates at 50 Hz with six-step proprioceptive history, outputting joint increments tracked by default low-level PD control. The GRF estimator is never updated in RL, ensuring full sim2real reward consistency. This hybrid reward structure permits simultaneous minimization of acoustic and mechanical impact while retaining maneuverability and stability at speeds up to 1.2 m/s.
Experimental Results
GRF Prediction Accuracy
Quantitative benchmarks on held-out robot datasets highlight the impact of PINN regularization: the proposed loss yields left/right RMSE of 14.5/14.0 N and R2 of 0.9887/0.9899, compared with 106.3/79.4 N RMSE and R2 of 0.39/0.67 for a purely supervised comparator. Ablations confirm that every component—dynamics residual, swing-phase masking, and smoothness—exerts a significant effect, but inverse-dynamics regularization is critical for reliable, physically consistent estimation.
Extensive real-robot trials measure mean (MNL) and peak (PNL) A-weighted sound pressure levels across multiple surfaces and footwear types. Against a baseline RL policy, the QuietWalk controller achieves an average MNL reduction of 7.17 dB and PNL reduction of 4.98 dB under barefoot walking across surfaces, without notable degradation in gait stability. Surface properties remain the dominant extrinsic factor, but footwear contact compliance and area significantly modulate both overall and peak noise.

Figure 5: Mean noise level (MNL) and peak noise level (PNL) of D1, D2, and D3 under the barefoot condition across four surface conditions.
Comparative results demonstrate that compliant footwear (sneakers, skate shoes) yields lower SPL than high-heeled shoes. Wood surfaces, with low damping, exacerbate noise—up to 20 dB higher in peak-to-peak comparisons between the most and least favorable footwear–surface pairings. These findings highlight the mechanism by which contact stiffness and damping at the shoe–surface interface governs impulsive energy transfer and thus acoustic emissions.

Figure 6: MNL and PNL of D2 across four footwear conditions and four surface conditions.
The policy's generalization capability is empirically validated under diverse outdoor terrains and footwear. Stable gaits and low failure rates are preserved under morphologically challenging configurations, including high heels on uneven and high-noise surfaces (gravel, cobblestone), underscoring effective adaptation in the face of significant distributional shift in contact dynamics.

Figure 7: Outdoor robustness evaluation under diverse footwear conditions.
Implications and Prospects
Strong claims supported by the numerical results: the integration of dynamics-consistent PINN-based GRF estimation enables high-fidelity, sensorless contact force awareness, and its embedding into RL reward structuring facilitates direct minimization of impact-induced noise across both simulated and hardware deployments. The ability to generalize to highly variable footwear and unpredictable surface conditions is demonstrated quantitatively, representing a substantial advance in both practical deployability and theoretical understanding of morphology–control interdependencies.
Practically, this approach eliminates the requirement for brittle and expensive force transducers, lowers acoustic footprints critical in human-centric deployments, and reduces mechanical wear via softer contacts. Theoretically, embedding explicit physics constraints within neural estimators and reward signals provides a promising direction for sim2real transfer in locomotion, especially where system morphology is variable or environmental characterization is incomplete.
Limitations: The approach is presently restricted to normal (vertical) GRF estimation; tangential and full 3D force estimation remains for future work. Acoustic evaluation is relative and tied to the specific recording setup, with further calibration needed for absolute comparability. Future extensions should include complete 3D contact modeling, improved acoustic sensing, and extended validation across heterogeneous robot platforms.
Conclusion
QuietWalk delivers a sensorless, physics-informed RL pipeline capable of minimizing GRF-induced acoustic and mechanical impact for humanoid robots under diverse and variable footwear conditions. The fusion of inverse-dynamics-constrained PINN estimation with impact-aware RL reward shaping yields robust, generalizable policies, validated empirically across footwear and surfaces with significant reductions in measured noise and impact. This work establishes new experimental and algorithmic benchmarks for physically-consistent, deployable, and general-purpose indoor humanoid locomotion.