Papers
Topics
Authors
Recent
Search
2000 character limit reached

Learning Locomotion on Discrete Terrain via Minimal Proximity Sensing

Published 30 Jun 2026 in cs.RO | (2606.31912v2)

Abstract: Learning-based control has revolutionized dynamic locomotion, yet navigating unstructured terrain remains limited by a robot's incomplete awareness of imminent ground contact. While global perception systems such as LiDARs and depth cameras provide environmental context, they are frequently plagued by latencies, occlusions, and the high computational cost of dense geometric reconstruction. On the other hand, proprioceptive feedback is purely reactive, initiating corrections only after impact has occurred. This work explores embedding a minimal suite of low-cost, high-frequency infrared proximity sensors directly into the feet of a quadrupedal robot. These sensors provide "pre-contact" feedback that is robust to self-occlusions and significantly less computationally demanding than conventional vision-based pipelines. By integrating these localized signals into a reinforcement learning framework, we enable the robot to anticipate terrain discontinuities such as gaps and stepping stones that are problematic for traditional perception stacks due to occlusions or state estimation drift. We demonstrate that such sparse, near-field sensing can be reliably modeled in simulation and transferred to the real world with high fidelity. Experimental results show that local proximity sensing substantially improves traversal robustness over discrete terrain and offers a low-power, low-latency alternative or complement to complex global perception suites in unpredictable environments. For more information about results and methods, please see the project website: https://sites.google.com/view/foot-tof/home.

Summary

  • The paper presents a novel strategy using foot-mounted ToF sensors for immediate underfoot feedback, enabling anticipatory control on discrete terrains.
  • It employs a two-stage curriculum with a recurrent (LSTM) policy optimized via PPO, effectively handling diverse challenges like gaps and uneven footholds.
  • Empirical evaluations show that the approach outperforms global perception methods in noise resilience and robustness, achieving up to 60 cm gap traversal.

Learning Locomotion on Discrete Terrain via Minimal Proximity Sensing

Introduction and Motivation

The paper "Learning Locomotion on Discrete Terrain via Minimal Proximity Sensing" (2606.31912) addresses a significant bottleneck in the field of legged robot locomotion: the challenge of perceiving and negotiating unstructured terrains characterized by discontinuities and sparse footholds. Traditional reliance on global exteroceptive sensors (LiDAR, depth cameras) is hampered by occlusions, perceptual aliasing, and substantial computational cost, while proprioceptive signals only allow reactive correction after loss of ground contact. The authors propose using sparse, low-power, foot-mounted time-of-flight (ToF) proximity sensors to provide immediate, underfoot geometrical feedback, enabling anticipatory control policies robust to noise, occlusion, and environmental uncertainty.

Figure 1

Figure 1: Local proximity sensing integrated on quadruped feet, yielding policies with superior noise-resilience and locomotion reliability over discrete terrain.

Sensor System and Characterization

The hardware solution integrates the STMicroelectronics VL53L5CX ToF infrared sensor into the feet of an ANYmal-D quadruped. This sensor provides a 4×44\times 4 grid of range measurements with a $\ang{45}\times\ang{45}$ FoV, achieving 1 mm1\,\mathrm{mm} resolution at 60 Hz60\,\mathrm{Hz} with measured latencies of 20−60 ms20-60\,\mathrm{ms}. The grid-based FoV ensures coverage of the projected foothold area during swing.

Figure 2

Figure 2

Figure 2: Foot-mounted proximity sensor with 4×44\times 4 grid providing dense underfoot coverage.

Comprehensive characterization established that sensor accuracy is optimal at short range and perpendicular incidence, with noise and outlier rates increasing for longer ranges and grazing angles.

Figure 3

Figure 3

Figure 3: Heatmaps quantifying ToF noise and dropped readings as functions of distance and incidence angle.

Sensor noise profiles and failure modes derived from empirical data were modeled in simulation to enable high-fidelity policy transfer under real-world conditions.

Learning-Based Locomotion Framework

The control architecture leverages a recurrent policy (LSTM) optimized via PPO, using Isaac Gym for scalable simulation. Observations consist of proximity sensor readings augmented with proprioceptive signals (joint states, body dynamics). The learning pipeline omits reliance on elevation maps, map-fusion, or expensive attention over 3D point clouds, directly mapping sparse sensor feedback to agile locomotion actions.

Training is structured in two phases:

  • Stage 1: Curriculum traverses 5 canonical terrains (dense/row/mixed stones, rough fields, gaps), with task difficulty escalated via foothold geometry, gap sizes, and height variations. This phase establishes core sensor-motor associations.
  • Stage 2: Encompasses more challenging variants (smaller footholds, larger gaps, steeper elevation changes) and introduces beam walking, stairs, and pyramid slopes for generalization.

Figure 4

Figure 4

Figure 4: Sequence of discrete terrain morphologies employed during two-stage curriculum learning.

Randomization of sensor degradations (offsets, Gaussian/proportional noise, drops, delays) is incorporated to prevent overfitting and to bridge the sim-to-real gap.

Empirical Robustness Analysis

Rigorous robustness evaluations compare foot-centric proximity sensing to two base-centric benchmarks:

  • Mock (Ideal Height Scan): Assumes noise-free, egocentric top-down terrain scans.
  • Mock (Aggre. Elevation Map): Fuses multiple chassis-mounted depth images using odometry for registration, simulating real-world elevation map aggregation.

Figure 5

Figure 5

Figure 5: Visualization of diverse observation configurations: foot-centric rays, egocentric height scan, and odometry-fused elevation map.

Performance was assessed by success rate (1000 trials) for terrain traversal under injected degradations:

  • Static joint position offsets: Base-centric approaches rapidly lose performance as kinematic errors misalign foot pose estimation, corrupting mapping from global height field to local contact.
  • Robot morphology mismatches (e.g., shank length changes): Sensitivity of base-centric mapping to inaccurate kinematic parameters is pronounced.
  • Odometry drift: Aggregated elevation map performance degrades severely as frame transformation errors between map updates accumulate, while foot-centric sensing remains stable.

Figure 6

Figure 6

Figure 6: Robustness of foot proximity sensor policies versus baselines under static joint bias, morphology mismatches, and odometry noise; local sensing consistently outperforms.

These results substantiate the claim that minimal, local sensor feedback provides superior resilience to state estimation errors and sim-to-real transfer.

Hardware Deployment and Behavioral Analysis

The optimized policy was deployed on ANYmal-D and evaluated on real-world challenging terrains: long gaps, sparse/irregular stepping stones with variable heights, and narrow beams.

Figure 7

Figure 7: Policy traversal on physical stepping stone arrangements of increasing difficulty using only underfoot proximity sensors.

Behavioral logs during traversal highlight emergent active perception strategies. The robot leverages foot proximity readings to dynamically sense, align, and target viable footholds, actively modifying swing trajectories in response to detected gaps or contact surfaces.

Figure 8

Figure 8: Deployment traces—proximity sensor readings, raw ToF data, and contact-related kinematic cues—demonstrate reflexive pre-contact adaptation.

Strong numerical results include successful gap traversal up to 60 cm60\,\mathrm{cm} and stable crossing of narrow, elevated, and irregularly distributed footholds at average speeds exceeding 0.5 m/s0.5\,\mathrm{m/s}. The policy achieved a high completion rate across all tested terrain morphologies, matching or exceeding benchmarks established by vision-based frameworks but with significantly reduced computational overhead and failure rate due to occlusion or perception stack latency.

Implications and Future Directions

The demonstrated local sensing paradigm represents a shift towards minimalist, task-informed sensor integration. By confining exteroceptive feedback to the immediate vicinity of the contact point, the control policy avoids the compounding of errors inherent in global perception and state estimation stacks. This leads to robust, reflexive locomotion without expensive sensor fusion or high-bandwidth mapping.

Theoretically, these findings challenge the necessity of global terrain modeling for agile locomotion on sparse terrain, advocating instead for high-temporal-bandwidth, spatially focused feedback mechanisms integrated at the limb level. Practically, this opens avenues for legged robotic systems to operate in visually challenging, computationally constrained, or cluttered environments with improved robustness.

Future work may address the scalability of this approach by:

  • Optimizing sensor morphology, number, and placement for holistic near-field awareness
  • Compensating for specific ToF limitations (e.g., adverse optical surface properties) via sensor fusion with tactile or force feedback
  • Investigating IR-transparent or self-cleaning mechanisms for enhanced long-term deployment durability
  • Exploring learned behaviors combining local and global perception to generalize across currently intractable scenarios (e.g., deformable surfaces, severe visibility constraints)

Conclusion

Integrating low-cost, low-latency proximity sensors at the robot’s foot enables robust, anticipatory traversals of complex discrete terrain with minimal reliance on global perception. This sensory minimalism—when paired with modern RL techniques—yields policies that outperform conventional vision-based stacks in robustness to noise, occlusion, and calibration error, realizing a path towards more deployable and adaptive legged robots (2606.31912).

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.