- The paper presents a novel strategy using foot-mounted ToF sensors for immediate underfoot feedback, enabling anticipatory control on discrete terrains.
- It employs a two-stage curriculum with a recurrent (LSTM) policy optimized via PPO, effectively handling diverse challenges like gaps and uneven footholds.
- Empirical evaluations show that the approach outperforms global perception methods in noise resilience and robustness, achieving up to 60 cm gap traversal.
Learning Locomotion on Discrete Terrain via Minimal Proximity Sensing
Introduction and Motivation
The paper "Learning Locomotion on Discrete Terrain via Minimal Proximity Sensing" (2606.31912) addresses a significant bottleneck in the field of legged robot locomotion: the challenge of perceiving and negotiating unstructured terrains characterized by discontinuities and sparse footholds. Traditional reliance on global exteroceptive sensors (LiDAR, depth cameras) is hampered by occlusions, perceptual aliasing, and substantial computational cost, while proprioceptive signals only allow reactive correction after loss of ground contact. The authors propose using sparse, low-power, foot-mounted time-of-flight (ToF) proximity sensors to provide immediate, underfoot geometrical feedback, enabling anticipatory control policies robust to noise, occlusion, and environmental uncertainty.

Figure 1: Local proximity sensing integrated on quadruped feet, yielding policies with superior noise-resilience and locomotion reliability over discrete terrain.
Sensor System and Characterization
The hardware solution integrates the STMicroelectronics VL53L5CX ToF infrared sensor into the feet of an ANYmal-D quadruped. This sensor provides a 4×4 grid of range measurements with a $\ang{45}\times\ang{45}$ FoV, achieving 1mm resolution at 60Hz with measured latencies of 20−60ms. The grid-based FoV ensures coverage of the projected foothold area during swing.


Figure 2: Foot-mounted proximity sensor with 4×4 grid providing dense underfoot coverage.
Comprehensive characterization established that sensor accuracy is optimal at short range and perpendicular incidence, with noise and outlier rates increasing for longer ranges and grazing angles.


Figure 3: Heatmaps quantifying ToF noise and dropped readings as functions of distance and incidence angle.
Sensor noise profiles and failure modes derived from empirical data were modeled in simulation to enable high-fidelity policy transfer under real-world conditions.
Learning-Based Locomotion Framework
The control architecture leverages a recurrent policy (LSTM) optimized via PPO, using Isaac Gym for scalable simulation. Observations consist of proximity sensor readings augmented with proprioceptive signals (joint states, body dynamics). The learning pipeline omits reliance on elevation maps, map-fusion, or expensive attention over 3D point clouds, directly mapping sparse sensor feedback to agile locomotion actions.
Training is structured in two phases:
- Stage 1: Curriculum traverses 5 canonical terrains (dense/row/mixed stones, rough fields, gaps), with task difficulty escalated via foothold geometry, gap sizes, and height variations. This phase establishes core sensor-motor associations.
- Stage 2: Encompasses more challenging variants (smaller footholds, larger gaps, steeper elevation changes) and introduces beam walking, stairs, and pyramid slopes for generalization.


Figure 4: Sequence of discrete terrain morphologies employed during two-stage curriculum learning.
Randomization of sensor degradations (offsets, Gaussian/proportional noise, drops, delays) is incorporated to prevent overfitting and to bridge the sim-to-real gap.
Empirical Robustness Analysis
Rigorous robustness evaluations compare foot-centric proximity sensing to two base-centric benchmarks:
- Mock (Ideal Height Scan): Assumes noise-free, egocentric top-down terrain scans.
- Mock (Aggre. Elevation Map): Fuses multiple chassis-mounted depth images using odometry for registration, simulating real-world elevation map aggregation.


Figure 5: Visualization of diverse observation configurations: foot-centric rays, egocentric height scan, and odometry-fused elevation map.
Performance was assessed by success rate (1000 trials) for terrain traversal under injected degradations:
- Static joint position offsets: Base-centric approaches rapidly lose performance as kinematic errors misalign foot pose estimation, corrupting mapping from global height field to local contact.
- Robot morphology mismatches (e.g., shank length changes): Sensitivity of base-centric mapping to inaccurate kinematic parameters is pronounced.
- Odometry drift: Aggregated elevation map performance degrades severely as frame transformation errors between map updates accumulate, while foot-centric sensing remains stable.


Figure 6: Robustness of foot proximity sensor policies versus baselines under static joint bias, morphology mismatches, and odometry noise; local sensing consistently outperforms.
These results substantiate the claim that minimal, local sensor feedback provides superior resilience to state estimation errors and sim-to-real transfer.
Hardware Deployment and Behavioral Analysis
The optimized policy was deployed on ANYmal-D and evaluated on real-world challenging terrains: long gaps, sparse/irregular stepping stones with variable heights, and narrow beams.

Figure 7: Policy traversal on physical stepping stone arrangements of increasing difficulty using only underfoot proximity sensors.
Behavioral logs during traversal highlight emergent active perception strategies. The robot leverages foot proximity readings to dynamically sense, align, and target viable footholds, actively modifying swing trajectories in response to detected gaps or contact surfaces.

Figure 8: Deployment traces—proximity sensor readings, raw ToF data, and contact-related kinematic cues—demonstrate reflexive pre-contact adaptation.
Strong numerical results include successful gap traversal up to 60cm and stable crossing of narrow, elevated, and irregularly distributed footholds at average speeds exceeding 0.5m/s. The policy achieved a high completion rate across all tested terrain morphologies, matching or exceeding benchmarks established by vision-based frameworks but with significantly reduced computational overhead and failure rate due to occlusion or perception stack latency.
Implications and Future Directions
The demonstrated local sensing paradigm represents a shift towards minimalist, task-informed sensor integration. By confining exteroceptive feedback to the immediate vicinity of the contact point, the control policy avoids the compounding of errors inherent in global perception and state estimation stacks. This leads to robust, reflexive locomotion without expensive sensor fusion or high-bandwidth mapping.
Theoretically, these findings challenge the necessity of global terrain modeling for agile locomotion on sparse terrain, advocating instead for high-temporal-bandwidth, spatially focused feedback mechanisms integrated at the limb level. Practically, this opens avenues for legged robotic systems to operate in visually challenging, computationally constrained, or cluttered environments with improved robustness.
Future work may address the scalability of this approach by:
- Optimizing sensor morphology, number, and placement for holistic near-field awareness
- Compensating for specific ToF limitations (e.g., adverse optical surface properties) via sensor fusion with tactile or force feedback
- Investigating IR-transparent or self-cleaning mechanisms for enhanced long-term deployment durability
- Exploring learned behaviors combining local and global perception to generalize across currently intractable scenarios (e.g., deformable surfaces, severe visibility constraints)
Conclusion
Integrating low-cost, low-latency proximity sensors at the robot’s foot enables robust, anticipatory traversals of complex discrete terrain with minimal reliance on global perception. This sensory minimalism—when paired with modern RL techniques—yields policies that outperform conventional vision-based stacks in robustness to noise, occlusion, and calibration error, realizing a path towards more deployable and adaptive legged robots (2606.31912).