Papers
Topics
Authors
Recent
Search
2000 character limit reached

TactileStep: Sole Tactile Learning for Regulating Foot-Terrain Interaction in Humanoid Locomotion

Published 24 Sep 2026 in cs.RO | (2609.28959v1)

Abstract: Humanoid parkour policies can traverse various terrains, but task completion may mask challenges of harsh landings, edge contacts, and unstable stance contacts. Humans naturally regulate foot-terrain interaction through tactile feedback, modulating contact compliance according to terrain stiffness. This highlights a key domain gap between humans and humanoid robots: the absence of rich tactile sensing in most humanoid systems. We address this problem with TactileStep, a deployable tactile learning framework that brings sole pressure sensing into humanoid locomotion control for softer touchdowns and more stable support. TactileStep aligns tactile simulation with the real pressure insole, allowing the policy to learn from the same contact features available on hardware. During training, we use tactile and motion cues to recognize different foot-contact phases and apply phase-aware rewards that encourage safer landing and more stable stance. Evaluated in simulation and on a Unitree G1 humanoid across diverse terrains, TactileStep reduces peak touchdown force by up to 48.8% and peak A-weighted impact noise by up to 30.1 dB over a strong perceptive baseline, while increasing stance contact area by up to 23.8%.

Summary

  • The paper introduces TactileStep, a learning-based approach that equips a Unitree G1 humanoid with thin plantar pressure insoles to regulate foot-sole pressure and optimize foot-terrain interaction across various gait phases.
  • TactileStep reduces impact force by up to 48.8% and peak noise by up to 30.1 dB and improves stance contact area by up to 23.8%.
  • The research shows that tactile feedback is crucial for reliable regulation of the contact state during humanoid locomotion, particularly on variable and challenging terrains.

Problem formulation and motivation

“TactileStep: Sole Tactile Learning for Regulating Foot-Terrain Interaction in Humanoid Locomotion” (2609.28959) addresses a specific deficiency in current perceptive humanoid locomotion: successful traversal does not imply physically desirable foot–terrain interaction. A vision-based policy may identify a feasible foothold and complete a stair or platform maneuver while still producing excessive touchdown forces, edge-dominated contacts, or unstable stance support. These effects are consequential for whole-body stability, actuator loading, structural durability, acoustic emissions, and operation near humans.

The paper’s central claim is that contact quality should be treated as a closed-loop control objective rather than as an indirect consequence of geometric foothold selection. Vision and depth sensing provide pre-contact information about terrain geometry, but they do not directly reveal the realized pressure distribution after touchdown. Proprioceptive measurements and estimated ground-reaction forces provide related signals, yet do not expose the spatial distribution of load across the sole. TactileStep therefore equips a Unitree G1 humanoid with thin plantar pressure insoles and provides the policy with deployable features describing the realized contact state.

The framework is organized around four phases for each foot: Swing, Pre-Landing, Landing, and Stance. This phase decomposition is important because the relevant control objectives differ temporally. During Pre-Landing, the controller should reduce excessive downward velocity and acceleration. During Landing, it should suppress impact force and force-rate transients. During Stance, it should establish broad support, maintain a centered center of pressure (CoP), and avoid abrupt pressure migration. The method consequently rejects a single undifferentiated contact penalty in favor of phase-conditioned reward terms.

Figure 1

Figure 1: TactileStep maps sole pressure to deployable contact features and routes phase-conditioned tactile objectives through the locomotion policy.

Deployable tactile representation

A major design decision is to avoid exposing the actor to simulator-specific contact quantities that are unavailable on hardware. Instead, TactileStep constructs a lightweight tactile simulator whose output is processed into the same low-dimensional features measured by the real insole. Each simulated sole contains 60 taxels. Raycasts estimate the distance and normal alignment between each taxel and the terrain, and the resultant rigid-body normal contact force is distributed across taxels according to proximity and orientation. A spatial diffusion step then spreads load over neighboring taxels to approximate the compliant load distribution of a real sole and to suppress isolated point contacts generated by rigid-body collision solvers.

The policy does not receive the full pressure map. It receives, for each foot, normalized normal force, contact-area ratio, and two-dimensional CoP. These features are physically interpretable and less sensitive than raw taxel readings to sensor layout, manufacturing variation, calibration, and sim-to-real mismatch. A two-frame tactile history is concatenated with proprioceptive history and depth-observation history. The actor therefore receives contact feedback while retaining the perceptual information required for foothold selection and terrain traversal.

This compression imposes an explicit information bottleneck. The representation discards fine-grained pressure topology, shear force, and other potentially useful tactile variables. The paper argues that force magnitude, occupied contact area, and CoP contain the principal information needed for touchdown regulation and stance support, but it does not establish that these features are sufficient for arbitrary contact geometries or deformable surfaces.

The simulator is deliberately not a soft-body model. It approximates compliant pressure formation through force weighting and diffusion, which permits massively parallel reinforcement learning at manageable computational cost. Feature-level comparisons under a tactile-independent baseline show similar ranges and temporal trends between simulation and hardware during stair descent, although residual discrepancies remain. This alignment strategy is therefore pragmatic rather than physically complete: it attempts to match the policy-relevant statistics instead of reproducing the full mechanics of plantar deformation.

Phase-aware reinforcement learning

TactileStep formulates locomotion as a POMDP and trains a PPO policy whose 29-dimensional action is interpreted as target joint positions for a PD controller. The total objective combines task tracking, regularization, safety, AMP motion-prior terms, and tactile contact-quality objectives. Training uses 2,048 simulated humanoids and 50,000 PPO iterations, with domain randomization applied to dynamics, actuation, observations, depth sensing, tactile force scale, taxel positions, and measurement delay.

The gait-phase estimator combines tactile force, contact-area activation, foot height, and vertical foot velocity. A foot without contact is classified as either Swing or Pre-Landing depending on whether it is descending toward the terrain. Once contact is detected, a fixed landing window distinguishes Landing from Stance. This estimator is hand-designed and depends on thresholds, but it gives the reward system a physically meaningful temporal structure without requiring an externally imposed periodic gait clock.

The Pre-Landing objective penalizes excessive downward velocity and acceleration. The Landing objective penalizes both the instantaneous normalized force and its positive increment, with additional peak statistics over the landing window. The Stance objective rewards contact area and CoP margin while penalizing abrupt CoP displacement. These terms encode a specific landing–support trade-off: reducing impact should not be achieved by producing a weak or poorly centered contact.

The method also uses dual critics. Dense rewards, such as velocity tracking, posture regulation, action smoothness, and continuously active tactile shaping, are assigned to one critic; sparse or event-triggered signals, including peak landing penalties, target completion, safety violations, and contact events, are assigned to another. Their advantages are mixed for the actor update. The rationale is to reduce value-estimation interference between continuous locomotion objectives and intermittent contact-quality events. The appendix reports that the dual-critic policy outperforms a single-critic variant in most contact metrics. For example, on stair descent, touchdown force decreases from 579.83 N with a single critic to 320.74 N with the dual critic, while contact area increases from 0.516 to 0.545. This result supports the architectural motivation, although it does not isolate whether the benefit arises specifically from critic separation or from associated optimization effects.

Simulation evaluation

Simulation experiments use 4,096 evaluation trials per policy and terrain, with six terrain categories: flat ground, slopes, stair ascent, stair descent, platform step-up, and platform drop-down.

Figure 2

Figure 2: Parallel simulation evaluation across flat, sloped, stair, and platform terrains.

TactileStep consistently improves contact quality relative to both the external perceptive baseline and the relevant ablations. For sustained support on flat ground, the contact-area ratio reaches 0.929, compared with 0.912 for the baseline and 0.750 without the stable-support reward. On slopes, the corresponding values are 0.805, 0.804, and 0.624. The CoP margin follows the same pattern: TactileStep achieves 35.79 mm on flat ground and 34.24 mm on slopes, while the no-stable-support variant reaches only 29.90 mm and 27.55 mm, respectively.

The gains are more pronounced on stairs, where foothold geometry makes post-contact support more difficult. During stair ascent, TactileStep obtains a contact-area ratio of 0.536, compared with 0.480 for the baseline and 0.462 without stable-support rewards. Its CoP margin is 29.05 mm, compared with 26.03 mm for the baseline and 25.56 mm for the ablation. During stair descent, the contact-area ratio is 0.545 versus 0.487 for the baseline, while the CoP margin is 29.43 mm versus 27.07 mm. These results indicate that the support reward has its greatest effect where geometric discontinuities make complete and centered sole contact difficult.

The impact-force ablation attributes touchdown improvement primarily to the soft-landing reward. Removing the soft-landing terms consistently increases simulated impact force across terrain types. Removing tactile observations also degrades the hardware-relevant contact metrics, demonstrating that the reward alone is insufficient if the deployed actor cannot observe the realized contact state.

The locomotion objective is not substantially sacrificed. Success rates are approximately 100% for TactileStep on all reported simulated terrain categories: 100% on stair ascent, stair descent, platform step-up, and platform drop-down, 99.98% on flat ground, and 99.63% on slopes. The baseline is comparable on most terrains but falls to 92.26% on platform drop-down. Velocity-tracking error and traversal time are generally slightly worse for TactileStep. The clearest cost is energy: on stair ascent, energy increases from 802.2 J for the baseline to 1,126.3 J for TactileStep; on stair descent, it increases from 887.4 J to 1,232.9 J. Thus, the method improves contact quality without reducing task success, but it does so with a nontrivial energetic penalty, plausibly caused by more active regulation and marginally longer or more corrective motions.

Hardware deployment

The hardware experiments deploy the learned controllers on a Unitree G1 with pressure insoles. Each condition uses 20 samples. Impact force is measured from the first rising edge after touchdown, while acoustic measurements use peak A-weighted sound pressure level. The insole operates wirelessly at 25 Hz; supplementary 100 Hz tests on flat ground and platform drop-down produce comparable force estimates, providing some validation of the measurement protocol.

Figure 3

Figure 3: Hardware touchdown-force extraction and acoustic measurement configuration during stair locomotion.

The strongest hardware evidence concerns impact reduction. On stair ascent, TactileStep reduces impact force from 499.9 N for the perceptive baseline to 349.8 N, a reduction of approximately 30.0%. On platform ascent, the reduction is from 695.0 N to 355.7 N, approximately 48.8%, which is the largest reported force improvement. On platform descent, force decreases from 608.9 N to 404.6 N, approximately 33.6%. The corresponding acoustic reductions are also substantial:

Terrain Baseline impact TactileStep impact Baseline noise TactileStep noise
Stair ascent 499.9 N 349.8 N 101.2 dB 72.2 dB
Stair descent 359.3 N 344.1 N 97.2 dB 67.1 dB
Platform ascent 695.0 N 355.7 N 110.1 dB 83.0 dB
Platform descent 608.9 N 404.6 N 90.8 dB 75.9 dB
Flat 201.3 N 191.3 N 90.6 dB 66.5 dB
Slope 255.0 N 202.0 N 83.4 dB 69.9 dB

The maximum acoustic improvement is 30.1 dB on stair descent. Because the background noise floor is approximately 55–60 dB and is dominated by the robot’s fan, these measurements concern impact-induced peaks rather than an isolated measurement of foot noise. Nevertheless, the consistent reduction in force and sound supports the claim that TactileStep produces physically softer touchdowns rather than merely changing the acoustic signature.

The support results show that lower impact does not require reduced load-bearing contact. On stair ascent, contact area increases from 0.470 for the baseline to 0.482; on stair descent, it increases from 0.483 to 0.598, a relative improvement of approximately 23.8%. On flat ground, the increase is from 0.453 to 0.510. Slope performance is similar but still favorable, increasing from 0.486 to 0.499. Contact-area measurements are not reported for platform step-up and drop-down because those trials emphasize a single transition rather than sustained stance.

The tactile-observation ablation is informative. On platform ascent, removing online tactile observations yields 371.5 N, compared with 355.7 N for the complete method; on stair descent, it yields 353.3 N versus 344.1 N. The differences in force are sometimes modest, but the acoustic and support gaps are more pronounced. On stair descent, the no-tactile-observation policy produces 84.9 dB and a contact-area ratio of 0.417, whereas TactileStep produces 67.1 dB and 0.598. This supports the paper’s stronger claim that deployment-time tactile feedback is necessary for reliable regulation of the contact state, not merely useful as a training-time proxy.

Limitations and open questions

The evaluation is broad in terrain type but bounded in operating regime. Commands are sampled around a forward speed of 0.45–0.55 m/s with variable yaw rate, so the method’s behavior at substantially higher speeds is unresolved. Faster locomotion would shorten the Landing phase and amplify the consequences of force-estimation latency, pressure-sensor sampling, and phase-classification errors.

The tactile simulator also assumes rigid terrain and approximates sole compliance through diffusion. Its validity on deformable, granular, slippery, or highly irregular surfaces is not established. The feature representation excludes shear forces and does not model the mechanical coupling between sole deformation and pressure redistribution. Consequently, the reported sim-to-real agreement should not be interpreted as validation of the simulator’s contact mechanics beyond the tested regime.

The gait-phase estimator depends on hand-designed thresholds for force, area, height, and vertical motion. This introduces a potential failure mode under unusual contact sequences, partial footholds, rapid reversals, or simultaneous multi-contact events. The experiments do not systematically report phase-estimation errors, policy failures, recovery behavior, or the distribution of unsuccessful trials beyond aggregate success rates.

Finally, long-term pressure-insole drift, wireless reliability, sensor durability, and recalibration are not evaluated. The hardware results establish short-horizon deployment feasibility, but not robustness over extended operation. The paper therefore leaves open whether the compact features remain calibrated under wear, temperature variation, repeated high-load impacts, or changes in footwear and sole mechanics.

Conclusion

TactileStep reframes humanoid locomotion contact as a controllable state rather than a passive outcome of visual foothold selection. Its principal technical contribution is the alignment of simulated and hardware plantar sensing through compact force, contact-area, and CoP features, combined with phase-aware rewards and a dual-critic PPO architecture. Across simulation and Unitree G1 experiments, the method preserves near-perfect traversal success while reducing hardware touchdown force by up to 48.8%, peak acoustic noise by up to 30.1 dB, and increasing stance contact area by up to 23.8%. The results establish that sole tactile feedback can improve both impact mitigation and post-touchdown support, although the gains entail higher simulated energy consumption and remain bounded by assumptions about terrain rigidity, sensing calibration, gait-phase estimation, and operating speed.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

Explain it Like I'm 14

1. What is this paper about?

This paper introduces TactileStep, a system that helps a humanoid robot walk and jump over difficult surfaces more safely.

Many robots can already walk across stairs, slopes, and platforms. However, simply reaching the destination does not mean every step is good. A robot might:

  • Hit the ground too hard.
  • Land on only the edge of its foot.
  • Make a loud noise.
  • Lose balance because its foot is not supporting it properly.

Humans use the sensitive skin on the bottoms of their feet to notice how they are standing and landing. TactileStep gives a robot a similar ability by placing thin pressure sensors inside its shoes.

The main idea is simple:

The robot should not only decide where to put its foot. It should also feel how the foot touches the ground and adjust its movement.

2. What questions did the researchers ask?

The researchers wanted to find out three main things:

  1. Can pressure sensing help the robot land more gently and quietly?
  2. Can it help the robot create a wider and more stable foot contact with the ground?
  3. Which parts of the new system are most useful?

For example, they tested whether the improvements came from:

  • Giving the robot pressure information while it walks.
  • Teaching it to reduce impact during landing.
  • Teaching it to keep its foot stable after landing.

3. How did the researchers build and test the system?

Pressure sensors in the robot’s feet

The robot used was a Unitree G1 humanoid robot. Thin pressure insoles were placed inside its feet. These insoles measured where and how strongly the foot was pressing against the ground.

The system summarized the pressure into three easy-to-use measurements:

  • Force: How strongly the foot is pushing down.
  • Contact area: How much of the sole is touching the ground.
  • Center of pressure: The average location where the pressure is concentrated.

The center of pressure is similar to finding the balance point of a seesaw. If it moves too close to the edge of the foot, the robot may become unstable.

A simplified virtual pressure sensor

Before testing the real robot, the researchers trained the system in a computer simulation. Training in simulation is faster and safer than immediately trying everything on a real robot.

They divided each virtual foot into 60 small areas, called taxels. The simulator estimated how much force each area would feel when the foot touched the ground. It then spread the force across nearby areas to imitate the slight softness of a real shoe sole.

Instead of giving the robot a complicated picture of all 60 pressure points, the researchers gave it the three simpler measurements: force, contact area, and center of pressure. This made it easier for information learned in simulation to work on the real robot.

Four parts of every step

TactileStep divided each foot’s movement into four phases:

  1. Swing: The foot is moving through the air.
  2. Pre-landing: The foot is close to the ground but has not touched it yet.
  3. Landing: The foot has just made contact.
  4. Stance: The foot is supporting the robot’s body.

This is important because different moments need different actions. For example, the robot should slow its foot before landing, reduce the impact during landing, and keep its pressure centered while standing.

Reinforcement learning

The researchers used reinforcement learning, which is a way of training by giving a computer program rewards and penalties.

An everyday analogy is teaching a dog a trick:

  • Good behavior receives a reward.
  • Bad behavior receives a smaller reward or a penalty.
  • After many attempts, the dog learns which actions work best.

In this paper, the robot received rewards for walking successfully and penalties for:

  • Moving its foot downward too quickly before landing.
  • Creating a large impact force.
  • Shifting pressure suddenly.
  • Touching the ground with only a small part of its foot.
  • Putting pressure too close to the edge of the foot.

The researchers also compared TactileStep with a strong vision-based walking system and with versions of their own system that had certain parts removed.

4. What did they discover?

Softer landings

TactileStep reduced the force of the robot’s foot hitting the ground. In the experiments, the peak landing force was reduced by as much as 48.8% compared with the main comparison system.

The robot also became quieter. Its peak impact noise was reduced by up to 30.1 decibels in some tests. This matters because loud impacts can be annoying to people and can also create vibrations that slowly damage the robot.

More stable support

The robot’s foot usually touched the ground over a larger area. In some tests, the contact area increased by as much as 23.8%.

A larger contact area is like standing with your whole foot flat on the floor instead of balancing on your toes or on the edge of your shoe. It gives the robot a stronger base and makes it less likely to wobble or fall.

The robot also kept its center of pressure farther from the edge of its foot. This gave it a larger safety margin.

The robot still completed difficult terrain tasks

The new system did not merely make the robot move slowly or avoid difficult surfaces. In simulation, it successfully crossed stairs, slopes, flat ground, and platforms at rates similar to or better than the comparison system.

However, the improved behavior used somewhat more energy. The robot sometimes took slightly longer and moved its joints more actively in order to control its landings and support.

Pressure feedback really mattered

When the researchers removed pressure information from the robot’s observations, performance became worse, especially for:

  • Reducing noise.
  • Increasing contact area.
  • Controlling the actual landing behavior.

This suggests that vision alone cannot tell the robot everything it needs to know. A camera can show the robot what the ground looks like before contact, but pressure sensors tell it what actually happened after the foot touched down.

5. Why is this research important?

Earlier walking systems mainly focused on one question:

Can the robot get across the terrain?

TactileStep asks a more detailed question:

Can the robot get across the terrain while landing gently and standing securely?

This could be useful for robots that work:

  • Around people.
  • Inside homes, offices, or hospitals.
  • On stairs and uneven surfaces.
  • For long periods, where repeated hard impacts could cause wear.
  • In places where loud noise is a problem.

The work also shows how human-like sensing can improve robots. People naturally feel pressure under their feet and adjust their movements without thinking about it. Giving robots similar information may help them become safer and more physically aware.

6. Limitations and future improvements

The researchers point out that TactileStep is not perfect yet. For example:

  • It has mainly been tested at a limited range of speeds.
  • Its virtual pressure model assumes fairly hard surfaces.
  • It has not been fully tested on sand, mud, or soft ground.
  • The long-term durability and accuracy of the pressure insoles are still unknown.
  • The system does not yet thoroughly study how the robot recovers after a failed step.
  • The four walking phases are identified using hand-designed rules rather than learned automatically.

Conclusion

TactileStep gives a humanoid robot pressure-sensitive feet and teaches it to use that information while walking. The robot learns to slow down before landing, reduce the force of impact, and place more of its foot firmly on the ground.

The experiments show that this can make humanoid locomotion softer, quieter, and more stable without preventing the robot from crossing challenging terrain. In the future, this type of tactile feedback could help humanoid robots operate more safely near people and last longer during everyday use.

Knowledge Gaps

Knowledge Gaps, Limitations, and Open Questions

  • Generalization to high-speed locomotion remains untested. The policy is evaluated within a bounded command range, so its performance under running, jumping, rapid stair traversal, and other motions with shorter contact durations and larger impact transients is unknown.
  • Recovery from poor or unexpected contacts is not characterized. The paper does not systematically evaluate whether tactile feedback enables recovery from edge landings, partial footholds, slipping, missed steps, or sudden changes in support.
  • Failure cases are not reported in sufficient detail. Success rates and average contact metrics do not reveal the types, frequencies, or consequences of failures, such as falls, foot scuffing, unstable post-contact oscillations, or repeated corrective steps.
  • The benefits of tactile feedback for balance recovery are unresolved. The controller improves landing and stance metrics, but the experiments do not test perturbation rejection, external pushes, sudden support loss, or dynamic balance after touchdown.
  • The tactile simulator does not model compliant or deformable contact. Its rigid-contact force distribution and heuristic spatial diffusion have not been validated on carpets, mats, soft flooring, sand, gravel, mud, grass, or other deformable and granular surfaces.
  • Tactile transfer across terrain materials is unexplored. The study does not determine whether the learned force, contact-area, and CoP representations remain reliable when friction, compliance, roughness, or surface damping changes substantially.
  • Shear and tangential forces are omitted. The representation contains only normal force, contact area, and CoP, leaving slip detection, lateral loading, tangential friction, and torsional contact behavior unmodeled.
  • The low-dimensional tactile representation may discard important spatial information. The paper does not establish whether full pressure maps, learned spatial embeddings, or additional features such as pressure gradients and contact topology would improve performance on partial or irregular footholds.
  • The robustness of feature extraction to sensor artifacts is unknown. The effects of taxel failure, saturation, dead zones, nonlinear response, temperature variation, wiring faults, and calibration errors are not evaluated.
  • Long-term pressure-insole durability is unvalidated. The paper does not report sensor lifetime, mechanical wear, hysteresis, drift, contamination, or performance degradation over extended locomotion and repeated impacts.
  • Online recalibration and maintenance procedures are missing. It remains unclear how frequently the insoles must be recalibrated, whether calibration can be performed automatically, and how the policy behaves when calibration changes over time.
  • The gait-phase estimator may be brittle outside the tested regime. Its hand-designed thresholds and finite-state logic could fail under unusual speeds, irregular gait patterns, delayed contact sensing, multi-contact events, or highly compliant terrain; these cases are not systematically tested.
  • The effect of phase-estimation errors is not quantified. The paper does not measure how false contact detections, missed contacts, or incorrect transitions between pre-landing, landing, and stance affect policy safety and performance.
  • The choice of phase definitions and thresholds is not compared against alternatives. There is no ablation of the four-phase design versus a two-phase clock, continuous phase variables, learned phase representations, or event-based contact segmentation.
  • Reward-weight sensitivity is not examined. The paper does not show whether the reported improvements are robust to changes in the relative weights for impact force, force growth, contact area, CoP margin, and CoP motion.
  • The trade-off between contact quality and locomotion efficiency is unresolved. TactileStep consumes substantially more simulated energy and can increase traversal time, but the paper does not identify the Pareto frontier between softness, stability, speed, energy, and task success.
  • The source of the increased energy consumption is unclear. It is not determined whether the additional cost comes from slower motion, higher joint torques, more corrective actions, altered foot placement, or the phase-aware reward structure.
  • The causal contribution of individual tactile features is incomplete. The reported ablations remove tactile observations or entire reward groups, but do not isolate normal force, contact area, CoP position, CoP margin, or tactile history separately.
  • The value of tactile history is not established. The study does not compare instantaneous tactile features with different history lengths, recurrent models, temporal filters, or event-based representations.
  • The sim-to-real alignment analysis is limited in scope. Alignment is illustrated using short traces from a single stair-descent segment and a tactile-independent policy, without quantitative distributional errors, broader terrains, repeated trials, or statistical confidence intervals.
  • The tactile simulator’s physical parameters are not fully identified or validated. The paper does not report systematic calibration or sensitivity analyses for taxel layout, diffusion strength, activation thresholds, normal-alignment thresholds, and force-distribution parameters.
  • The evaluation uses a single humanoid platform. Results on the Unitree G1 do not establish whether the method transfers to robots with different foot geometries, actuator characteristics, sole compliance, body mass, kinematics, or tactile-sensor layouts.
  • Sensor placement and sole geometry generalization are unexplored. It is unknown whether the compact features remain reliable when pressure insoles are repositioned, partially covered, installed in different footwear, or used with nonplanar and differently sized feet.
  • The comparison with the external baseline is not fully controlled. The paper does not clarify whether the baseline has equivalent training budgets, motion priors, perception inputs, command distributions, hardware tuning, and deployment conditions.
  • Hardware sample sizes are small. Real-world evaluation uses only 20 samples per condition, and the paper does not report confidence intervals, statistical significance, trial randomization, or effects of operator and environmental variability.
  • Acoustic measurements may not generalize beyond the experimental setup. Noise is measured with a microphone mounted approximately 6 cm above the lower joint, so the reported peak A-weighted levels may depend strongly on microphone placement, robot structure, room acoustics, and fan noise.
  • The relationship between acoustic noise and mechanical impact is not isolated. The study reports correlated force and sound reductions but does not separate foot-ground impact noise from actuator, structure-borne, fan, or environmental noise.
  • Contact-area improvement is not linked to whole-body stability quantitatively. Larger sole contact area and CoP margin are treated as proxies for stability, but the paper does not measure zero-moment-point margins, angular momentum, slip probability, fall probability, or center-of-mass recovery.
  • The safety implications of softer landings are not fully assessed. Lower peak force may be achieved through altered timing or force redistribution, but joint loads, peak torque, impulse, vibration transmission, and cumulative mechanical stress are not reported.
  • Human-centered deployment remains hypothetical. The paper motivates quiet and gentle locomotion near people but does not evaluate perceived safety, human comfort, collision risks, proximity to pedestrians, or performance in realistic shared environments.
  • The method’s behavior under unexpected terrain changes is unknown. The experiments use predefined terrain classes and do not systematically test abrupt changes in height, slope, friction, compliance, obstacle location, or terrain geometry after policy execution begins.
  • The policy’s dependence on depth perception is not separated from the tactile contribution. The study does not evaluate degraded, delayed, occluded, or noisy depth sensing, making it unclear whether tactile feedback can compensate for failures in visual terrain perception.
  • Multi-contact interactions are not investigated. The framework focuses on sole–terrain contact and does not address simultaneous hand, knee, toe, or other body contacts that may occur during humanoid parkour or recovery.
  • The approach is not evaluated during prolonged operation. Long-horizon effects such as cumulative sensor drift, thermal changes, repeated-impact fatigue, policy degradation, and changes in robot dynamics remain open.
  • No formal safety or stability guarantees are provided. The learned controller is evaluated empirically, but there are no theoretical guarantees that tactile rewards or CoP margins prevent falls, slipping, excessive loads, or unsafe contact configurations.

Practical Applications

Immediate Applications

  • Quieter and safer humanoid operation in human-centered environments — Robotics, healthcare, hospitality, offices, homes. The TactileStep controller can be integrated into humanoid robots that already use depth sensing, proprioception, and joint-level PD control. Sole-pressure features—normal force, contact area, and center of pressure (CoP)—can be used to reduce harsh touchdowns and acoustic noise during walking on stairs, platforms, slopes, and indoor floors. This is immediately relevant to service robots operating near people, where reported reductions in impact force and peak A-weighted noise can improve comfort and perceived safety. Potential products/workflows: a “quiet walking” locomotion mode, terrain-specific gait profiles, and contact-quality dashboards for robot operators. Dependencies: compatible sole-pressure hardware, reliable calibration, sufficient onboard compute, and validation on the target robot rather than assuming direct transfer from the Unitree G1.
  • Retrofit tactile sensing for commercial humanoids — Robotics hardware and controls. Thin pressure insoles or embedded plantar sensor arrays can be added to existing humanoid platforms without redesigning the entire foot. The paper’s low-dimensional representation is particularly suitable for deployment because it abstracts away from the exact taxel layout. Manufacturers could use the features as an additional observation channel for existing perceptive locomotion policies. Potential products/workflows: standardized foot-sensor modules exposing force, contact-area, and CoP APIs; calibration tools that map raw sensor values to normal force; and middleware for tactile observations in ROS-based controllers. Dependencies: sufficient sensor durability, sampling rate, waterproofing, mechanical protection, and stable sensor-to-foot registration.
  • Improved stair, curb, and platform traversal — Construction, logistics, inspection, and public-space robotics. The phase-aware controller can be deployed for robots that frequently encounter height discontinuities. Pre-landing feedback can limit downward foot velocity, landing feedback can reduce force spikes, and stance feedback can encourage broader, more centered support. This is especially actionable for stair climbing and descending, where the paper reports larger advantages than on flat terrain. Potential workflows: tactile-aware stair-climbing controllers for warehouse robots, building-inspection robots, and delivery or security humanoids. Dependencies: terrain geometry must remain within the trained command and motion range; the current approach has not been validated systematically on loose, deformable, or granular surfaces.
  • Contact-quality monitoring and predictive maintenance — Industrial robotics and fleet operations. The same signals used for control can serve as operational health indicators. Repeatedly high touchdown forces, abnormal CoP shifts, reduced contact area, or asymmetric loading could flag worn soles, damaged actuators, misalignment, or deteriorating pressure sensors. Fleet software could log these metrics by terrain and mission. Potential products/workflows: maintenance alerts, impact-history logs, per-foot sensor diagnostics, and automatic policies that switch to a conservative gait after detecting degraded contact quality. Dependencies: long-term sensor drift and insole durability must be characterized; thresholds need to be calibrated against actual mechanical failure modes.
  • Benchmarking and evaluation tools for locomotion research — Academic robotics and industrial R&D. The paper provides practical metrics beyond binary task success: touchdown impact force, contact-area ratio, CoP margin, peak acoustic noise, velocity-tracking error, traversal time, and energy consumption. Research groups can add these metrics to locomotion benchmarks to distinguish policies that merely complete a route from those that establish safe and stable contact. Potential tools: standardized test suites in simulation and hardware, tactile sim-to-real alignment reports, and ablation protocols comparing vision-only, proprioceptive, and tactile policies. Dependencies: measurement procedures must be standardized across robot models, sensor layouts, floor materials, and microphones.
  • Simulation-based training for tactile locomotion — Robotics software and reinforcement learning. The lightweight tactile simulator can be incorporated into Isaac Sim/Isaac Lab pipelines. Rather than modeling full soft-body mechanics, researchers can distribute rigid-body contact forces across virtual sole taxels, apply spatial diffusion, and expose compact features to the policy. This offers a practical route to train policies at scale while reducing dependence on simulator-only contact signals. Potential tools: reusable tactile-simulation plugins, feature-level sim-to-real calibration modules, and policy libraries for phase-conditioned reward shaping. Dependencies: the approximation is most credible for rigid or moderately compliant contact and may require reparameterization for deformable terrain or different foot geometries.
  • Phase-aware reward design for other legged robots — Quadrupeds, bipeds, exoskeletons, and mobile manipulators. The four phases—swing, pre-landing, landing, and stance—provide a transferable structure for routing rewards or safety constraints to the moments when they are physically meaningful. Similar designs could regulate foot placement, impact, and support in quadrupeds, powered exoskeletons, or rehabilitation robots. Dependencies: phase thresholds currently rely on hand-designed force, area, height, and vertical-velocity cues; different morphologies and gait patterns will require new phase logic.
  • Safety-oriented robot operation and policy constraints — Industrial policy and certification. Contact metrics can be incorporated into operational limits: maximum allowable touchdown force, minimum stance contact area, or minimum CoP margin. A robot could automatically slow down, reject a foothold, or trigger recovery when these limits are violated. This supports risk assessments for robots working near people or expensive equipment. Dependencies: the paper demonstrates improved average metrics but does not provide systematic failure-rate, recovery, or safety-certification analysis; formal limits require broader testing.

Long-Term Applications

  • Humanoid robots for homes, hospitals, and eldercare — Healthcare and domestic robotics. Quiet, low-impact walking could enable humanoids to operate around sleeping patients, elderly users, children, and fragile household objects. Tactile feedback could also help detect uncertain support on rugs, thresholds, ramps, and partially obstructed floors. Potential products/workflows: bedside assistance robots, hospital delivery and support robots, and domestic assistants with automatic “gentle mode” locomotion. Dependencies: substantially broader testing is needed for carpets, wet floors, deformable mats, clutter, unexpected obstacles, human contact, and emergency recovery. Sensor hygiene, sterilization, and long-term reliability are also important in healthcare.
  • Adaptive locomotion across deformable, granular, or slippery terrain — Search and rescue, mining, agriculture, planetary exploration. Extending the pressure representation to include shear forces, friction estimates, pressure redistribution, and terrain compliance could allow robots to adapt their foot loading to sand, mud, gravel, snow, soft soil, or damaged infrastructure. The current work establishes a foundation but validates only a rigid-contact approximation. Potential products/workflows: terrain-adaptive foothold selection, slip detection, and online adjustment of compliance or step timing. Dependencies: richer tactile simulation, calibrated contact mechanics, shear sensing, and extensive sim-to-real testing are required.
  • Closed-loop foothold selection combining vision and touch — Advanced robotics and autonomous navigation. A future planner could use vision or depth to propose candidate footholds and then use tactile feedback to verify whether the realized contact is safe. If contact area or CoP margin is poor, the robot could redistribute load, reposition the foot, or initiate a recovery step. This would address the gap between geometric foothold feasibility and actual post-touchdown support. Potential tools: tactile-aware model-predictive control, contact-state estimators, and planners that rank footholds by predicted landing force and support margin. Dependencies: low-latency tactile processing, reliable contact-state prediction, and integration with whole-body balance and recovery controllers.
  • High-speed parkour and dynamic humanoid motion — Agile robotics, defense, and emergency response. Tactile regulation could make running, jumping, obstacle negotiation, and rapid descent more robust by controlling large impact transients. However, the paper explicitly limits evaluation to a bounded command range and does not establish behavior at substantially higher speeds. Potential products/workflows: impact-aware running controllers, safe jump-landing modules, and dynamic obstacle-crossing policies. Dependencies: faster sensors and control loops, actuator and structural limits, more accurate impact modeling, and explicit recovery behavior for missed or partial footholds.
  • Learning phase representations rather than relying on hand-designed rules — Machine learning for robotics. The current phase estimator uses thresholds on force, contact area, foot height, and vertical velocity. A learned phase representation could adapt to irregular gait timing, speed changes, damaged sensors, and unusual terrain. It could also support continuous phase uncertainty rather than discrete labels. Potential tools: recurrent contact-state estimators, probabilistic phase classifiers, and self-supervised tactile encoders. Dependencies: sufficient labeled or self-supervised data, robustness to sensor failure, and mechanisms preventing phase-estimation errors from destabilizing the policy.
  • Standardized tactile interfaces across robot platforms — Robotics ecosystem and academia–industry transfer. Because the policy uses normalized force, contact-area ratio, and CoP rather than raw taxel maps, different foot-sensing systems could potentially share a common observation interface. This could support portable locomotion policies and common evaluation datasets across humanoid platforms. Potential products/workflows: hardware-agnostic tactile APIs, shared calibration datasets, and pretrained contact-regulation policies. Dependencies: consistent definitions and normalization procedures, sensor cross-calibration, variation in sole geometry, and validation across multiple robot morphologies.
  • Energy-aware contact regulation — Energy-efficient robotics and fleet economics. The results show that improved contact quality can come with higher energy consumption, likely because of longer traversal times and more active joint regulation. Future work could optimize impact reduction, support stability, traversal time, and energy jointly, producing terrain- and mission-specific trade-offs. Potential products/workflows: “quiet,” “balanced,” and “energy-saving” locomotion modes; fleet-level optimization of battery use versus hardware wear. Dependencies: multi-objective reward tuning, reliable energy and actuator-wear models, and mission-specific definitions of acceptable noise and impact.
  • Human–robot interaction based on contact-aware behavior — Social robotics and workplace safety. Lower noise and gentler foot impacts could improve user acceptance, while pressure-based detection of unstable support could reduce unexpected stumbling near people. Contact-quality signals might also be combined with whole-body contact sensing to distinguish terrain interaction from collisions or human contact. Dependencies: the current study addresses foot–terrain interaction only; human-contact detection, collision avoidance, legal safety requirements, and predictable failure recovery require separate research.
  • Clinical and biomechanical analysis of robotic gait — Rehabilitation, prosthetics, and biomechanics. The force, contact-area, and CoP features could be adapted for robotic prostheses, lower-limb exoskeletons, or rehabilitation devices to monitor plantar loading and improve balance assistance. The phase-aware structure is compatible with gait-cycle analysis and could help personalize assistance during heel strike, foot flat, and stance. Dependencies: human biomechanics differ from humanoid robot dynamics; clinical validation, patient-specific calibration, medical-device regulation, and richer sensing—including shear and regional pressure—would be necessary.

Glossary

  • A-weighted sound level: A frequency-weighted measure of sound pressure level that approximates human hearing sensitivity. “peak A-weighted impact noise”
  • Adversarial motion prior: A learned motion-style constraint that encourages generated behavior to resemble a reference motion distribution. “adversarial motion prior terms”
  • Center of pressure (CoP): The point on a support surface where the resultant pressure or ground-reaction force acts. “center of pressure (CoP)”
  • CoP margin: The distance between the center of pressure and the boundary of the available support region. “the center-of-pressure margin to the support boundary”
  • Contact-area ratio: The fraction of sensing elements or sole area that is actively supporting the robot. “Contact area ratio AcA_c is measured during stance”
  • Contact compliance: The degree to which a contact interaction deforms or yields in response to applied forces. “modulating contact compliance according to terrain stiffness”
  • Contact transient: A short-lived change in force or motion occurring when contact is initiated. “impact transients can induce vibration”
  • Critic: A reinforcement-learning function that estimates expected future reward from a state or observation. “We use two critics with the same privileged observation”
  • Depth observation: A sensor-derived representation of the environment’s distance structure, commonly obtained from depth cameras or range sensors. “Ht\mathcal{H}_{t} is the depth-observation history”
  • Dense reward: A reinforcement-learning signal provided frequently and continuously during an episode. “dense terms provide frequent, continuous per-step feedback”
  • Domain gap: The discrepancy between data distributions or physical behavior in two domains, such as simulation and reality. “This highlights a key domain gap between humans and humanoid robots”
  • Foot–terrain interaction: The mechanical contact, force exchange, and support relationship between a robot’s foot and the ground. “regulating foot–terrain interaction in humanoid locomotion”
  • Foothold: A location or surface region on which a robot places its foot to support locomotion. “stable locomotion also depends on the support formed after touchdown”
  • Foothold feasibility: The degree to which a candidate foot-placement location can safely and physically support the robot. “These objectives serve as effective geometric proxies for foothold feasibility”
  • Ground reaction force (GRF): The force exerted by the ground on a contacting robot foot, typically opposing the force applied by the robot. “proprioception and estimated ground reaction forces (GRFs) provide impact-related feedback”
  • Gait phase: A temporally defined portion of a locomotion cycle, such as swing, landing, or stance. “its gait phase: swing, pre-landing, landing or stance”
  • Impact transient: A brief, high-magnitude force response produced at the moment of collision or touchdown. “regulating impact at touchdown”
  • Isaac Lab: A simulation and robot-learning framework built for large-scale physics-based training. “using parallel simulation for RL”
  • Landing window: The time interval used to measure or penalize forces associated with foot touchdown. “within the landing window Wlandf\mathcal{W}^{f}_{\mathrm{land}}”
  • Markov decision process (MDP): A sequential decision model in which the current state contains all information needed to predict future transitions and rewards. “We formulate the control problem as a Partially Observable Markov Decision Process (POMDP)”
  • Normal force: The component of a contact force perpendicular to the contacting surface. “normal force, contact area, and center of pressure (CoP)”
  • Normal-alignment threshold: The minimum allowed alignment between a foot’s surface normal and the terrain normal for a taxel to contribute strongly to force estimation. “where ηn\eta_n is the minimum normal-alignment threshold”
  • Partially Observable Markov Decision Process (POMDP): A Markov decision process in which the agent cannot directly observe the complete underlying state. “We formulate the control problem as a Partially Observable Markov Decision Process (POMDP)”
  • Perceptive locomotion: Robot locomotion that uses environmental sensing, such as vision or depth measurements, to adapt movement to terrain. “Recent learning-based controllers can handle diverse, complex terrains using onboard perception”
  • Phase-conditioned reward: A reward whose calculation or activation depends on the current phase of a movement cycle. “phase-conditioned rewards apply each objective where it is physically meaningful”
  • Plantar pressure: Pressure distributed across the sole of the foot. “humans rely on rich plantar pressure feedback”
  • Policy: A mapping from observations or states to actions in a control or reinforcement-learning system. “the policy combines tactile features with proprioception and depth”
  • Privileged observation: Information available during training but intentionally withheld from the deployed policy. “other privileged quantities are used only by the critics and reward computation”
  • Proprioception: Internal sensing of a robot’s body configuration, motion, and joint states. “the policy combines tactile features with proprioception and depth”
  • Proximal Policy Optimization (PPO): A policy-gradient reinforcement-learning algorithm that limits policy updates to improve training stability. “use the Proximal Policy Optimization (PPO)”
  • Raycasting: A computational technique that traces rays through a simulated scene to detect intersections and distances. “Raycasting estimates taxel–terrain gaps and terrain normals for force distribution”
  • Reinforcement learning (RL): A learning paradigm in which an agent learns actions through interaction and reward feedback. “This design is computationally efficient for large-scale RL”
  • Reward shaping: The design of intermediate reward signals to guide an agent toward desired behavior. “tactile reward shaping and sim-to-real alignment”
  • Sparse reward: A reward signal that is provided only at occasional events, milestones, or constraint violations. “sparse terms become informative only at discrete events, gates, or constraint violations”
  • Sim-to-real transfer: The process of transferring a policy or model learned in simulation to a physical robot. “Because of the large sim-to-real gap in high-dimensional raw tactile data”
  • Spatial diffusion: The spreading or smoothing of a signal across neighboring spatial elements. “The tactile model consists of three stages: force distribution, spatial diffusion, and feature extraction”
  • Stance: The phase of locomotion during which a foot remains in contact with and supports the body. “During Stance, the reward encourages broad, centered, and temporally stable support”
  • Support margin: The available stability distance between the effective support point and the boundary of the support region. “a small contact area or an offset center of pressure (CoP) can reduce the effective support margin”
  • Taxel: An individual tactile sensing element in a pressure or force-sensing array. “For each foot, we place M=60M=60 taxels on the sole”
  • Touchdown: The moment when a moving foot first contacts the terrain. “At touchdown, impact transients can induce vibration and generate noise”
  • Whole-body instability: Loss of stable coordinated control involving the robot’s entire body rather than a single joint or foot. “these local contact errors can quickly propagate into whole-body instability”
  • Zero-moment point (ZMP): A point on the support surface at which the net moment from inertial and contact forces is zero, commonly used in bipedal balance analysis. “plantar contact information supports support-region estimation, balance stabilization, and locomotion control under partial or uncertain contacts”

Open Problems

We're still in the process of identifying open problems mentioned in this paper. Please check back in a few minutes.

Tweets

Sign up for free to view the 1 tweet with 176 likes about this paper.