Papers
Topics
Authors
Recent
Search
2000 character limit reached

TactileStep Framework for Humanoid Locomotion

Updated 27 September 2026
  • TactileStep is a framework for regulating foot-terrain interaction in humanoid robots using wireless sole-pressure insoles to enhance stability and reduce impact.
  • The framework applies the principle of feature-level tactile alignment where the proportional, center of pressure, and loaded-area features during locomotion are applied.
  • The framework operates with a four-expert mixture-of-experts encoder with tactile reward components such as movement pre-landing, landing force, and stance support regulation.

TactileStep is a deployable tactile-learning framework for regulating foot–terrain interaction in humanoid locomotion. It equips a 29-degree-of-freedom Unitree G1 humanoid with wireless sole-pressure insoles and trains a policy to use normalized force, center of pressure (CoP), and loaded-area features during locomotion. The framework aligns simulated tactile pressure with the physical insole, recognizes four foot-contact phases, and applies phase-aware rewards for pre-landing compliance, touchdown-force reduction, and stable stance support. In simulation and hardware experiments across flat ground, slopes, stairs, and platforms, TactileStep reduced peak touchdown force by up to 48.8%, peak A-weighted impact noise by up to 30.1 dB, and increased stance contact area by up to 23.8% relative to a strong perceptive baseline (Wang et al., 24 Sep 2026).

1. Problem definition and research context

TactileStep addresses a contact-regulation problem in humanoid locomotion. Perceptive locomotion policies can traverse stairs, slopes, platforms, and other terrains while still producing harsh touchdowns, concentrated edge contacts, unstable stance regions, and excessive vibration or noise. Visual and depth sensing provide terrain geometry before contact, but they do not directly reveal the realized pressure distribution beneath the sole. Proprioception and estimated ground-reaction forces provide indirect information, whereas plantar pressure exposes contact initiation, load growth, support extent, and CoP migration.

The framework is motivated by the distinction between geometric feasibility and contact quality. A foothold that appears feasible from vision may generate a high-impact landing or a narrow, unstable support region once the foot actually contacts the terrain. TactileStep therefore treats foot–terrain interaction as an active control variable. The policy is trained to modify pre-landing motion, absorb touchdown, and establish a broad, centered stance contact.

The central sim-to-real principle is feature-level tactile alignment. Rather than transferring raw pressure maps, the framework extracts quantities available in both simulation and hardware:

  • normalized normal force;
  • CoP coordinates;
  • normalized loaded-area ratio;
  • short tactile histories for both feet.

This design is consistent with broader tactile-locomotion research. A vision-based tactile foot has previously been used to estimate foot and ground inclination for single-leg stabilization, with testing errors of approximately 0.477∘0.477^\circ for foot angle and 0.458∘0.458^\circ for ground angle (Zhang et al., 2021). Other work has demonstrated that foot-mounted vibration sensing can classify ground materials and construct location-linked maps, achieving fused within-user F1 of 96.1%96.1\% and cross-user F1 of 88.5%88.5\% (Ying-Lei et al., 12 Apr 2025). TactileStep differs by using sole-pressure features directly within a humanoid locomotion policy rather than primarily performing terrain recognition or post hoc mapping.

2. Sole-pressure sensing and tactile representation

The physical platform is a Unitree G1 humanoid with 29 actuated degrees of freedom and thin deployable pressure insoles. The wireless insole samples at 25 Hz, while wired acquisition can reach 100 Hz. Manufacturer-provided calibration data are used to calibrate sensor readings to force with an MLP.

For each foot f∈{L,R}f\in\{L,R\}, the high-dimensional insole signal is compressed into:

xtf=[Fˉtf, ptcop,f, Aˉtf],\mathbf{x}^{f}_t = \left[ \bar F^{f}_t,\, \mathbf p^{\mathrm{cop},f}_t,\, \bar A^{f}_t \right],

where Fˉtf\bar F^{f}_t is normalized normal force, ptcop,f=(px,py)\mathbf p^{\mathrm{cop},f}_t=(p_x,p_y) is CoP, and Aˉtf\bar A^{f}_t is the fraction of the sole identified as loaded. The practical actor representation is:

xtacf=[Af, F~f, p~xf, p~yf].\mathbf{x}^{f}_{\mathrm{tac}} = \left[ A^f,\, \tilde F^f,\, \tilde p_x^f,\, \tilde p_y^f \right].

Normal force is normalized by body weight 0.458∘0.458^\circ0:

0.458∘0.458^\circ1

where 0.458∘0.458^\circ2 is the force assigned to taxel 0.458∘0.458^\circ3. The contact-area ratio is clipped to 0.458∘0.458^\circ4, while CoP coordinates are normalized by the positive and negative dimensions of the foot outline and clipped to 0.458∘0.458^\circ5. When 0.458∘0.458^\circ6, CoP is set to zero because its estimate is unreliable at near-zero load.

The actor receives a two-frame tactile history for both feet:

0.458∘0.458^\circ7

tactile dimensions. The representation intentionally omits sensor-specific pressure-map details while retaining locomotion-relevant physical variables. Force represents transmitted load; force growth indicates impact impulsiveness; loaded area indicates support completeness; and CoP indicates whether pressure is centered or edge-concentrated.

The framework does not use raw pressure images as policy inputs. This reduces dependence on taxel layout, insole mounting, manufacturing variation, calibration, drift, and sole compliance. It also enables the same actor-level observation structure in simulation and hardware.

3. Tactile simulation and sim-to-real alignment

TactileStep uses a lightweight pressure approximation in Isaac Sim rather than simulating a deformable sole. Each simulated foot contains 0.458∘0.458^\circ8 virtual taxels distributed across the sole. Raycasts estimate the gap between each taxel and the terrain and the local terrain normal.

For taxel 0.458∘0.458^\circ9, a proximity- and alignment-dependent weight is formed from the foot normal, terrain normal, and taxel–terrain gap. The resultant normal contact force 96.1%96.1\%0 is distributed according to:

96.1%96.1\%1

Because rigid-body contact solvers can generate isolated point contacts, neighboring-taxel diffusion approximates load spreading through a compliant sole:

96.1%96.1\%2

where 96.1%96.1\%3 controls diffusion and 96.1%96.1\%4 is the neighborhood of taxel 96.1%96.1\%5.

The simulated tactile features are then calculated as:

96.1%96.1\%6

96.1%96.1\%7

and

96.1%96.1\%8

The area ratio is the fraction of active taxels, and CoP is the force-weighted average taxel position.

The framework applies tactile domain randomization during training:

  • per-taxel force scaling 96.1%96.1\%9;
  • taxel-coordinate perturbations with standard deviation 88.5%88.5\%0 m and clipping at 88.5%88.5\%1 m;
  • up to one frame of measurement delay, with probability approximately 88.5%88.5\%2–88.5%88.5\%3;
  • nonnegative force clipping.

The simulated actor therefore learns to tolerate calibration, spatial-registration, and timing errors expected from the physical insole. The authors compare a tactile-independent baseline in simulation and hardware to reduce the possibility that differences in learned policy behavior are confused with differences in tactile representation. Simulated and measured traces exhibit comparable force, area, and CoP-margin ranges and trends, although residual discrepancies remain.

The feature-level strategy contrasts with systems that attempt direct raw-map transfer. A tactile foot using a compliant optical skin and dense optical flow, for example, learns contact-relative foot and terrain angles from image deformation rather than using a pressure representation (Zhang et al., 2021). TactileStep instead prioritizes compact force-support descriptors that can be calculated with comparatively inexpensive sole-pressure hardware.

4. Policy architecture and contact-phase recognition

TactileStep formulates locomotion as a partially observable Markov decision process and trains the controller with proximal policy optimization (PPO). The actor receives histories of proprioception, tactile features, and depth observations:

88.5%88.5\%4

The proprioceptive variables include base angular velocity, projected gravity, commanded velocity, joint position and velocity, and the previous action. 88.5%88.5\%5 denotes depth-observation history. The deployed actor does not receive simulator-only quantities such as exact contact impulses or privileged vertical foot velocity.

Two critics receive the actor observation plus privileged information consisting of foot vertical velocity and one-hot contact-phase variables. This asymmetric actor–critic arrangement allows the critics to estimate values using information unavailable to the deployed actor.

The policy outputs 29 target joint positions:

88.5%88.5\%6

A low-level proportional–derivative controller converts these targets into torques:

88.5%88.5\%7

The actor uses a four-expert mixture-of-experts encoder. Actor and critic MLPs have widths 88.5%88.5\%8, while depth observations are processed by a CNN with a 128-dimensional output.

Four-state foot-contact model

A binary swing/stance state is insufficient for separating approach, impact, and settled support. TactileStep assigns each foot one of four states:

88.5%88.5\%9

The estimator uses tactile force f∈{L,R}f\in\{L,R\}0, tactile area f∈{L,R}f\in\{L,R\}1, foot vertical velocity f∈{L,R}f\in\{L,R\}2, foot height f∈{L,R}f\in\{L,R\}3, the previous contact state f∈{L,R}f\in\{L,R\}4, and a remaining landing-window counter f∈{L,R}f\in\{L,R\}5.

Contact is detected according to:

f∈{L,R}f\in\{L,R\}6

Pre-landing is defined by:

f∈{L,R}f\in\{L,R\}7

The phase update uses:

f∈{L,R}f\in\{L,R\}8

The resulting phase is:

f∈{L,R}f\in\{L,R\}9

The phase semantics are operational:

  • Swing: the foot is airborne and not yet near the terrain.
  • PreLanding: the foot descends toward the expected contact region.
  • Landing: new contact has been detected and transient impact is occurring.
  • Stance: contact has settled and the foot should provide stable support.

This separation permits phase-specific regulation rather than applying a single undifferentiated contact penalty throughout locomotion.

5. Phase-aware tactile rewards and optimization

The tactile reward is gated by foot and phase:

xtf=[Fˉtf, ptcop,f, Aˉtf],\mathbf{x}^{f}_t = \left[ \bar F^{f}_t,\, \mathbf p^{\mathrm{cop},f}_t,\, \bar A^{f}_t \right],0

where xtf=[Fˉtf, ptcop,f, Aˉtf],\mathbf{x}^{f}_t = \left[ \bar F^{f}_t,\, \mathbf p^{\mathrm{cop},f}_t,\, \bar A^{f}_t \right],1 is the set of four contact phases.

Pre-landing regulation

During pre-landing, the policy is penalized for excessive downward velocity and acceleration:

xtf=[Fˉtf, ptcop,f, Aˉtf],\mathbf{x}^{f}_t = \left[ \bar F^{f}_t,\, \mathbf p^{\mathrm{cop},f}_t,\, \bar A^{f}_t \right],2

The implementation uses thresholded versions of these terms. The intended behavior is anticipatory compliance: the robot begins moderating descent before touchdown rather than waiting for a large force spike.

Landing-force regulation

During landing, the policy is penalized for both high force and rapid force increase:

xtf=[Fˉtf, ptcop,f, Aˉtf],\mathbf{x}^{f}_t = \left[ \bar F^{f}_t,\, \mathbf p^{\mathrm{cop},f}_t,\, \bar A^{f}_t \right],3

with

xtf=[Fˉtf, ptcop,f, Aˉtf],\mathbf{x}^{f}_t = \left[ \bar F^{f}_t,\, \mathbf p^{\mathrm{cop},f}_t,\, \bar A^{f}_t \right],4

Peak terms operate over the landing window and penalize peak normalized force and peak positive force rate. Dense landing shaping uses:

xtf=[Fˉtf, ptcop,f, Aˉtf],\mathbf{x}^{f}_t = \left[ \bar F^{f}_t,\, \mathbf p^{\mathrm{cop},f}_t,\, \bar A^{f}_t \right],5

xtf=[Fˉtf, ptcop,f, Aˉtf],\mathbf{x}^{f}_t = \left[ \bar F^{f}_t,\, \mathbf p^{\mathrm{cop},f}_t,\, \bar A^{f}_t \right],6

where xtf=[Fˉtf, ptcop,f, Aˉtf],\mathbf{x}^{f}_t = \left[ \bar F^{f}_t,\, \mathbf p^{\mathrm{cop},f}_t,\, \bar A^{f}_t \right],7. A warmup factor gradually introduces these penalties:

xtf=[Fˉtf, ptcop,f, Aˉtf],\mathbf{x}^{f}_t = \left[ \bar F^{f}_t,\, \mathbf p^{\mathrm{cop},f}_t,\, \bar A^{f}_t \right],8

This schedule allows the policy to acquire basic locomotion before strong contact penalties shape touchdown behavior.

Stance-support regulation

During stance, the reward favors broad area, centered CoP, and stable pressure distribution:

xtf=[Fˉtf, ptcop,f, Aˉtf],\mathbf{x}^{f}_t = \left[ \bar F^{f}_t,\, \mathbf p^{\mathrm{cop},f}_t,\, \bar A^{f}_t \right],9

The area term favors broader sole–terrain contact. The CoP-margin term favors pressure away from the support boundary. The CoP-difference term suppresses abrupt pressure migration.

These objectives distinguish a soft touchdown from a safe stance. Reducing force alone could produce underloaded or unstable support; TactileStep instead seeks lower impact followed by broader and more centered contact.

PPO and dual critics

The total reward combines task, regularization, safety, and motion-prior terms:

Fˉtf\bar F^{f}_t0

Dense and sparse reward groups are separated for value estimation. The dense critic handles continuous locomotion feedback, while the sparse critic handles touchdown peaks, contact violations, target completion, and other event-driven signals. The actor uses a mixed advantage with reported weights Fˉtf\bar F^{f}_t1.

The motion prior uses an AMP reward:

Fˉtf\bar F^{f}_t2

where Fˉtf\bar F^{f}_t3 is the discriminator output for the policy motion state.

Training uses PPO with AdamW, adaptive KL control targeting Fˉtf\bar F^{f}_t4, Fˉtf\bar F^{f}_t5, Fˉtf\bar F^{f}_t6, clip range Fˉtf\bar F^{f}_t7, entropy coefficient Fˉtf\bar F^{f}_t8, five epochs per update, and 50,000 iterations.

6. Training, evaluation, results, and limitations

The training pipeline uses Isaac Sim and Isaac Lab with 2,048 parallel agents and 24 rollout steps per environment, producing 49,152 transitions per PPO update. Terrain curricula include rough flat ground, rough standing, stair ascent, stair descent, platform step-up, and platform drop-down.

The principal evaluation metrics are peak landing force, contact-area ratio, CoP margin, and peak A-weighted acoustic level. Simulation evaluation uses 4,096 environments or trials per policy and terrain; hardware experiments use 20 samples per condition.

Simulation support quality

TactileStep produces higher contact-area ratios than both a baseline and an ablation without stable-support rewards:

Terrain TactileStep Without stable-support rewards Baseline
Flat Fˉtf\bar F^{f}_t9 ptcop,f=(px,py)\mathbf p^{\mathrm{cop},f}_t=(p_x,p_y)0 ptcop,f=(px,py)\mathbf p^{\mathrm{cop},f}_t=(p_x,p_y)1
Slope ptcop,f=(px,py)\mathbf p^{\mathrm{cop},f}_t=(p_x,p_y)2 ptcop,f=(px,py)\mathbf p^{\mathrm{cop},f}_t=(p_x,p_y)3 ptcop,f=(px,py)\mathbf p^{\mathrm{cop},f}_t=(p_x,p_y)4
Stair ascent ptcop,f=(px,py)\mathbf p^{\mathrm{cop},f}_t=(p_x,p_y)5 ptcop,f=(px,py)\mathbf p^{\mathrm{cop},f}_t=(p_x,p_y)6 ptcop,f=(px,py)\mathbf p^{\mathrm{cop},f}_t=(p_x,p_y)7
Stair descent ptcop,f=(px,py)\mathbf p^{\mathrm{cop},f}_t=(p_x,p_y)8 ptcop,f=(px,py)\mathbf p^{\mathrm{cop},f}_t=(p_x,p_y)9 Aˉtf\bar A^{f}_t0

CoP margins are also higher, particularly on stairs:

Terrain TactileStep Without stable-support rewards Baseline
Flat Aˉtf\bar A^{f}_t1 mm Aˉtf\bar A^{f}_t2 mm Aˉtf\bar A^{f}_t3 mm
Slope Aˉtf\bar A^{f}_t4 mm Aˉtf\bar A^{f}_t5 mm Aˉtf\bar A^{f}_t6 mm
Stair ascent Aˉtf\bar A^{f}_t7 mm Aˉtf\bar A^{f}_t8 mm Aˉtf\bar A^{f}_t9 mm
Stair descent xtacf=[Af, F~f, p~xf, p~yf].\mathbf{x}^{f}_{\mathrm{tac}} = \left[ A^f,\, \tilde F^f,\, \tilde p_x^f,\, \tilde p_y^f \right].0 mm xtacf=[Af, F~f, p~xf, p~yf].\mathbf{x}^{f}_{\mathrm{tac}} = \left[ A^f,\, \tilde F^f,\, \tilde p_x^f,\, \tilde p_y^f \right].1 mm xtacf=[Af, F~f, p~xf, p~yf].\mathbf{x}^{f}_{\mathrm{tac}} = \left[ A^f,\, \tilde F^f,\, \tilde p_x^f,\, \tilde p_y^f \right].2 mm

Traversal success remains approximately unchanged or improves slightly:

Terrain TactileStep Baseline
Stair up 100% 99.98%
Stair down 100% 99.98%
Platform up 100% 99.93%
Platform down 100% 92.26%
Flat 99.98% 99.19%
Slope 99.63% 99.63%

The improved contact regulation has an energetic cost. Simulated energy consumption is higher for TactileStep on stair ascent, stair descent, platform transitions, flat terrain, and slopes. This is consistent with additional active joint regulation for impact absorption and stance stabilization.

Unitree G1 impact and support results

On hardware, TactileStep reduces impact force relative to the perceptive Hiking baseline:

Terrain Baseline Without tactile observation TactileStep
Stair up xtacf=[Af, F~f, p~xf, p~yf].\mathbf{x}^{f}_{\mathrm{tac}} = \left[ A^f,\, \tilde F^f,\, \tilde p_x^f,\, \tilde p_y^f \right].3 N xtacf=[Af, F~f, p~xf, p~yf].\mathbf{x}^{f}_{\mathrm{tac}} = \left[ A^f,\, \tilde F^f,\, \tilde p_x^f,\, \tilde p_y^f \right].4 N xtacf=[Af, F~f, p~xf, p~yf].\mathbf{x}^{f}_{\mathrm{tac}} = \left[ A^f,\, \tilde F^f,\, \tilde p_x^f,\, \tilde p_y^f \right].5 N
Stair down xtacf=[Af, F~f, p~xf, p~yf].\mathbf{x}^{f}_{\mathrm{tac}} = \left[ A^f,\, \tilde F^f,\, \tilde p_x^f,\, \tilde p_y^f \right].6 N xtacf=[Af, F~f, p~xf, p~yf].\mathbf{x}^{f}_{\mathrm{tac}} = \left[ A^f,\, \tilde F^f,\, \tilde p_x^f,\, \tilde p_y^f \right].7 N xtacf=[Af, F~f, p~xf, p~yf].\mathbf{x}^{f}_{\mathrm{tac}} = \left[ A^f,\, \tilde F^f,\, \tilde p_x^f,\, \tilde p_y^f \right].8 N
Platform up xtacf=[Af, F~f, p~xf, p~yf].\mathbf{x}^{f}_{\mathrm{tac}} = \left[ A^f,\, \tilde F^f,\, \tilde p_x^f,\, \tilde p_y^f \right].9 N 0.458∘0.458^\circ00 N 0.458∘0.458^\circ01 N
Platform down 0.458∘0.458^\circ02 N 0.458∘0.458^\circ03 N 0.458∘0.458^\circ04 N
Flat 0.458∘0.458^\circ05 N 0.458∘0.458^\circ06 N 0.458∘0.458^\circ07 N
Slope 0.458∘0.458^\circ08 N 0.458∘0.458^\circ09 N 0.458∘0.458^\circ10 N

The largest baseline-relative reduction is 48.8% on platform ascent. The no-tactile-observation policy often improves over the baseline, demonstrating that altered training alone contributes to performance, but TactileStep remains better in the reported conditions.

Acoustic impact reductions are substantial. On stair descent, peak A-weighted noise falls from 0.458∘0.458^\circ11 dB for the baseline to 0.458∘0.458^\circ12 dB for TactileStep, a reduction of approximately 30.1 dB. The largest reductions occur on stair ascent, stair descent, and platform ascent.

TactileStep also increases sustained contact area:

Terrain Baseline Without tactile observation TactileStep
Stair up 0.458∘0.458^\circ13 0.458∘0.458^\circ14 0.458∘0.458^\circ15
Stair down 0.458∘0.458^\circ16 0.458∘0.458^\circ17 0.458∘0.458^\circ18
Flat 0.458∘0.458^\circ19 0.458∘0.458^\circ20 0.458∘0.458^\circ21
Slope 0.458∘0.458^\circ22 0.458∘0.458^\circ23 0.458∘0.458^\circ24

The maximum area increase is 23.8% on stair descent. This result indicates that lower impact is not achieved simply by avoiding load; the policy subsequently establishes a broader support region.

Ablation findings

The ablations distinguish the contributions of sensing, rewards, and value estimation:

  • Without tactile observations: performance generally improves over the perceptive baseline but remains inferior to TactileStep for impact noise and stance contact area. This indicates that tactile sensing must remain in the deployed observation loop.
  • Without soft-landing rewards: simulated touchdown impact increases across terrains, showing that tactile observations alone do not produce the full landing benefit.
  • Without stable-support rewards: contact area and CoP margin decrease, demonstrating the role of area, CoP, and CoP-smoothness terms in stance regulation.
  • Single critic: impact increases and support quality generally decreases relative to the dual-critic design.

Measurement and methodological limitations

The wireless insole operates at 25 Hz, although wired acquisition reaches 100 Hz. Hardware validation compared 25-Hz and 100-Hz measurements: on flat ground, impact was 0.458∘0.458^\circ25 N at 100 Hz and 0.458∘0.458^\circ26 N at 25 Hz; on platform drop-down, it was 0.458∘0.458^\circ27 N and 0.458∘0.458^\circ28 N, respectively. These comparisons support the reported measurement procedure but do not establish that 25 Hz captures all possible high-frequency impact phenomena.

The pressure simulator approximates compliance through taxel diffusion rather than deformable-contact mechanics. Deformable, granular, highly compliant, or dynamically changing terrain is not validated. The phase estimator relies on hand-designed thresholds involving force, area, height, and velocity. Long-term insole drift, durability, recalibration, and failure recovery are not systematically evaluated. Faster locomotion beyond the bounded command range is untested, and higher energy consumption indicates a trade-off between contact quality and actuation efficiency.

TactileStep’s principal contribution is therefore a closed-loop, phase-structured contact-regulation architecture rather than a standalone pressure sensor. It aligns simulated and real sole features, exposes those features to the deployed policy, and assigns distinct objectives to pre-landing, touchdown, and stance. Within the tested Unitree G1 scenarios, this combination reduces impact and acoustic emission while improving contact-area coverage and CoP support. The remaining research questions concern dynamic and deformable terrain, higher-speed locomotion, sensor durability, phase-estimator robustness, energy-efficient regulation, and integration with richer sole tactile measurements such as distributed shear, slip, and contact-wrench estimation.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to TactileStep.