Papers
Topics
Authors
Recent
Search
2000 character limit reached

SenseGlove Exoskeleton for Robotic Dexterity

Updated 22 June 2026
  • SenseGlove exoskeleton is a passive, sensorized glove that captures detailed hand motions for dexterous robotic manipulation.
  • It uses distributed Hall-effect sensors and AR-tag tracking to accurately map human joint angles to the Shadow DEX-EE robotic hand.
  • Integration with dynamics filtering, simulation-based RL, and vision-based distillation ensures robust zero-shot transfer of complex manipulation skills.

The SenseGlove exoskeleton, as realized in ExoStart, is a passive, sensorized, low-cost wearable device designed for capturing high-fidelity human hand demonstrations to facilitate learning dexterous robotic manipulation. Specifically tailored to kinematically match the Shadow DEX-EE robotic hand, it enables direct mapping of human intent to robotic execution, driving robust policy learning and zero-shot sim-to-real transfer of complex manipulation skills (Si et al., 13 Jun 2025).

1. Exoskeleton Hardware Architecture and Sensing

The exoskeleton is realized as a fully passive, 3D-printed glove, with each of its three fingers (two opposing one in a tripod configuration) providing four revolute degrees of freedom: metacarpophalangeal (MCP) ab/adduction and flexion, proximal interphalangeal (PIP), and distal interphalangeal (DIP). There are no actuators; all motion originates from human input. The glove architecture is directly scaled to individual users, with a 0.8 scale factor reported for optimal ergonomic fit. Soft silicone pads on the fingertips and middle phalanges offer tactile feedback.

Sensing is accomplished by a distributed array of 12 joint-angle sensors, where each sensor comprises a small ring magnet and a 3 mm Hall-effect sensor; the latter linearly maps magnet field strength to joint angle. The hand pose is tracked using an AR-tagged cubic structure (13.1 mm tags) affixed to the glove’s wrist frame and localized by external cameras. During demonstrations, 3D-printed object replicas are also AR-tagged for pose capture. Five fixed “basket” cameras capture demonstration scenes; for policy execution, two wrist-mounted cameras are utilized. Data acquisition rates are 200 Hz for glove kinematics (via a custom microcontroller) and 10 Hz for robot arm state, hand pose, and object pose, with all data streams synchronized in ROS.

The exoskeleton’s link lengths and joint axes are fabricated to identically match those of the Shadow DEX-EE hand, enabling a direct mapping θi↔qi\theta_i \leftrightarrow q_i (human joint to robot joint), including identical zero positions and joint limits. Calibration for each Hall sensor is performed via:

θi=ki(Vi−Vi0)\theta_i = k_i (V_i - V_i^0)

where ViV_i is the measured Hall voltage, Vi0V_i^0 is the zero-offset, and kik_i is a gain found by traversing the joint’s known range (Si et al., 13 Jun 2025).

2. Demonstration Processing: Dynamics Filtering and Trajectory Generation

Raw exoskeleton-captured demonstrations are subject to both sensor noise and violations of simulated physics constraints (notably penetration ungainly to MuJoCo). To mitigate this, ExoStart applies a short-horizon, sampling-based trajectory optimizer (MJPC) to enforce dynamic feasibility.

Optimization Formulation

At each time tt, the solver optimizes over a control sequence horizon HH:

U∗=argmin⁡U={ut,…,ut+H−1}J(xsim,xdemo)U^* = \operatorname{argmin}_{U = \{u_t, \dots, u_{t+H-1}\}} J(x^{sim}, x^{demo})

subject to the system transition xk+1=f(xk,uk)x_{k+1} = f(x_k, u_k) and joint/torque constraints. The demonstration target xkdemox_k^{demo} is temporally interpolated from raw demonstration streams; θi=ki(Vi−Vi0)\theta_i = k_i (V_i - V_i^0)0 represents the current simulator state. No Denavit–Hartenberg parameter tables are provided; published DEX-EE parameters are used.

Cost Function

The loss θi=ki(Vi−Vi0)\theta_i = k_i (V_i - V_i^0)1 is a weighted linear combination comprising:

  • End-effector position tracking: θi=ki(Vi−Vi0)\theta_i = k_i (V_i - V_i^0)2
  • End-effector orientation tracking: θi=ki(Vi−Vi0)\theta_i = k_i (V_i - V_i^0)3
  • Object pose (position & orientation): θi=ki(Vi−Vi0)\theta_i = k_i (V_i - V_i^0)4
  • Arm joint position error: θi=ki(Vi−Vi0)\theta_i = k_i (V_i - V_i^0)5
  • Fingertip–object keypoint proximity: θi=ki(Vi−Vi0)\theta_i = k_i (V_i - V_i^0)6
  • Finger joint velocity reg.: θi=ki(Vi−Vi0)\theta_i = k_i (V_i - V_i^0)7

The overall stage cost is:

θi=ki(Vi−Vi0)\theta_i = k_i (V_i - V_i^0)8

with typical weights: θi=ki(Vi−Vi0)\theta_i = k_i (V_i - V_i^0)9, ViV_i0, ViV_i1, ViV_i2, ViV_i3, ViV_i4.

Solver Settings

Key parameters: ViV_i5 (50 steps at 200 Hz), 40 control sequences per iteration with 3 spline knots each, exploration noise ViV_i6, and use of the cross-entropy method. Optimization typically runs for 3 hours per demonstration, yielding 25–150 feasible trajectories per manipulation task (Si et al., 13 Jun 2025).

3. Learning Pipeline for Dexterous Manipulation

The policy learning process is staged, beginning with demonstrations processed by the dynamics filter, followed by simulation-based RL, and concluding with student policy distillation for vision-based execution.

Simulation-Based Teacher Training

The state space encompasses arm joints (positions/velocities), end-effector pose (position and rotation), finger joints (positions/velocities/torques), and fingertip poses, each stacked over three timesteps. Actions comprise 6D Cartesian velocities and 12D finger joint targets. Training utilizes a sparse reward: ViV_i7 if task goal is achieved at episode end. Initial states are sampled from both native (uniform random) and demo-corrective sets, with auto-curriculum sampling and zero-variance filtering across four rollouts to select informative initializations. The actor-learner framework is based on IMPALA. Training applies heavy physical domain randomization (varying friction, mass, inertia, damping, and perturbation forces), proceeding for approximately ViV_i8 actor steps until simulation success exceeds 95% for most tasks.

Vision-Based Student Distillation

A vision-conditioned student policy is trained to clone the simulated teacher policy. Observations comprise RGB from five fixed cameras (222×296×3), arm/finger positions, and end-effector pose; actions are identical to the teacher. Behavioral cloning is performed using Action-Chunking Transformers (ACT), with image augmentation (including Gaussian noise, random brightness, contrast, hue, and saturation). Training proceeds for several hundred epochs with a learning rate ViV_i9 and a batch size Vi0V_i^00 until the behavioral cloning loss plateaus (Si et al., 13 Jun 2025).

4. Zero-Shot Transfer: Empirical Results and Ablations

Performance is measured by executing 50 real-world episodes per task. Success rates for representative tasks:

Task Success Rate (%)
Key Lock 56
Nut Unscrew 50
Peg Insertion 54
Box Stand 94
Cube Flip 62
Case Open 56
Bulb Install 2

Direct exoskeleton demonstrations outperform simulated teleoperation both in speed and reliability: e.g., Peg Insertion is completed in 8 s (10/10) with exoskeleton demonstrations, versus 48 s (9/10) with teleoperation. Omitting the dynamics filter severely degrades convergence and data efficiency; Box Stand fails entirely and Peg Insertion requires sixfold more RL updates. For Key Lock, a single demonstration yields 17/50 success; 13 demonstrations yield 28/50. Physical and visual domain randomization, as well as simulated external perturbations, are key to robust policy transfer (Si et al., 13 Jun 2025).

5. Sim-to-Real Transfer and Policy Robustness

Robust zero-shot sim-to-real transfer is realized through extensive physical parameter randomization (object/robot mass, friction, inertia, damping) and aggressive visual randomization (textures, lighting, camera positions). Simulated disturbances (external pushes to the object during grasp) further reinforce policy robustness, enabling performance generalization to unmodeled real-world conditions (Si et al., 13 Jun 2025).

Observed limitations include failures due to simulator fidelity (e.g., Bulb Install achieves only 2% due to MuJoCo’s soft-contact penetration); policies may also exploit demonstration-specific shortcuts unavailable to the robot. Mitigations involve improving simulation contact models and incorporating penalty terms or task constraints. Morphological generalization is limited by exoskeleton–robot kinematic matching; a new hardware design is necessary for each robot hand morphology, though future work may leverage generic gloves and post-hoc retargeting via dynamics filtering. Tactile feedback is limited: the human demonstrator perceives contact, but no force is sensed or digitized; adding force sensors could improve demonstration richness (Si et al., 13 Jun 2025).

6. Significance and Prospective Directions

The SenseGlove exoskeleton, as employed in ExoStart, demonstrates that passive, low-cost, sensorized exoskeletons can enable data-efficient collection of high-quality human dexterity demonstrations. When combined with simulation-based dynamics filtering and RL-based auto-curriculum learning pipelines, such exoskeletons can bootstrap complex, in-hand manipulation skills transferred zero-shot to physical robotic hands.

A plausible implication is that the method’s success in capturing human hand motion without actuators, and direct kinematic mapping to robot hands, supports future translation of human-rich manipulation strategies into robotic platforms. Challenges remain in model fidelity, force feedback capture, and morphological adaptation. Directions for further research include integrating tactile/force sensors, refining simulation contact models, and developing more generic, adaptive hardware for cross-morphology retargeting (Si et al., 13 Jun 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to SenseGlove Exoskeleton.