Papers
Topics
Authors
Recent
Search
2000 character limit reached

YUBI: Yielding Universal Bidigital Interface for Bimanual Dexterous Manipulation at Scale

Published 8 Jun 2026 in cs.RO and cs.AI | (2606.10244v1)

Abstract: We introduce Yielding Universal Bidigital Interface (YUBI), a finger-aligned gripper designed to enable intuitive, ergonomic, and scalable data collection for bimanual dexterous manipulation. While handheld data collection systems such as Universal Manipulation Interface (UMI) enable affordable data collection, their bulky pistol-grip designs can pose ergonomic and usability challenges for fine-grained, dexterous manipulation tasks. To address this, YUBI presents a distinct design principle: yielding, finger-driven actuation that directly maps human finger movements to gripper jaw motion. Using the YUBI devices, we set up a data collection system with integrated VR-based 6 DoF tracking of the gripper, ensuring high-fidelity trajectory data acquisition. We curate a UMI-based dataset of unprecedented scale: 8,434 hours across 1.20M episodes and 119 tasks. Experiments show that YUBI offers advantages over the UMI gripper in versatility for complex bimanual tasks, dexterity, and operational efficiency. A single policy trained on the YUBI dataset transfers across multiple bimanual robots (UR, Franka, and ELEY) simply by mounting the gripper on each platform, confirming that the collected data are directly executable as policy supervision. We release the gripper hardware, data-collection software, and dataset as one integrated stack, offering the open community a reproducible path to large-scale data acquisition for advancing robotic foundation models.

Summary

  • The paper introduces a 319-gram, finger-aligned handheld gripper and rig-mounted VR tracking system designed to improve dexterity, reduce fatigue, and scale demonstration collection.
  • YUBI produces 8,434 hours of data across 1.20 million episodes, 119 tasks, 179 operators, and seven domains, making it the largest reported UMI-based dataset.
  • A single π₀.₅ policy trained with YUBI data transfers across UR, Franka, and Toyota ELEY robots, achieving task success rates from 45% to 100% while outperforming diffusion policy on challenging tasks.

YUBI (Yielding Universal Bidigital Interface) is a handheld data-collection system for bimanual dexterous manipulation, developed by a consortium including AIRoA, Toyota Motor Corporation, and AIST (2606.10244). The work addresses two persistent bottlenecks in scalable demonstration collection: the ergonomic and dexterity limitations of pistol-grip UMI-style grippers, and the tracking-versus-fatigue trade-off inherent to VR-based pose estimation. The paper's central deliverable is an integrated stack — gripper hardware, collection software, and a dataset of 8434 hours across 1.20M episodes and 119 tasks — which the authors position as the largest UMI-based dataset reported to date.

Motivation and positioning

The authors argue that Vision-Language-Action (VLA) policy quality is bounded by demonstration volume and diversity, and that the open research community lacks access to datasets at the scale available to frontier industrial labs. Existing paradigms each carry structural drawbacks: leader-follower teleoperation is expensive and low-throughput; human video or wearable-capture data suffer from an embodiment gap; and handheld interfaces such as UMI bridge embodiment by having operators manipulate the robot's actual end-effector, but prior designs impose their own constraints.

Two specific deficiencies motivate YUBI's design. First, pistol-grip UMI devices place the operator's fingers mechanically offset from the gripper pinch point, degrading haptic transparency and fine motor control; Fin-Ray-type compliant fingertips further exhibit poor positional repeatability and insufficient force for payloads at or above 2 kg. Second, tracking presents a trade-off: SLAM-based wrist tracking drifts under fast motion or textureless scenes, while VR-based systems such as ActiveUMI and exUMI improve 6 DoF fidelity but require wearing the headset, inducing neck fatigue that limits sessions to roughly 30 minutes per manufacturer guidance.

Gripper design

YUBI replaces the pistol grip with yielding, finger-driven actuation: one jaw is actuated by the thumb and the opposing jaw by the coordinated index and middle fingers, so the aperture directly follows the operator's natural pinch without motor resistance or gear backlash. A support grip acts as a mechanical fulcrum, distributing load across the hand and enabling a design payload near 2 kg while retaining access to confined spaces via minimized moment arms. The handheld mass is reduced to approximately 319 g (200 g gripper plus 119 g Quest controller), compared with roughly 780 g for original UMI and over 900 g for VR-integrated variants — a reduction the authors link directly to lower wrist fatigue and less trajectory noise during long sessions.

A secondary design consequence concerns transferability: because the support grip sits below the jaws, operators naturally adopt overhead approaches rather than lateral sweeps near table level, producing trajectories that remain above the workspace and are more safely reproducible by robot arms. Each unit is fabricable from 3D-printed parts for under $200 USD (excluding the Quest 3S).

Operation setup

The stationary configuration places two YUBI devices on a desk rig integrating a rig-mounted Meta Quest 3S for 6 DoF controller tracking (fusing IR LED constellation observations with onboard IMU), a fixed RealSense D435 top-view stereo camera at 30 Hz used for post-hoc quality filtering and annotation, a task UI laptop, and a foot pedal for hands-free action segmentation. Mounting the HMD on the rig rather than the operator's head removes neck strain while preserving tracking coverage. Wrist cameras, controller poses, and magnetic-encoder jaw angles stream at 100, 80, and 100 Hz respectively over ROS 2 with source-side timestamping.

A portable mode detaches the system from the desk: the HMD is chest-mounted, an egocentric fisheye camera replaces the top view, sub-action boundaries are triggered by double-clicking the gripper, and the laptop hub is carried in a shoulder bag. This enables whole-body tasks such as dishwasher loading and shelf placement, with the data schema preserved across modes.

Dataset

Collection ran 24/7 over two months on 22 desks with 179 operators (125 male, 54 female), yielding 8434 hours, 1.20M episodes, 6.80M video-language-action triplets (via per-episode sub-action annotation averaging 7.99 actions per task), and 119 tasks spanning seven domains (industrial, kitchen, toy, desk work, clothing, appliance, personal care). Episodes average 39.8 s. Relative to prior UMI-style releases, the scale difference is substantial: FastUMI-100K provides 92.8K demonstrations and ~600 hours, whereas YUBI offers roughly 73× more demonstrations and 14× more hours; against DexWild (9.5K demos), the ratio approaches 720×. All trajectories are expressed in a shared table frame via ChArUco-board extrinsic calibration, converted to LeRobot format at 30 Hz, and passed through a filtering cascade removing short episodes, frozen-pose/aperture signals, kinematically implausible jumps, and frames flagged as degraded by the Quest tracker's occlusion indicator.

Usability study

A gender-balanced group of 10 novice operators compared YUBI against original UMI. In a single-attempt pick-and-place dexterity test over hex nuts M10–M3, both devices saturate on large nuts (≥94% at M8–M10), but diverge sharply at small sizes: YUBI leads by +20 pp at M6 and +10 pp at M5, and achieves 44% versus UMI's 14% on M3 — roughly a 3× improvement. An anomalous dip at M4 for YUBI is attributed to a size-specific geometric mismatch between nut diameter and fingertip curvature rather than a precision loss, supported by a similar dip pattern in UMI. In an efficiency test across five tasks, YUBI was consistently faster than UMI, with speed-ups from 1.37× (domino arrangement) to 4.19× (phone charging), narrowing the gap to direct hand operation. These results support the claim that finger-aligned actuation improves both precision and throughput, though the sample of 10 operators limits statistical generalization.

Policy deployment

The authors fine-tune π₀.₅ on YUBI wrist-camera data with relative end-effector trajectory action chunks, conditioned on language instructions, executing at a downsampled 10 Hz control rate with per-robot IK conversion. Because supervision lives in end-effector space, a single policy transfers without retargeting across three kinematically distinct bimanual platforms — UR, Franka, and Toyota's semi-humanoid ELEY — each fitted with a motorized YUBI gripper whose flange alignment places the pinch point on the wrist rotation axis. Over 20 rollouts per task, success rates were 20/20 (ball in basket, UR), 13/20 (stack cup pyramid, UR), 9/20 (unfold glasses, UR), 18/20 (pick-and-place socks, Franka), 18/20 (tape in box, Franka), and 11/20 (cup placement, ELEY).

An ablation against a diffusion-policy baseline (frozen CLIP encoders, from-scratch diffusion decoder) shows both saturating on ball in basket, but DP failing categorically on unfold glasses (0/20 vs 9/20) and trailing on stack cup pyramid (9/20 vs 13/20), motivating the choice of a pretrained VLA backbone for complex bimanual contact-rich tasks. Notably, deployment uses only wrist cameras and trajectories, not the top-view stream.

Limitations and open questions

The paper concedes several constraints explicitly. Sub-millimeter precision and tactile-sensitive tasks — tight cable insertion, fragile material handling — remain out of reach, and addressing them would require dedicated curation, multi-modal sensing, and task-specific post-training; YUBI itself carries no tactile sensing, unlike exUMI or Touch-in-the-Wild. The optimal recipe for combining YUBI data with complementary sources (in-the-wild demonstrations, real robot data) is left unresolved. Most significantly, the full 8434 hours have not yet been used for large-scale VLA pretraining — the deployment experiments rely on comparatively small per-task fine-tuning sets (194–3985 demonstrations), so the dataset's value as pretraining corpus remains unvalidated. Deployment success rates also show headroom: two of six tasks sit at or below 55%. Finally, portable-mode whole-body trajectories are proposed as a data source for mobile manipulators and humanoids, but no policy experiments on those embodiments are reported.

Conclusion

YUBI contributes a finger-aligned, yielding gripper, a fatigue-mitigated VR tracking setup, and an open ecosystem whose dataset substantially exceeds all prior UMI-style releases in hours, episodes, and task diversity. Cross-platform deployment of a single π₀.₅ policy validates that end-effector-space supervision collected with the device transfers across kinematically distinct bimanual robots sharing the YUBI end-effector. The principal open question is whether this scale of handheld-collected data yields commensurate gains when used for large-scale VLA pretraining rather than per-task fine-tuning.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 2 tweets with 4 likes about this paper.