Papers
Topics
Authors
Recent
Search
2000 character limit reached

XRZero-G0: Pushing the Frontier of Dexterous Robotic Manipulation with Interfaces, Quality and Ratios

Published 14 Apr 2026 in cs.RO | (2604.13001v2)

Abstract: The acquisition of high-quality, action-aligned demonstration data remains a fundamental bottleneck in scaling foundation models for dexterous robot manipulation. Although robot-free human demonstrations (e.g., the UMI paradigm) offer a scalable alternative to traditional teleoperation, current systems are constrained by sub-optimal hardware ergonomics, open-loop workflows, and a lack of systematic data-mixing strategies. To address these limitations, we present XRZero-G0, a hardware-software co-designed system for embodied data collection and policy learning. The system features an ergonomic, virtual reality interface equipped with a top-view camera and dual specialized grippers to directly improve collection efficiency. To ensure dataset reliability, we propose a closed-loop collection, inspection, training, and evaluation pipeline for non-proprioceptive data. This workflow achieves an 85% data validity rate and establishes a transparent mechanism for quality control. Furthermore, we investigate the empirical scaling behaviors and optimal mixing ratios of robot-free data. Extensive experiments indicate that combining a minimal volume of real-robot data with large-scale robot-free data (e.g., a 10:1 ratio) achieves performance comparable to exclusively real-robot datasets, while reducing acquisition costs by a factor of twenty. Utilizing XRZero-G0, we construct a 2,000-hour robot-free dataset that enables zero-shot cross-embodiment transfer to a target physical robot, demonstrating a highly scalable methodology for generalized real-world manipulation.Our project repository: https://github.com/X-Square-Robot/XRZero-G0

Authors (42)

Summary

  • The paper presents a novel system integrating an ergonomic VR interface, closed-loop quality assurance, and data mixing strategies to scale dexterous manipulation tasks.
  • It achieves an 85% data validity rate and reduces real-robot data requirements by up to 20ร— through effective human demonstration decoupling.
  • Experimental results demonstrate robust zero-shot policy transfer across heterogeneous robots and cost-efficient scaling for complex, long-horizon tasks.

XRZero-G0: Scalable, High-Fidelity Data Acquisition for Dexterous Robotic Manipulation

Introduction

XRZero-G0 presents a comprehensive hardware-software co-designed framework targeting the principal limitations that have constrained generalist robotic embodimentโ€”namely, efficient collection of high-quality, action-aligned data, robustness against noisy demonstrations, and economically viable scaling of real-robot and robot-free datasets. The system integrates an untethered VR-based ergonomic interface, a closed-loop quality assurance pipeline, and empirically validated data-mixing strategies, culminating in the G0-Dataset exceeding 2,000 hours and supporting 3,000 manipulation tasks. This infrastructure enables generalist vision-language-action (VLA) policy pre-training and robust zero-shot cross-embodiment transfer, driving cost-efficient advancements in real-world, long-horizon manipulation. Figure 1

Figure 1: The XRZero-G0 system allows scalable, robot-free data collection through a VR wearable, enabling immediate cross-embodiment mapping.

System Architecture

XRZero-G0 is anchored in a tightly integrated hardware and software pipeline. The system comprises a backpack-powered VR interface with an egocentric RGB camera, dual-wrist cameras, and two physically heterogeneous grippersโ€”the H-shaped (press-actuated, macroscopic interaction) and G-shaped (dexterous, finger-driven) devices. This configuration supports robust, drift-resistant spatial tracking and enables ergonomic demonstrations unconstrained by a robotโ€™s kinematic limitations or workspace. Figure 2

Figure 2: XRZero-G0 system architecture featuring an ergonomic, multi-view VR interface and closed-loop data verification.

A real-time edge computing unit synchronizes high-frequency 6-DoF controller trajectories and RGB streams, transmitting spatiotemporally aligned packets to a backend server for immediate ingestion into the quality pipeline. The closed-loop Collection-Inspection-Training-Evaluation pipeline encompasses visual frame cleansing, motion downsampling, kinematically valid retargeting (via URDF/IK), open-loop robot playback, and semantic subtask annotation. This process achieves an 85% data validity rate, systematically eliminating the degradation phenomena observed in prior open-loop, human-centric data regimes. Figure 3

Figure 3: Backpack-powered XRZero-G0 data collection rig delivers rapid, ergonomic acquisition with customized physical grippers.

Dataset Composition and Distribution

The G0-Dataset collected with XRZero-G0 spans 2,000+ hours and captures over 3,000 distinct tasks. Distribution analysis shows a pronounced long-tail: the dataset includes abundant repeatable primitives (e.g., folding, sorting) while extensively covering rare, semantically specialized skills. This compositional structure enhances policy generalization by providing both robust core alignment and task-specific operational diversity. Egocentric viewpoints and physical gripper alternation enrich the dataโ€™s morphology, with semantic bounding boxes granting explicit supervision for multimodal learning. Figure 4

Figure 4: Overview of G0-Dataset with long-tail task distribution and diverse egocentric visual frames, including semantic annotations.

XRZero-G0 attains peak data acquisition rates exceeding 93 episodes/hour, which far outstrip conventional teleoperation and VR-based baselines, consolidating its position as a throughput-optimized pipeline for large-scale embodied datasets.

Experimental Results

Data Collection Efficiency

Empirical benchmarks demonstrate XRZero-G0โ€™s significant improvements in human operator throughput and task usability. Compared to master-slave teleoperation, mean episode duration reduces by 2.33ร— for simple, 1.88ร— for medium, and 1.71ร— for complex tasks. These gains are attributed to decoupling human demonstration from rigid robotic kinematics, reducing cognitive-motor mapping friction and leveraging proprioceptive feedback from physical grippers. Figure 5

Figure 5: XRZero-G0 outperforms traditional teleoperation and VR interfaces in data collection speed across multiple task complexities.

Cross-Embodiment Transfer and Robot-Free Data Scaling

A core result is the empirical validation of pure robot-free data for direct policy inference. Robot-free demonstrations can be kinematically mapped and replayed 1:1 on heterogeneous dual-arm robots, demonstrating functional equivalence to real-robot teleoperation data. Scaling the pure robot-free dataset from 300 to 500 episodes yields a strong linear uplift in grasping task performance (up to 75% success in challenging tasks using Wall-OSS and ฯ€0\pi_0), while scaling further for complex long-horizon dual-arm tasks (up to 2,000 episodes) shows robust generalization across variation in robot height, a setting where fixed-base teleoperation data typically overfits catastrophically. Figure 6

Figure 6: Success rates scale linearly with more robot-free data for both basic and long-horizon dual-arm tasks, substantiating robust spatial generalization.

Data Mixing Laws and Cost-Efficiency

Central to XRZero-G0 is the establishment of optimal data-mixing strategies. Empirical analysis reveals that augmenting a moderate-size real-robot anchor with equivalent robot-free data (1:1 mixture) not only increases the asymptotic performance ceiling but also amplifies sample efficiency in data-abundant regimes. Remarkably, a 10:1 mixtureโ€”replacing 90% of expensive real-robot data with robot-free episodesโ€”matches or is statistically indistinguishable from real-robot-only baselines across key tasks. This substantiates the Few-Shot Physical Anchoring effect: minimal real-robot data provides embodiment-specific kinetic priors after generalized environment and affordance learning from abundant robot-free demonstrations. Figure 7

Figure 7

Figure 7: Data mixing experiments show that 90% cost substitution with robot-free data incurs no performance drop when anchored with a small volume of real-robot trajectories.

Qualitative experiments on structurally disparate robots confirm stable, zero-shot transfer using policies trained with all mixing ratios, reliably executing intricate behaviors such as dexterous manipulation, tool use, and handling of highly deformable objects.

Implications and Future Prospects

XRZero-G0โ€™s results have both practical and theoretical import. Practically, the demonstrated 20ร— reduction in real-robot data requirements alongside robust performance scaling directly enables democratized embodied policy learning for organizations without extensive hardware infrastructure. The closed-loop quality regime ameliorates the recurrent brittleness in prior human-centric datasets, while modular, ergonomic interfaces offer extensibility to future sensor fusion (tactile/auditory) and further miniaturization.

Theoretically, the findings affirm that: (1) Large-scale, morphologically diverse human demonstrationsโ€”when rigorously verifiedโ€”are sufficient to drive direct cross-embodiment policy learning; (2) A small fraction of embodiment-specific anchoring data is adequate for task transfer without catastrophic distribution shift or spatial overfitting; (3) Data scaling laws and mixing ratios can be operationalized for resource-constrained settings.

As embodied policies progress toward end-to-end world modeling and long-horizon cognition, the G0-Dataset structure and pipeline exemplified here provide the fundamental data infrastructure for open-world generalization, robust spatial invariance, and policy pre-training across arbitrary morphologies.

Conclusion

XRZero-G0 advances the field of dexterous robotic manipulation by solving the core data bottlenecks through hardware-software codesign. The system achieves record data throughput, high fidelity (85% validity), ergonomic usability, and empirically validated data mixing strategies that enable massive cost reduction and stable cross-embodiment transfer. This framework authoritatively demonstrates that the era of scalable, cost-efficient, and robust robot-free data for generalist policy learning is operationally attainable, paving the way for broader deployment of embodied intelligence in practical environments.

(2604.13001)

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Collections

Sign up for free to add this paper to one or more collections.