Leveraging complementary teleoperation and in-the-wild data collection

Determine how best to leverage the complementary nature of DITTO’s in-the-wild data collection and teleoperation modalities, including their respective roles in large-scale pretraining and high-quality behavior alignment.

Background

DITTO supports two data-collection modalities: handheld, in-the-wild collection, which provides scalable demonstrations with direct physical interaction, and bilateral teleoperation, which captures robot-native actuator behavior and provides joint-level force feedback. The paper reports preliminary evidence that combining these sources can improve policy performance, particularly for grasp retention and lifting in the test-tube uncapping task.

The authors explicitly identify the unresolved challenge of determining how these modalities should be combined most effectively. They suggest using in-the-wild data for large-scale pretraining and teleoperation data for high-quality behavior alignment, while noting that the co-training requirements at scale—especially for Vision-Language-Action Models—remain to be understood.

References

On the data collection side, how to best leverage the complementary nature of our two modalities remains an open question.

— DITTO: Dexterous Interface for Transparent TeleOperation  (2609.19196 - Palacios et al., 16 Sep 2026) in Section Conclusions and Limitations