Papers
Topics
Authors
Recent
Search
2000 character limit reached

Get a Grip: Reconstructing Hand-Object Stable Grasps in Egocentric Videos

Published 25 Dec 2023 in cs.CV | (2312.15719v2)

Abstract: We propose the task of Hand-Object Stable Grasp Reconstruction (HO-SGR), the reconstruction of frames during which the hand is stably holding the object. We first develop the stable grasp definition based on the intuition that the in-contact area between the hand and object should remain stable. By analysing the 3D ARCTIC dataset, we identify stable grasp durations and showcase that objects in stable grasps move within a single degree of freedom (1-DoF). We thereby propose a method to jointly optimise all frames within a stable grasp, minimising object motions to a latent 1-DoF. Finally, we extend the knowledge to in-the-wild videos by labelling 2.4K clips of stable grasps. Our proposed EPIC-Grasps dataset includes 390 object instances of 9 categories, featuring stable grasps from videos of daily interactions in 141 environments. Without 3D ground truth, we use stable contact areas and 2D projection masks to assess the HO-SGR task in the wild. We evaluate relevant methods and our approach preserves significantly higher stable contact area, on both EPIC-Grasps and stable grasp sub-sequences from the ARCTIC dataset.

Authors (2)
Definition Search Book Streamline Icon: https://streamlinehq.com
References (56)
  1. Thor-net: End-to-end graformer-based realistic two hands and object reconstruction with self-supervision. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 1001–1010, 2023.
  2. Contactpose: A dataset of grasps with object contact and hand pose. In Proceedings of the European Conference on Computer Vision (ECCV), pages 361–378, 2020.
  3. The Yale human grasping dataset: Grasp, object, and task data in household and machine shop environments. International Journal of Robotics Research, 34(3):251–255, 2015.
  4. A hand-centric classification of human and robot dexterous manipulation. IEEE Transactions on Haptics, 6(2):129–144, 2013.
  5. Reconstructing hand-object interactions in the wild. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 12417–12426, 2021.
  6. Dexycb: A benchmark for capturing hand grasping of objects. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9044–9053, 2021.
  7. Alignsdf: Pose-aligned signed distance fields for hand-object reconstruction. In Proceedings of the European Conference on Computer Vision (ECCV), pages 231–248, 2022.
  8. Blender Online Community. Blender - a 3D modelling and rendering package. Blender Foundation, Stichting Blender Foundation, Amsterdam, 2018.
  9. M.R. Cutkosky. On grasp choice, grasp models, and the design of hands for manufacturing tasks. IEEE Transactions on Robotics and Automation, 5(3):269–279, 1989.
  10. Scaling egocentric vision: The epic-kitchens dataset. In Proceedings of the European Conference on Computer Vision (ECCV), pages 720–736, 2018.
  11. Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100. International Journal of Computer Vision (IJCV), 130:33–55, 2022.
  12. Epic-kitchens visor benchmark: Video segmentations and object relations. In Proceedings of the Neural Information Processing Systems (NeurIPS) Track on Datasets and Benchmarks, 2022.
  13. ARCTIC: A dataset for dexterous bimanual hand-object manipulation. In Proceedings IEEE Conference on Computer Vision and Pattern Recognition, 2023.
  14. The GRASP Taxonomy of Human Grasp Types. IEEE Transactions on Human-Machine Systems, 46(1):66–77, 2016.
  15. First-person hand action benchmark with RGB-D videos and 3d hand pose annotations. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 409–419, 2018.
  16. Contactopt: Optimizing contact to improve grasps. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1471–1481, 2021.
  17. Honnotate: A method for 3d annotation of hand and object poses. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3193–3203, 2020.
  18. Leveraging photometric consistency over time for sparsely supervised hand-object reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 568–577, 2020.
  19. Towards unconstrained joint hand-object reconstruction from rgb videos. In 2021 International Conference on 3D Vision (3DV), pages 659–668, 2021.
  20. Learning joint reconstruction of hands and manipulated objects. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 11807–11816, 2019.
  21. Epos: Estimating 6d pose of objects with symmetries. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11703–11712, 2020.
  22. Reconstructing Hand-Held Objects from Monocular Video. In Proceedings of SIGGRAPH Asia 2022 Conference Papers, 2022.
  23. Grasping field: Learning implicit representations for human grasps. In 2020 International Conference on 3D Vision (3DV), pages 333–344, 2020.
  24. Neural 3d mesh renderer. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3907–3916, 2018.
  25. Learning analysis-by-synthesis for 6d pose estimation in rgb-d images. In Proceedings of the IEEE international conference on computer vision, pages 954–962, 2015.
  26. H2o: Two hands manipulating objects for first person interaction recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10138–10148, 2021.
  27. Cosypose: Consistent multi-view multi-object 6d pose estimation. In Proceedings of the European Conference on Computer Vision (ECCV), pages 574–591, 2020.
  28. Semi-supervised 3d hand-object poses estimation with interactions in time. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14687–14697, 2021.
  29. Catre: Iterative point clouds alignment for category-level object pose refinement. In Proceedings of the European Conference on Computer Vision (ECCV), pages 499–516, 2022.
  30. Hoi4d: A 4d egocentric dataset for category-level human-object interaction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21013–21022, 2022.
  31. Core50: a new dataset and benchmark for continuous object recognition. In Conference on Robot Learning, pages 17–26, 2017.
  32. An Invitation to 3-D Vision: From Images to Geometric Models. SpringerVerlag, 2003.
  33. Learning to imitate object interactions from internet videos. arXiv preprint arXiv:2211.13225, 2022.
  34. Accelerating 3d deep learning with pytorch3d. arXiv preprint arXiv:2007.08501, 2020.
  35. Understanding everyday hands in action from RGB-D images. In Proceedings of International Conference on Computer Vision, 2015, pages 3889–3897, 2015.
  36. Embodied Hands: Modeling and Capturing Hands and Bodies Together. ACM Trans. Graph, 36:17, 2017.
  37. Frankmocap: A monocular 3d whole-body pose estimation system via regression and integration. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshop, pages 1749–1759, 2021.
  38. Assembly101: A large-scale multi-view video dataset for understanding procedural activities. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21096–21106, 2022.
  39. NodeSLAM: Neural Object Descriptors for Multi-View Shape Reconstruction. In Proceedings of the International Conference on 3D Vision (3DV), 2020.
  40. Showme: Benchmarking object-agnostic hand-object 3d reconstruction. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshop, pages 1935–1944, 2023.
  41. Grab: A dataset of whole-body human grasping of objects. In Proceedings of the European Conference on Computer Vision (ECCV), pages 581–600, 2020.
  42. H+O: unified egocentric recognition of 3d hand-object poses and interactions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4511–4520, 2019.
  43. Collaborative learning for hand and object reconstruction with attention-guided graph convolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1664–1674, 2022.
  44. Occlusion-Aware Self-Supervised Monocular 6D Object Pose Estimation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021.
  45. Normalized object coordinate space for category-level 6d object pose and size estimation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2642–2651, 2019.
  46. Interacting hand-object pose estimation via dense mutual attention. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 5735–5745, 2023.
  47. Posecnn: A convolutional neural network for 6d object pose estimation in cluttered scenes. arXiv preprint arXiv:1711.00199, 2017.
  48. Artiboost: Boosting articulated 3d hand-object pose estimation via online exploration and synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2750–2760, 2022.
  49. Oakink: A large-scale knowledge repository for understanding hand-object interaction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20953–20962, 2022.
  50. Cpf: Learning a contact potential field to model the hand-object interaction. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 11097–11106, 2021.
  51. What’s in your hands? 3d reconstruction of generic objects in hands. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3895–3905, 2022.
  52. Diffusion-guided reconstruction of everyday hand-object interaction clips. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 19717–19728, 2023.
  53. ManipNet: Neural Manipulation Synthesis with a Hand-Object Spatial Representation. ACM Transactions on Graphics, 40(4), 2021.
  54. Perceiving 3d human-object spatial arrangements from a single image in the wild. In Proceedings of the European Conference on Computer Vision (ECCV), pages 34–51, 2020.
  55. Stability-driven contact reconstruction from monocular color images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1643–1653, 2022.
  56. Toch: Spatio-temporal object-to-hand correspondence for motion refinement. In Proceedings of the European Conference on Computer Vision (ECCV), pages 1–19, 2022.
Citations (4)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.