Papers
Topics
Authors
Recent
Search
2000 character limit reached

TanDepth: Leveraging Global DEMs for Metric Monocular Depth Estimation in UAVs

Published 8 Sep 2024 in cs.CV | (2409.05142v2)

Abstract: Aerial scene understanding systems face stringent payload restrictions and must often rely on monocular depth estimation for modeling scene geometry, which is an inherently ill-posed problem. Moreover, obtaining accurate ground truth data required by learning-based methods raises significant additional challenges in the aerial domain. Self-supervised approaches can bypass this problem, at the cost of providing only up-to-scale results. Similarly, recent supervised solutions which make good progress towards zero-shot generalization also provide only relative depth values. This work presents TanDepth, a practical scale recovery method for obtaining metric depth results from relative estimations at inference-time, irrespective of the type of model generating them. Tailored for Unmanned Aerial Vehicle (UAV) applications, our method leverages sparse measurements from Global Digital Elevation Models (GDEM) by projecting them to the camera view using extrinsic and intrinsic information. An adaptation to the Cloth Simulation Filter is presented, which allows selecting ground points from the estimated depth map to then correlate with the projected reference points. We evaluate and compare our method against alternate scaling methods adapted for UAVs, on a variety of real-world scenes. Considering the limited availability of data for this domain, we construct and release a comprehensive, depth-focused extension to the popular UAVid dataset to further research.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (49)
  1. F. Nex, C. Armenakis, M. Cramer, D. A. Cucci, M. Gerke, E. Honkavaara, A. Kukko, C. Persello, and J. Skaloud, “UAV in the advent of the twenties: Where we stand and what is next,” ISPRS journal of photogrammetry and remote sensing, vol. 184, pp. 215–242, 2022.
  2. M. A. Akhloufi, A. Couturier, and N. A. Castro, “Unmanned aerial vehicles for wildland fires: Sensing, perception, cooperation and assistance,” Drones, vol. 5, no. 1, p. 15, 2021.
  3. M. Lyu, Y. Zhao, C. Huang, and H. Huang, “Unmanned aerial vehicles for search and rescue: A survey,” Remote Sensing, vol. 15, no. 13, p. 3266, 2023.
  4. B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng, “Nerf: Representing scenes as neural radiance fields for view synthesis,” Communications of the ACM, vol. 65, no. 1, pp. 99–106, 2021.
  5. R. Ranftl, K. Lasinger, D. Hafner, K. Schindler, and V. Koltun, “Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer,” IEEE transactions on pattern analysis and machine intelligence, vol. 44, no. 3, pp. 1623–1637, 2020.
  6. L. Yang, B. Kang, Z. Huang, X. Xu, J. Feng, and H. Zhao, “Depth anything: Unleashing the power of large-scale unlabeled data,” arXiv preprint arXiv:2401.10891, 2024.
  7. L. Yang, B. Kang, Z. Huang, Z. Zhao, X. Xu, J. Feng, and H. Zhao, “Depth anything V2,” 2024. [Online]. Available: https://arxiv.org/abs/2406.09414
  8. H. Florea and S. Nedevschi, “Survey on monocular depth estimation for unmanned aerial vehicles using deep learning,” in 2022 IEEE 18th International Conference on Intelligent Computer Communication and Processing (ICCP).   IEEE, 2022, pp. 319–326.
  9. T. Zhou, M. Brown, N. Snavely, and D. G. Lowe, “Unsupervised learning of depth and ego-motion from video,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 1851–1858.
  10. P. Rizzoli, M. Martone, C. Gonzalez, C. Wecklich, D. B. Tridon, B. Bräutigam, M. Bachmann, D. Schulze, T. Fritz, M. Huber et al., “Generation and performance assessment of the global TanDEM-X digital elevation model,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 132, pp. 119–139, 2017.
  11. W. Zhang, J. Qi, P. Wan, H. Wang, D. Xie, X. Wang, and G. Yan, “An easy-to-use airborne LiDAR data filtering method based on cloth simulation,” Remote sensing, vol. 8, no. 6, p. 501, 2016.
  12. Y. Lyu, G. Vosselman, G.-S. Xia, A. Yilmaz, and M. Y. Yang, “UAVid: A semantic segmentation dataset for UAV imagery,” ISPRS journal of photogrammetry and remote sensing, vol. 165, pp. 108–119, 2020.
  13. V. Licăret, V. Robu, A. Marcu, D. Costea, E. Sluşanschi, R. Sukthankar, and M. Leordeanu, “Ufo depth: Unsupervised learning with flow-based odometry optimization for metric depth estimation,” in 2022 International Conference on Robotics and Automation (ICRA).   IEEE, 2022, pp. 6526–6532.
  14. H. Florea, V.-C. Miclea, and S. Nedevschi, “WildUAV: Monocular UAV dataset for depth estimation tasks,” in 2021 IEEE 17th International Conference on Intelligent Computer Communication and Processing (ICCP).   IEEE, 2021, pp. 291–298.
  15. R. Ranftl, A. Bochkovskiy, and V. Koltun, “Vision transformers for dense prediction,” in Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 12 179–12 188.
  16. V. Guizilini, I. Vasiljevic, D. Chen, R. Ambru\textcommabelows, and A. Gaidon, “Towards zero-shot scale-aware monocular depth estimation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 9233–9243.
  17. L. Piccinelli, Y.-H. Yang, C. Sakaridis, M. Segu, S. Li, L. Van Gool, and F. Yu, “UniDepth: Universal monocular metric depth estimation,” arXiv preprint arXiv:2403.18913, 2024.
  18. H. Fu, M. Gong, C. Wang, K. Batmanghelich, and D. Tao, “Deep ordinal regression network for monocular depth estimation,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 2002–2011.
  19. S. F. Bhat, I. Alhashim, and P. Wonka, “Adabins: Depth estimation using adaptive bins,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 4009–4018.
  20. ——, “Localbins: Improving depth estimation by learning local distributions,” in European Conference on Computer Vision.   Springer, 2022, pp. 480–496.
  21. S. F. Bhat, R. Birkl, D. Wofk, P. Wonka, and M. Müller, “Zoedepth: Zero-shot transfer by combining relative and metric depth,” arXiv preprint arXiv:2302.12288, 2023.
  22. V.-C. Miclea and S. Nedevschi, “Monocular depth estimation with improved long-range accuracy for UAV environment perception,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–15, 2021.
  23. ——, “Dynamic semantically guided monocular depth estimation for uav environment perception,” IEEE Transactions on Geoscience and Remote Sensing, 2023.
  24. L. Madhuanand, F. Nex, M. Yang, N. Paparoditis, C. Mallet, F. Lafarge, F. Remondino, I. Toschi, T. Fuse et al., “Deep learning for monocular depth estimation from UAV images,” ISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Sciences, vol. 2, 2020.
  25. M. Fonder, D. Ernst, and M. Van Droogenbroeck, “M4depth: A motion-based approach for monocular depth estimation on video sequences,” arXiv preprint arXiv:2105.09847, 2021.
  26. R. Garg, V. K. Bg, G. Carneiro, and I. Reid, “Unsupervised cnn for single view depth estimation: Geometry to the rescue,” in European conference on computer vision.   Springer, 2016, pp. 740–756.
  27. C. Godard, O. Mac Aodha, M. Firman, and G. J. Brostow, “Digging into self-supervised monocular depth estimation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 3828–3838.
  28. J.-W. Bian, H. Zhan, N. Wang, Z. Li, L. Zhang, C. Shen, M.-M. Cheng, and I. Reid, “Unsupervised scale-consistent depth learning from video,” International Journal of Computer Vision, pp. 1–17, 2021.
  29. L. Madhuanand, F. Nex, and M. Y. Yang, “Self-supervised monocular depth estimation from oblique UAV videos,” ISPRS journal of photogrammetry and remote sensing, vol. 176, pp. 1–14, 2021.
  30. M. Hermann, B. Ruf, and M. Weinmann, “Real-time dense 3D reconstruction from monocular video data captured by low-cost UAVs,” arXiv preprint arXiv:2104.10515, 2021.
  31. F. Xue, G. Zhuo, Z. Huang, W. Fu, Z. Wu, and M. H. Ang, “Toward hierarchical self-supervised monocular absolute depth estimation for autonomous driving applications,” in 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS).   IEEE, 2020, pp. 2330–2337.
  32. B. Wagstaff and J. Kelly, “Self-supervised scale recovery for monocular depth and egomotion estimation,” arXiv preprint arXiv:2009.03787, 2020.
  33. H. Li, Y. Ma, Y. Gu, K. Hu, Y. Liu, and X. Zuo, “RadarCam-Depth: Radar-camera fusion for depth estimation with learned metric scale,” arXiv preprint arXiv:2401.04325, 2024.
  34. V. Guizilini, R. Ambrus, S. Pillai, A. Raventos, and A. Gaidon, “3d packing for self-supervised monocular depth estimation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 2485–2494.
  35. K. Swami, A. Muduli, U. Gurram, and P. Bajpai, “Do what you can, with what you have: Scale-aware and high quality monocular depth estimation without real world labels,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 988–997.
  36. M. Pirvu, V. Robu, V. Licaret, D. Costea, A. Marcu, E. Slusanschi, R. Sukthankar, and M. Leordeanu, “Depth distillation: unsupervised metric depth estimation for UAVs by finding consensus between kinematics, optical flow and deep learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 3215–3223.
  37. Y. Pan, B. Liu, Z. Liu, H. Shen, J. Xu, W. Fu, and T. Yang, “MoNA Bench: A benchmark for monocular depth estimation in navigation of autonomous unmanned aircraft system,” Drones, vol. 8, no. 2, 2024. [Online]. Available: https://www.mdpi.com/2504-446X/8/2/66
  38. T. G. Farr and M. Kobrick, “Shuttle Radar Topography Mission produces a wealth of data,” Eos, Transactions American Geophysical Union, vol. 81, no. 48, pp. 583–585, 2000.
  39. T. Tachikawa, M. Kaku, A. Iwasaki, D. B. Gesch, M. J. Oimoen, Z. Zhang, J. J. Danielson, T. Krieger, B. Curtis, J. Haase et al., “ASTER global digital elevation model version 2-summary of validation results,” NASA, Tech. Rep., 2011.
  40. T. Tadono, H. Ishida, F. Oda, S. Naito, K. Minakawa, and H. Iwamoto, “Precise global DEM generation by ALOS PRISM,” ISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Sciences, vol. 2, pp. 71–76, 2014.
  41. S. Shah, D. Dey, C. Lovett, and A. Kapoor, “Airsim: High-fidelity visual and physical simulation for autonomous vehicles,” in Field and service robotics.   Springer, 2018, pp. 621–635.
  42. M. Fonder and M. Van Droogenbroeck, “Mid-air: A multi-modal dataset for extremely low altitude drone flights,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops, 2019, pp. 0–0.
  43. W. Wang, D. Zhu, X. Wang, Y. Hu, Y. Qiu, C. Wang, Y. Hu, A. Kapoor, and S. Scherer, “TartanAir: A dataset to push the limits of visual SLAM,” in 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2020.
  44. M. Hermann, M. Weinmann, F. Nex, E. Stathopoulou, F. Remondino, B. Jutzi, and B. Ruf, “Depth estimation and 3D reconstruction from UAV-borne imagery: Evaluation on the UseGeo dataset,” ISPRS Open Journal of Photogrammetry and Remote Sensing, p. 100065, 2024.
  45. J. Bian, Z. Li, N. Wang, H. Zhan, C. Shen, M.-M. Cheng, and I. Reid, “Unsupervised scale-consistent depth and ego-motion learning from monocular video,” Advances in neural information processing systems, vol. 32, pp. 35–45, 2019.
  46. S. Katz, A. Tal, and R. Basri, “Direct visibility of point sets,” in ACM SIGGRAPH 2007 papers, 2007, pp. 24–es.
  47. B. Wessel, M. Huber, C. Wohlfart, U. Marschalk, D. Kosmann, and A. Roth, “Accuracy assessment of the global TanDEM-X digital elevation model with GPS data,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 139, pp. 171–182, 2018.
  48. J. L. Schönberger and J.-M. Frahm, “Structure-from-Motion Revisited,” in Conference on Computer Vision and Pattern Recognition (CVPR), 2016.
  49. J. L. Schönberger, E. Zheng, M. Pollefeys, and J.-M. Frahm, “Pixelwise View Selection for Unstructured Multi-View Stereo,” in European Conference on Computer Vision (ECCV), 2016.
Citations (1)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.