Papers
Topics
Authors
Recent
Search
2000 character limit reached

Zone Evaluation: Revealing Spatial Bias in Object Detection

Published 20 Oct 2023 in cs.CV | (2310.13215v2)

Abstract: A fundamental limitation of object detectors is that they suffer from "spatial bias", and in particular perform less satisfactorily when detecting objects near image borders. For a long time, there has been a lack of effective ways to measure and identify spatial bias, and little is known about where it comes from and what degree it is. To this end, we present a new zone evaluation protocol, extending from the traditional evaluation to a more generalized one, which measures the detection performance over zones, yielding a series of Zone Precisions (ZPs). For the first time, we provide numerical results, showing that the object detectors perform quite unevenly across the zones. Surprisingly, the detector's performance in the 96% border zone of the image does not reach the AP value (Average Precision, commonly regarded as the average detection performance in the entire image zone). To better understand spatial bias, a series of heuristic experiments are conducted. Our investigation excludes two intuitive conjectures about spatial bias that the object scale and the absolute positions of objects barely influence the spatial bias. We find that the key lies in the human-imperceptible divergence in data patterns between objects in different zones, thus eventually forming a visible performance gap between the zones. With these findings, we finally discuss a future direction for object detection, namely, spatial disequilibrium problem, aiming at pursuing a balanced detection ability over the entire image zone. By broadly evaluating 10 popular object detectors and 5 detection datasets, we shed light on the spatial bias of object detectors. We hope this work could raise a focus on detection robustness. The source codes, evaluation protocols, and tutorials are publicly available at https://github.com/Zzh-tju/ZoneEval.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (81)
  1. https://www.kaggle.com/datasets/parot99/face-mask-detection-yolo-darknet-format.
  2. https://www.kaggle.com/datasets/eunpyohong/fruit-object-detection.
  3. https://www.kaggle.com/datasets/vodan37/yolo-helmethead/metadata.
  4. Mind the pad–cnns can develop blind spots. In ICLR, 2021.
  5. Why do deep convolutional networks generalize so poorly to small image transformations? Journal of Machine Learning Research, 20(184):1–25, 2019.
  6. Self-driving cars: A survey. Expert Systems with Applications, 165:113816, 2021.
  7. Network dissection: Quantifying interpretability of deep visual representations. In CVPR, 2017.
  8. Cascade R-CNN: Delving into high quality object detection. In CVPR, 2018.
  9. Prime sample attention in object detection. In CVPR, 2020.
  10. End-to-end object detection with transformers. In ECCV, 2020.
  11. Truly shift-invariant convolutional neural networks. In CVPR, 2021.
  12. MMDetection: Open mmlab detection toolbox and benchmark. arXiv preprint arXiv:1906.07155, 2019.
  13. Disentangle your dense object detector. In ACM MM, 2021.
  14. Explaining knowledge distillation by quantifying the knowledge. In CVPR, 2020.
  15. Gaussian YOLOv3: An accurate and fast object detector using localization uncertainty for autonomous driving. In ICCV, pages 502–511, 2019.
  16. Toward spatially unbiased generative models. In ICCV, 2021.
  17. Remix: Rebalanced mixup. In ECCV 2020 Workshops, pages 95–110, 2020.
  18. Feature space augmentation for long-tailed data. In ECCV, 2020.
  19. Class-balanced loss based on effective number of samples. In CVPR, 2019.
  20. Class rectification hard mining for imbalanced deep learning. In ICCV, pages 1851–1860, 2017.
  21. The pascal visual object classes (voc) challenge. International Journal of Computer Vision, 88(2):303–338, 2010.
  22. Accurate, large minibatch sgd: Training imagenet in 1 hour. arXiv preprint arXiv:1706.02677, 2017.
  23. Mask R-CNN. In ICCV, 2017.
  24. Deep residual learning for image recognition. In CVPR, 2016.
  25. Learning deep representation for imbalanced classification. In CVPR, pages 5375–5384, 2016.
  26. Fire detection in video surveillances using convolutional neural networks and wavelet transform. Engineering Applications of Artificial Intelligence, 110:104737, 2022.
  27. Composition loss for counting, density map estimation and localization in dense crowds. In ECCV, pages 532–546, 2018.
  28. Global pooling, more than meets the eye: Position information is encoded channel-wise in cnns. In ICCV, 2021.
  29. Position, padding and predictions: A deeper look at position information in cnns. arXiv preprint arXiv:2101.12322, 2021.
  30. ultralytics/yolov5: v6.2 - YOLOv5 Classification Models, Apple M1, Reproducibility, ClearML and Deci.ai integrations, August 2022.
  31. Decoupling representation and classifier for long-tailed recognition. In ICLR, 2020.
  32. A style-based generator architecture for generative adversarial networks. In CVPR, 2019.
  33. Analyzing and improving the image quality of stylegan. In CVPR, 2020.
  34. Osman Semih Kayhan and Jan C van Gemert. On translation invariance in cnns: Convolutional layers can exploit absolute spatial location. In CVPR, 2020.
  35. M2m: Imbalanced classification via major-to-minor translation. In CVPR, pages 13896–13905, 2020.
  36. The open images dataset v4. International Journal of Computer Vision, 128(7):1956–1981, 2020.
  37. Gradient harmonized single-stage detector. In AAAI, 2019.
  38. A dual weighting label assignment scheme for object detection. In CVPR, 2022.
  39. Generalized Focal Loss: learning qualified and distributed bounding boxes for dense object detection. In NeurIPS, 2020.
  40. Overcoming classifier imbalance for long-tail object detection with balanced group softmax. In CVPR, 2020.
  41. Feature pyramid networks for object detection. In CVPR, 2017.
  42. Focal loss for dense object detection. In ICCV, 2017.
  43. Microsoft coco: Common objects in context. In ECCV, 2014.
  44. Ssd: Single shot multibox detector. In ECCV, 2016.
  45. A convnet for the 2020s. In CVPR, 2022.
  46. Large-scale long-tailed recognition in an open world. In CVPR, pages 2537–2546, 2019.
  47. Improving robustness without sacrificing accuracy with patch gaussian augmentation. In ICML workshop, 2019.
  48. Exploring the limits of weakly supervised pretraining. In ECCV, 2018.
  49. Marco Manfredi and Yu Wang. Shift equivariance in object detection. In ECCV Workshops, 2020.
  50. Imbalance problems in object detection: A review. IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(10):3388–3415, 2020.
  51. Factors in finetuning deep model for object detection with long-tail distribution. In CVPR, 2016.
  52. Libra R-CNN: Towards balanced learning for object detection. In CVPR, 2019.
  53. Yolov3: An incremental improvement. arXiv preprint arXiv:1804.02767, 2018.
  54. Faster R-CNN: Towards real-time object detection with region proposal networks. In NeurIPS, 2015.
  55. Real-time video fire/smoke detection based on cnn in antifire surveillance systems. Journal of Real-Time Image Processing, 18:889–900, 2021.
  56. Objects365: A large-scale, high-quality dataset for object detection. In ICCV, 2019.
  57. Training region-based object detectors with online hard example mining. In CVPR, 2016.
  58. Rethinking counting and localization in crowds: A purely point-based framework. In ICCV, pages 3365–3374, 2021.
  59. Sparse R-CNN: End-to-end object detection with learnable proposals. In CVPR, 2021.
  60. Mitigating the bias of centered objects in common datasets. In ICPR, 2022.
  61. Unbiased look at dataset bias. In CVPR, pages 1521–1528, 2011.
  62. A generalized loss function for crowd counting and localization. In CVPR, pages 1974–1983, 2021.
  63. Region proposal by guided anchoring. In CVPR, 2019.
  64. Adaptive class suppression loss for long-tail object detection. In CVPR, 2021.
  65. Pyramid Vision Transformer: A versatile backbone for dense prediction without convolutions. In ICCV, 2021.
  66. Learning to model the tail. Advances in neural information processing systems, 30, 2017.
  67. Implicit semantic data augmentation for deep networks. In NeurIPS, 2019.
  68. Aggregated residual transformations for deep neural networks. In CVPR, 2017.
  69. Positional encoding as spatial inductive bias in gans. In CVPR, 2021.
  70. Dense label encoding for boundary discontinuity free rotation detection. In CVPR, 2021.
  71. Rethinking rotated object detection with gaussian wasserstein distance loss. In ICML, 2021.
  72. Learning high-precision bounding box for rotated object detection via kullback-leibler divergence. In NIPS, 2021.
  73. Dino: Detr with improved denoising anchor boxes for end-to-end object detection. In ICLR, 2023.
  74. Varifocalnet: An iou-aware dense object detector. In CVPR, 2021.
  75. Examining cnn representations with respect to dataset bias. In AAAI, 2018.
  76. Richard Zhang. Making convolutional networks shift-invariant again. In ICML, 2019.
  77. Bridging the gap between anchor-based and anchor-free detection via adaptive training sample selection. In CVPR, 2020.
  78. Deep long-tailed learning: A survey. arXiv preprint arXiv:2110.04596, 2021.
  79. Single-image crowd counting via multi-column convolutional neural network. In CVPR, pages 589–597, 2016.
  80. Training cost-sensitive neural networks with methods addressing the class imbalance problem. IEEE Transactions on Knowledge and Data Engineering, 18(1):63–77, 2005.
  81. Deformable DETR: Deformable transformers for end-to-end object detection. arXiv preprint arXiv:2010.04159, 2020.
Citations (2)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.