Papers
Topics
Authors
Recent
Search
2000 character limit reached

Expand-and-Quantize: Unsupervised Semantic Segmentation Using High-Dimensional Space and Product Quantization

Published 12 Dec 2023 in cs.CV | (2312.07342v1)

Abstract: Unsupervised semantic segmentation (USS) aims to discover and recognize meaningful categories without any labels. For a successful USS, two key abilities are required: 1) information compression and 2) clustering capability. Previous methods have relied on feature dimension reduction for information compression, however, this approach may hinder the process of clustering. In this paper, we propose a novel USS framework called Expand-and-Quantize Unsupervised Semantic Segmentation (EQUSS), which combines the benefits of high-dimensional spaces for better clustering and product quantization for effective information compression. Our extensive experiments demonstrate that EQUSS achieves state-of-the-art results on three standard benchmarks. In addition, we analyze the entropy of USS features, which is the first step towards understanding USS from the perspective of information theory.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (48)
  1. Multi-objects detection and segmentation for scene understanding based on Texton forest and kernel sliding perceptron. Journal of Electrical Engineering & Technology, 16: 1143–1150.
  2. Deep semantic segmentation of natural and medical images: a review. Artificial Intelligence Review, 54: 137–178.
  3. Estimating or propagating gradients through stochastic neurons for conditional computation. arXiv preprint arXiv:1308.3432.
  4. Coco-stuff: Thing and stuff classes in context. In Proceedings of the IEEE conference on computer vision and pattern recognition, 1209–1218.
  5. Deep clustering for unsupervised learning of visual features. In Proceedings of the European conference on computer vision (ECCV), 132–149.
  6. Emerging properties in self-supervised vision transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 9650–9660.
  7. A Vector Quantized Approach for Text to Speech Synthesis on Real-World Spontaneous Speech. arXiv preprint arXiv:2302.04215.
  8. Improved baselines with momentum contrastive learning. arXiv preprint arXiv:2003.04297.
  9. Picie: Unsupervised semantic segmentation using invariance and equivariance in clustering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 16794–16804.
  10. The cityscapes dataset for semantic urban scene understanding. In Proceedings of the IEEE conference on computer vision and pattern recognition, 3213–3223.
  11. Cover, T. M. 1965. Geometrical and statistical properties of systems of linear inequalities with applications in pattern recognition. IEEE transactions on electronic computers, (3): 326–334.
  12. Damerau, F. J. 1964. A technique for computer detection and correction of spelling errors. Communications of the ACM, 7(3): 171–176.
  13. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, 248–255. Ieee.
  14. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929.
  15. Deep multi-modal object detection and semantic segmentation for autonomous driving: Datasets, methods, and challenges. IEEE Transactions on Intelligent Transportation Systems, 22(3): 1341–1360.
  16. Understanding the difficulty of training deep feedforward neural networks. In Proceedings of the thirteenth international conference on artificial intelligence and statistics, 249–256. JMLR Workshop and Conference Proceedings.
  17. The information bottleneck problem and its applications in machine learning. IEEE Journal on Selected Areas in Information Theory, 1(1): 19–38.
  18. Towards High-Quality Neural TTS for Low-Resource Languages by Learning Compact Speech Representations. arXiv preprint arXiv:2210.15131.
  19. Unsupervised Semantic Segmentation by Distilling Feature Correspondences. In International Conference on Learning Representations.
  20. Unetr: Transformers for 3d medical image segmentation. In Proceedings of the IEEE/CVF winter conference on applications of computer vision, 574–584.
  21. Deep residual learning for image recognition. CoRR, abs/1512, 3385: 2.
  22. Kernel methods in machine learning.
  23. LoRA: Low-Rank Adaptation of Large Language Models. In International Conference on Learning Representations.
  24. EPQuant: A Graph Neural Network compression approach based on product quantization. Neurocomputing, 503: 49–61.
  25. Learning visual groups from co-occurrences in space and time. arXiv preprint arXiv:1511.06811.
  26. Self-supervised product quantization for deep unsupervised image retrieval. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 12085–12094.
  27. Product quantization for nearest neighbor search. IEEE transactions on pattern analysis and machine intelligence, 33(1): 117–128.
  28. Aggregating local descriptors into a compact image representation. In 2010 IEEE computer society conference on computer vision and pattern recognition, 3304–3311. IEEE.
  29. Invariant information clustering for unsupervised image classification and segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 9865–9874.
  30. Semantic feature extraction for generalized zero-shot learning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, 1166–1173.
  31. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980.
  32. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114.
  33. End-to-end supervised product quantization for image search and retrieval. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 5041–5050.
  34. Efficient inference in fully connected crfs with gaussian edge potentials. Advances in neural information processing systems, 24.
  35. Compressing unknown images with product quantizer for efficient zero-shot classification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 5463–5472.
  36. ACSeg: Adaptive Conceptualization for Unsupervised Semantic Segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 7162–7172.
  37. Lowe, D. G. 1999. Object recognition from local scale-invariant features. In Proceedings of the seventh IEEE international conference on computer vision, volume 2, 1150–1157. Ieee.
  38. Meletis, P. 2022. Towards holistic scene understanding: Semantic segmentation and beyond. arXiv preprint arXiv:2201.07734.
  39. Autoregressive unsupervised image segmentation. In European Conference on Computer Vision, 142–158. Springer.
  40. Grad-cam: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE international conference on computer vision, 618–626.
  41. Leveraging Hidden Positives for Unsupervised Semantic Segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 19540–19549.
  42. Shannon, C. E. 1948. A mathematical theory of communication. The Bell system technical journal, 27(3): 379–423.
  43. A comparative study of real-time semantic segmentation for autonomous driving. In Proceedings of the IEEE conference on computer vision and pattern recognition workshops, 587–597.
  44. Training with Quantization Noise for Extreme Model Compression. In International Conference on Learning Representations.
  45. The information bottleneck method. arXiv preprint physics/0004057.
  46. TransFGU: a top-down approach to fine-grained unsupervised semantic segmentation. In European Conference on Computer Vision, 73–89. Springer.
  47. Neural network language model compression with product quantization and soft binarization. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 28: 2438–2449.
  48. On compressing deep models by low rank and sparse decomposition. In Proceedings of the IEEE conference on computer vision and pattern recognition, 7370–7379.
Citations (1)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.