Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
41 tokens/sec
GPT-4o
59 tokens/sec
Gemini 2.5 Pro Pro
41 tokens/sec
o3 Pro
7 tokens/sec
GPT-4.1 Pro
50 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

T2ICount: Enhancing Cross-modal Understanding for Zero-Shot Counting (2502.20625v2)

Published 28 Feb 2025 in cs.CV

Abstract: Zero-shot object counting aims to count instances of arbitrary object categories specified by text descriptions. Existing methods typically rely on vision-LLMs like CLIP, but often exhibit limited sensitivity to text prompts. We present T2ICount, a diffusion-based framework that leverages rich prior knowledge and fine-grained visual understanding from pretrained diffusion models. While one-step denoising ensures efficiency, it leads to weakened text sensitivity. To address this challenge, we propose a Hierarchical Semantic Correction Module that progressively refines text-image feature alignment, and a Representational Regional Coherence Loss that provides reliable supervision signals by leveraging the cross-attention maps extracted from the denosing U-Net. Furthermore, we observe that current benchmarks mainly focus on majority objects in images, potentially masking models' text sensitivity. To address this, we contribute a challenging re-annotated subset of FSC147 for better evaluation of text-guided counting ability. Extensive experiments demonstrate that our method achieves superior performance across different benchmarks. Code is available at https://github.com/cha15yq/T2ICount.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (8)
  1. Yifei Qian (7 papers)
  2. Zhongliang Guo (14 papers)
  3. Bowen Deng (30 papers)
  4. Chun Tong Lei (3 papers)
  5. Shuai Zhao (116 papers)
  6. Chun Pong Lau (26 papers)
  7. Xiaopeng Hong (59 papers)
  8. Michael P. Pound (3 papers)
X Twitter Logo Streamline Icon: https://streamlinehq.com