Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
119 tokens/sec
GPT-4o
56 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
6 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Single Shot Scene Text Retrieval (1808.09044v1)

Published 27 Aug 2018 in cs.CV

Abstract: Textual information found in scene images provides high level semantic information about the image and its context and it can be leveraged for better scene understanding. In this paper we address the problem of scene text retrieval: given a text query, the system must return all images containing the queried text. The novelty of the proposed model consists in the usage of a single shot CNN architecture that predicts at the same time bounding boxes and a compact text representation of the words in them. In this way, the text based image retrieval task can be casted as a simple nearest neighbor search of the query text representation over the outputs of the CNN over the entire image database. Our experiments demonstrate that the proposed architecture outperforms previous state-of-the-art while it offers a significant increase in processing speed.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (4)
  1. Marçal Rusiñol (20 papers)
  2. Dimosthenis Karatzas (80 papers)
  3. Lluís Gómez (3 papers)
  4. Andrés Mafla (4 papers)
Citations (46)

Summary

We haven't generated a summary for this paper yet.