TextBlockV2: Towards Precise-Detection-Free Scene Text Spotting with Pre-trained Language Model (2403.10047v1)

Published 15 Mar 2024 in cs.CV

Abstract: Existing scene text spotters are designed to locate and transcribe texts from images. However, it is challenging for a spotter to achieve precise detection and recognition of scene texts simultaneously. Inspired by the glimpse-focus spotting pipeline of human beings and impressive performances of Pre-trained LLMs (PLMs) on visual tasks, we ask: 1) "Can machines spot texts without precise detection just like human beings?", and if yes, 2) "Is text block another alternative for scene text spotting other than word or character?" To this end, our proposed scene text spotter leverages advanced PLMs to enhance performance without fine-grained detection. Specifically, we first use a simple detector for block-level text detection to obtain rough positional information. Then, we finetune a PLM using a large-scale OCR dataset to achieve accurate recognition. Benefiting from the comprehensive language knowledge gained during the pre-training phase, the PLM-based recognition module effectively handles complex scenarios, including multi-line, reversed, occluded, and incomplete-detection texts. Taking advantage of the fine-tuned LLM on scene recognition benchmarks and the paradigm of text block detection, extensive experiments demonstrate the superior performance of our scene text spotter across multiple public benchmarks. Additionally, we attempt to spot texts directly from an entire scene image to demonstrate the potential of PLMs, even LLMs.

PDF HTML Abstract

Summarize PDF Markdown Bookmark Chat (Pro)

References (75)

Authors (7)

Jiahao Lyu (9 papers)
Jin Wei (16 papers)
Gangyan Zeng (6 papers)
Zeng Li (24 papers)
Enze Xie (84 papers)
Wei Wang (1793 papers)
Yu Zhou (335 papers)

Citations (1)

View on Semantic Scholar

Tweets

https://twitter.com/CSVisionPapers/status/1769879459413229953

TextBlockV2: Towards Precise-Detection-Free Scene Text Spotting with Pre-trained Language Model (2403.10047v1)

Related Papers

Tweets