Efficiently Leveraging Linguistic Priors for Scene Text Spotting
Abstract: Incorporating linguistic knowledge can improve scene text recognition, but it is questionable whether the same holds for scene text spotting, which typically involves text detection and recognition. This paper proposes a method that leverages linguistic knowledge from a large text corpus to replace the traditional one-hot encoding used in auto-regressive scene text spotting and recognition models. This allows the model to capture the relationship between characters in the same word. Additionally, we introduce a technique to generate text distributions that align well with scene text datasets, removing the need for in-domain fine-tuning. As a result, the newly created text distributions are more informative than pure one-hot encoding, leading to improved spotting and recognition performance. Our method is simple and efficient, and it can easily be integrated into existing auto-regressive-based approaches. Experimental results show that our method not only improves recognition accuracy but also enables more accurate localization of words. It significantly improves both state-of-the-art scene text spotting and recognition pipelines, achieving state-of-the-art results on several benchmarks.
- A new method for text detection and recognition in indoor scene for assisting blind people. In International Conference on Machine Vision, 2017.
- A robot object recognition method based on scene text reading in home environments. Sensors (Basel, Switzerland), 21, 2021.
- Recognizing text-based traffic signs. IEEE Transactions on Intelligent Transportation Systems, 16:1360–1369, 2015.
- Dictionary-guided scene text recognition. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 7379–7388, 2021.
- Character region awareness for text detection. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 9357–9366, 2019.
- Optimal boxes: Boosting end-to-end scene text recognition by adjusting annotated bounding boxes via reinforcement learning. In European Conference on Computer Vision, pages 233–248. Springer, 2022.
- Textsnake: A flexible representation for detecting text of arbitrary shapes. In Proceedings of the European conference on computer vision (ECCV), pages 20–36, 2018.
- Deep relational reasoning graph network for arbitrary shape text detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9699–9708, 2020.
- Fourier contour embedding for arbitrary-shaped text detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3123–3131, 2021.
- Real-time scene text detection with differentiable binarization. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pages 11474–11481, 2020a.
- Pyramid mask text detector. arXiv preprint arXiv:1903.11800, 2019a.
- Panet: Few-shot image semantic segmentation with prototype alignment. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9197–9206, 2019a.
- Toward understanding wordart: Corner-guided transformer for scene text recognition. In European Conference on Computer Vision, 2022.
- Scene text recognition with permuted autoregressive sequence models. ArXiv, abs/2207.06966, 2022.
- Multi-granularity prediction for scene text recognition. In European Conference on Computer Vision, 2022.
- Background-insensitive scene text recognition with text semantic segmentation. In European Conference on Computer Vision, 2022.
- Open-set text recognition via character-context decoupling. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 4513–4522, 2022a.
- Towards the unseen: Iterative text recognition by distilling from errors. 2021 IEEE/CVF International Conference on Computer Vision (ICCV), pages 14930–14939, 2021.
- Robust scene text recognition with automatic rectification. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4168–4176, 2016a.
- Gated recurrent convolution neural network for ocr. Advances in Neural Information Processing Systems, 30, 2017.
- Aon: Towards arbitrarily-oriented text recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 5571–5579, 2018.
- Icdar2019 robust reading challenge on arbitrary-shaped text-rrc-art. In 2019 International Conference on Document Analysis and Recognition (ICDAR), pages 1571–1576. IEEE, 2019.
- Total-text: A comprehensive dataset for scene text detection and recognition. 2017 14th IAPR International Conference on Document Analysis and Recognition (ICDAR), 01:935–942, 2017.
- Curved scene text detection via transverse and longitudinal sequence connection. Pattern Recognit., 90:337–345, 2019b.
- Icdar 2015 competition on robust reading. 2015 13th International Conference on Document Analysis and Recognition (ICDAR), pages 1156–1160, 2015.
- Accurate scene text recognition based on recurrent neural network. In Asian conference on computer vision, pages 35–48. Springer, 2014.
- An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition. IEEE transactions on pattern analysis and machine intelligence, 39(11):2298–2304, 2016b.
- Star-net: a spatial attention residue network for scene text recognition. In BMVC, volume 2, page 7, 2016.
- Scatter: selective context attentional scene text recognizer. In proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11962–11972, 2020.
- Char-net: A character-aware neural network for distorted scene text recognition. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 32, 2018.
- Read like humans: Autonomous, bidirectional and iterative language modeling for scene text recognition. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 7094–7103, 2021.
- Levenshtein ocr. ArXiv, abs/2209.03594, 2022.
- From two to one: A new scene text recognizer with visual language modeling network. 2021 IEEE/CVF International Conference on Computer Vision (ICCV), pages 14174–14183, 2021a.
- Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2(7), 2015.
- Training data-efficient image transformers & distillation through attention. In ICML, 2021.
- Canine: Pre-training an efficient tokenization-free encoder for language representation. Transactions of the Association for Computational Linguistics, 10:73–91, 2022.
- Synthetic data for text localisation in natural images. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 2315–2324, 2016.
- Adam: A method for stochastic optimization. CoRR, abs/1412.6980, 2014.
- Scene text recognition using higher order language priors. In British Machine Vision Conference, 2009. URL https://api.semanticscholar.org/CorpusID:9695967.
- End-to-end scene text recognition. 2011 International Conference on Computer Vision, pages 1457–1464, 2011. URL https://api.semanticscholar.org/CorpusID:14136313.
- Icdar 2013 robust reading competition. 2013 12th International Conference on Document Analysis and Recognition, pages 1484–1493, 2013.
- Recognizing text with perspective distortion in natural scenes. 2013 IEEE International Conference on Computer Vision, pages 569–576, 2013. URL https://api.semanticscholar.org/CorpusID:5619635.
- A robust arbitrary text detection system for natural scene images. Expert Syst. Appl., 41:8027–8048, 2014. URL https://api.semanticscholar.org/CorpusID:15559857.
- Icdar2017 robust reading challenge on multi-lingual scene text detection and script identification - rrc-mlt. 2017 14th IAPR International Conference on Document Analysis and Recognition (ICDAR), 01:1454–1459, 2017. URL https://api.semanticscholar.org/CorpusID:4761162.
- Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 8798–8808, 2021. URL https://api.semanticscholar.org/CorpusID:234469662.
- Convolutional character networks. 2019 IEEE/CVF International Conference on Computer Vision (ICCV), pages 9125–9135, 2019.
- Swintextspotter: Scene text spotting via better synergy between text detection and text recognition. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 4583–4593, 2022.
- Pan++: Towards efficient and accurate end-to-end spotting of arbitrarily-shaped text. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44:5349–5367, 2021b.
- Glass: Global to local attention for scene-text spotting. In European Conference on Computer Vision, 2022.
- Towards weakly-supervised text spotting using a multi-task transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4604–4613, 2022.
- Character region attention for text spotting. ArXiv, abs/2007.09629, 2020.
- Spts: single-point text spotting. In Proceedings of the 30th ACM International Conference on Multimedia, pages 4272–4281, 2022.
- Text spotting transformers. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 9509–9518, 2022.
- Abinet++: Autonomous, bidirectional and iterative language modeling for scene text spotting. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45:7123–7141, 2022.
- Mask textspotter v3: Segmentation proposal network for robust scene text spotting. ArXiv, abs/2007.09482, 2020b.
- Abcnet v2: Adaptive bezier-curve network for real-time end-to-end text spotting. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44:8048–8064, 2022b.
- Mango: A mask attention guided one-stage scene text spotter. In AAAI, 2021.
- Deepsolo: Let transformer decoder with explicit points solo for text spotting. ArXiv, abs/2305.19957, 2022.
- Aster: An attentional scene text recognizer with flexible rectification. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41:2035–2048, 2019. URL https://api.semanticscholar.org/CorpusID:206769320.
- Decoupled attention network for text recognition. In AAAI Conference on Artificial Intelligence, 2019b. URL https://api.semanticscholar.org/CorpusID:209444482.
- Robustscanner: Dynamically enhancing positional clues for robust text recognition. In European Conference on Computer Vision, 2020. URL https://api.semanticscholar.org/CorpusID:220525922.
- Show, attend and read: A simple and strong baseline for irregular text recognition. ArXiv, abs/1811.00751, 2018. URL https://api.semanticscholar.org/CorpusID:53301402.
- Seed: Semantics enhanced encoder-decoder framework for scene text recognition. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 13525–13534, 2020a. URL https://api.semanticscholar.org/CorpusID:218862702.
- Visual semantics allow for textual reasoning better in scene text recognition. In AAAI Conference on Artificial Intelligence, 2021. URL https://api.semanticscholar.org/CorpusID:245502387.
- Self-supervised implicit glyph attention for text recognition. 2022. URL https://api.semanticscholar.org/CorpusID:251589365.
- Textdragon: An end-to-end framework for arbitrary shaped text spotting. 2019 IEEE/CVF International Conference on Computer Vision (ICCV), pages 9075–9084, 2019.
- Text perceptron: Towards end-to-end arbitrary-shaped text spotting. In AAAI Conference on Artificial Intelligence, 2020b.
Paper Prompts
Sign up for free to create and run prompts on this paper.