Papers
Topics
Authors
Recent
Search
2000 character limit reached

PixelLink: Detecting Scene Text via Instance Segmentation

Published 4 Jan 2018 in cs.CV | (1801.01315v1)

Abstract: Most state-of-the-art scene text detection algorithms are deep learning based methods that depend on bounding box regression and perform at least two kinds of predictions: text/non-text classification and location regression. Regression plays a key role in the acquisition of bounding boxes in these methods, but it is not indispensable because text/non-text prediction can also be considered as a kind of semantic segmentation that contains full location information in itself. However, text instances in scene images often lie very close to each other, making them very difficult to separate via semantic segmentation. Therefore, instance segmentation is needed to address this problem. In this paper, PixelLink, a novel scene text detection algorithm based on instance segmentation, is proposed. Text instances are first segmented out by linking pixels within the same instance together. Text bounding boxes are then extracted directly from the segmentation result without location regression. Experiments show that, compared with regression-based methods, PixelLink can achieve better or comparable performance on several benchmarks, while requiring many fewer training iterations and less training data.

Authors (4)
Citations (546)

Summary

  • The paper presents PixelLink, a segmentation-based approach that eliminates the need for bounding box regression in scene text detection.
  • It leverages a DNN architecture built on VGG16 for pixel-wise text classification and link prediction, streamlining training with less data and iterations.
  • Experiments on benchmarks like IC15 demonstrate an F-score of 83.7, evidencing state-of-the-art accuracy and enhanced computational efficiency.

The paper “PixelLink: Detecting Scene Text via Instance Segmentation” presents a novel methodology for scene text detection predicated on the paradigm of instance segmentation. Traditional text detection approaches rely heavily on bounding box regression, prompting predictions regarding whether pixels represent text or simply background, in addition to determining precise text locations. This methodology sometimes falters when text instances are spatially proximal. Addressing this challenge, the authors propose PixelLink, which obviates the need for bounding box regression by directly extracting text locations from segmented instances.

Methodology

PixelLink employs a Deep Neural Network (DNN) architecture configured for two distinct pixel-wise predictions: text/non-text classification and link prediction, using a fully convolutional network built upon the VGG16 model. For each pixel, the network predicts whether the pixel belongs to text or not, and simultaneously, for each neighbor, it predicts if they belong to the same text instance (link prediction). Instance segmentation is achieved through the concept of pixel linking, whereby connected components (CCs) are formed by grouping pixels based on their predicted links, subsequently deriving bounding boxes from these CCs.

Results and Evaluation

The experiments conducted delineate the efficacy of PixelLink, showcasing performance on par or superior to prevalent regression-based text detection methods across several benchmarks, including IC13, IC15, and MSRA-TD500. On the IC15 benchmark, PixelLink achieved a F-score of 83.7, surpassing methods like EAST and SegLink under comparable conditions. Notably, PixelLink requires fewer training iterations and less data, courtesy of its reduced requirement for large receptive fields and simplified model learning tasks, marking significant efficiency improvements.

The study highlights that PixelLink's strategy of direct bounding box extraction from instance segmentation circumvents potential prediction inaccuracies associated with regression-based methods. Moreover, the proposal of an Instance-Balanced Cross-Entropy Loss mitigates disparities introduced by varied text instance sizes, ensuring robust training by balancing the importance of each instance.

Implications and Future Directions

PixelLink underscores a paradigm shift toward segmentation-based approaches in text detection, eliciting potential enhancements in accuracy and computational efficiency. It exemplifies how traditional regression tasks can transition into segmentation endeavors, harnessing existing convolutional architectures, such as VGG16, without the necessity of bounding box regression.

Future research could extend PixelLink’s framework to other domains necessitating instance segmentation, suggesting applications in general object detection or even more specialized tasks demanding object differentiation at the pixel level. Potential exploration into different backbone architectures beyond VGG16 holds promise for improving both detection accuracy and operational speed. Additionally, deeper investigations into link prediction granularity and its impact on text detection can yield enhancements in delineating closely situated or intricately shaped text regions.

In conclusion, by reframing the text detection task through instance segmentation, PixelLink not only presents a viable alternative to regression-dependent methodologies but also broadens the scope for employing similar techniques across various vision-based applications.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.