跳到主要导航 跳到搜索 跳到主要内容

TextRSR: Enhanced Arbitrary-Shaped Scene Text Representation via Robust Subspace Recovery

  • Zhiwen Shao
  • , Shengtian Jiang*
  • , Hancheng Zhu*
  • , Xuehuai Shi*
  • , Canlin Li
  • , Lizhuang Ma
  • , Dit Yan Yeung
  • *此作品的通讯作者
  • Ministry of Education of the People's Republic of China
  • Hong Kong University of Science and Technology
  • Shanghai Jiao Tong University
  • Nanjing University of Posts and Telecommunications
  • Zhengzhou University of Light Industry

科研成果: 期刊稿件文章同行评审

摘要

In recent years, scene text detection research has increasingly focused on arbitrary-shaped texts, where text representation is a fundamental problem. However, most existing methods still struggle to separate adjacent or overlapping texts due to ambiguous spatial positions of points or segmentation masks. Besides, the time efficiency of the entire pipeline is often neglected, resulting in sub-optimal inference speed. To tackle these problems, we first propose a novel text representation method based on robust subspace recovery, which robustly represents complex text shapes by combining orthogonal basis vectors learned from labeled text contours. These basis vectors capture basis contour patterns with distinct information, enabling clearer boundaries even in densely populated text scenarios. Moreover, we propose a dynamic sparse assignment scheme for positive samples that adaptively adjusts their weights during training, which not only accelerates inference speed by eliminating redundant predictions but also enhances feature learning by providing sufficient supervision signals. Building on these innovations, we present TextRSR, an accurate and efficient scene text detection network. Extensive experiments on challenging benchmarks demonstrate the superior accuracy and efficiency of TextRSR compared to state-of-the-art methods. Particularly, TextRSR achieves an F-measure of 88.5% at 37.8 frames per second (FPS) for CTW1500 dataset and an F-measure of 89.1% at 23.1 FPS for Total-Text dataset.

源语言英语
页(从-至)2550-2563
页数14
期刊IEEE Transactions on Multimedia
28
DOI
出版状态已出版 - 2026
已对外发布

指纹

探究 'TextRSR: Enhanced Arbitrary-Shaped Scene Text Representation via Robust Subspace Recovery' 的科研主题。它们共同构成独一无二的指纹。

引用此