跳到主要导航 跳到搜索 跳到主要内容

Enhancing scene text script identification through multi-task self-supervised learning

  • Jin Huang
  • , Li Liu*
  • , Yue Lu
  • , Ching Y. Suen
  • *此作品的通讯作者
  • Nanchang University
  • Concordia University

科研成果: 期刊稿件文章同行评审

摘要

This paper proposes a multi-task self-supervised learning framework for scene text script identification, aimed at addressing the challenges posed by diverse fonts, complex backgrounds, low resolutions, and frequent distortions in natural scenes. By leveraging unlabeled data, our approach learns robust image representations tailored for script identification. Three complementary tasks are designed: an advanced Jigsaw puzzle task to capture both local and global features, a spatial alignment-enhanced SwAV task to generalize across transformations while retaining spatial details, and a rotation prediction task to enhance spatial reasoning. Experiments on four benchmark datasets demonstrate that our method achieves state-of-the-art results, outperforming existing approaches. Ablation studies confirm the effectiveness of each module within our framework, showcasing its potential to reduce reliance on labeled data and enhance script identification in real-world applications. The code can be accessed at https://github.com/jin-or-king/Enhancing_scene_text_script_identification_through_multi-task_self-supervised_learning.

源语言英语
页(从-至)9571-9586
页数16
期刊Visual Computer
41
12
DOI
出版状态已出版 - 9月 2025

指纹

探究 'Enhancing scene text script identification through multi-task self-supervised learning' 的科研主题。它们共同构成独一无二的指纹。

引用此