跳到主要导航 跳到搜索 跳到主要内容

Scene Text Recognition with Temporal Convolutional Encoder

  • Shanghai Key Laboratory of Multidimensional Information Processing
  • East China Normal University
  • Videt Lab

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Texts from scene images typically consist of several characters and exhibit a characteristic sequence structure. Existing methods capture the structure with the sequence-to-sequence models by an encoder to have the visual representations and then a decoder to translate the features into the label sequence. In this paper, we study text recognition framework by considering the long-term temporal dependencies in the encoder stage. We demonstrate that the proposed Temporal Convolutional Encoder with increased sequential extents improves the accuracy of text recognition. We also study the impact of different attention modules in convolutional blocks for learning accurate text representations. We conduct comparisons on seven datasets and the experiments demonstrate the effectiveness of our proposed approach.

源语言英语
主期刊名2020 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2020 - Proceedings
出版商Institute of Electrical and Electronics Engineers Inc.
2383-2387
页数5
ISBN(电子版)9781509066315
DOI
出版状态已出版 - 5月 2020
已对外发布
活动2020 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2020 - Barcelona, 西班牙
期限: 4 5月 20208 5月 2020

出版系列

姓名ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings
2020-May
ISSN(印刷版)1520-6149

会议

会议2020 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2020
国家/地区西班牙
Barcelona
时期4/05/208/05/20

学术指纹

探究 'Scene Text Recognition with Temporal Convolutional Encoder' 的科研主题。它们共同构成独一无二的学术指纹。

引用此