Scene Text Recognition with Image-Text Matching-Guided Dictionary

Jiajun Wei, Hongjian Zhan, Xiao Tu, Yue Lu*, Umapada Pal

*Corresponding author for this work

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Employing a dictionary can efficiently rectify the deviation between the visual prediction and the ground truth in scene text recognition methods. However, the independence of the dictionary on the visual features may lead to incorrect rectification of accurate visual predictions. In this paper, we propose a new dictionary language model leveraging the Scene Image-Text Matching(SITM) network, which avoids the drawbacks of the explicit dictionary language model: 1) the independence of the visual features; 2) noisy choice in candidates etc. The SITM network accomplishes this by using Image-Text Contrastive (ITC) Learning to match an image with its corresponding text among candidates in the inference stage. ITC is widely used in vision-language learning to pull the positive image-text pair closer in feature space. Inspired by ITC, the SITM network combines the visual features and the text features of all candidates to identify the candidate with the minimum distance in the feature space. Our lexicon method achieves better results(93.8% accuracy) than the ordinary method results(92.1% accuracy) on six mainstream benchmarks. Additionally, we integrate our method with ABINet and establish new state-of-the-art results on several benchmarks.

Original languageEnglish
Title of host publicationDocument Analysis and Recognition – ICDAR 2023 - 17th International Conference, Proceedings
EditorsGernot A. Fink, Rajiv Jain, Koichi Kise, Richard Zanibbi
PublisherSpringer Science and Business Media Deutschland GmbH
Pages54-69
Number of pages16
ISBN (Print)9783031417306
DOIs
StatePublished - 2023
Event17th International Conference on Document Analysis and Recognition, ICDAR 2023 - San José, United States
Duration: 21 Aug 202326 Aug 2023

Publication series

NameLecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Volume14192 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference17th International Conference on Document Analysis and Recognition, ICDAR 2023
Country/TerritoryUnited States
CitySan José
Period21/08/2326/08/23

Keywords

  • Dictionary Language Model
  • Image-Text Contrastive Learning
  • Scene Image-Text Matching
  • Scene Text Recognition

Fingerprint

Dive into the research topics of 'Scene Text Recognition with Image-Text Matching-Guided Dictionary'. Together they form a unique fingerprint.

Cite this