跳到主要导航 跳到搜索 跳到主要内容

Cross-domain document layout analysis using document style guide

  • Xingjiao Wu
  • , Luwei Xiao
  • , Xiangcheng Du
  • , Yingbin Zheng
  • , Xin Li
  • , Tianlong Ma
  • , Cheng Jin*
  • , Liang He
  • *此作品的通讯作者
  • Fudan University
  • East China Normal University
  • Videt Lab

科研成果: 期刊稿件文章同行评审

摘要

Document layout analysis (DLA) is a crucial computer vision task that involves partitioning document images into high-level semantic regions such as figures, tables, backgrounds, and texts. Deep learning models for DLA typically require a large amount of labeled data, which can be expensive. Though some researchers use generated data for training, a substantial style gap exists between the generated and target data. Moreover, it is necessary to improve the quality of the generated samples to achieve better control. To address these challenges, we propose a cross-domain DLA framework called DL-DSG, which leverages document-style guidance. DL-DSG comprises three components: the document layout generator (DLG) responsible for generating document element locations, the document element decorator (DED) for filling the elements, and the document style discriminator (DSD) for style guidance. In addition to generating controlled documents, we also focus on bridging the gap between the generated and target samples. To this end, we introduce a novel strategy that transforms document style judgment into the document cross-domain style guidance component. We evaluate the effectiveness of DL-DSG on popular DLA datasets, including PubLayNet, DSSE-200, CS-150, and CDSSE, and demonstrate its superior performance.

源语言英语
文章编号123039
期刊Expert Systems with Applications
245
DOI
出版状态已出版 - 1 7月 2024

学术指纹

探究 'Cross-domain document layout analysis using document style guide' 的科研主题。它们共同构成独一无二的学术指纹。

引用此