跳到主要导航 跳到搜索 跳到主要内容

LSTMVAEF: Vivid Layout via LSTM-Based Variational Autoencoder Framework

  • Jie He
  • , Xingjiao Wu
  • , Wenxin Hu
  • , Jing Yang*
  • *此作品的通讯作者
  • East China Normal University

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

The lack of training data is still a challenge in the Document Layout Analysis task (DLA). Synthetic data is an effective way to tackle this challenge. In this paper, we propose an LSTM-based Variational Autoencoder framework (LSTMVAF) to synthesize layouts for DLA. Compared with the previous method, our method can generate more complicated layouts and only need training data from DLA without extra annotation. We use LSTM models as basic models to learn the potential representing of class and position information of elements within a page. It is worth mentioning that we design a weight adaptation strategy to help model train faster. The experiment shows our model can generate more vivid layouts that only need a few real document pages.

源语言英语
主期刊名Document Analysis and Recognition – ICDAR 2021 - 16th International Conference, Proceedings
编辑Josep Lladós, Daniel Lopresti, Seiichi Uchida
出版商Springer Science and Business Media Deutschland GmbH
176-189
页数14
ISBN(印刷版)9783030863302
DOI
出版状态已出版 - 2021
活动16th International Conference on Document Analysis and Recognition, ICDAR 2021 - Lausanne, 瑞士
期限: 5 9月 202110 9月 2021

出版系列

姓名Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
12822 LNCS
ISSN(印刷版)0302-9743
ISSN(电子版)1611-3349

会议

会议16th International Conference on Document Analysis and Recognition, ICDAR 2021
国家/地区瑞士
Lausanne
时期5/09/2110/09/21

学术指纹

探究 'LSTMVAEF: Vivid Layout via LSTM-Based Variational Autoencoder Framework' 的科研主题。它们共同构成独一无二的学术指纹。

引用此