跳到主要导航 跳到搜索 跳到主要内容

UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

  • Jiabo Ye
  • , Anwen Hu
  • , Haiyang Xu
  • , Qinghao Ye
  • , Ming Yan*
  • , Guohai Xu
  • , Chenliang Li
  • , Junfeng Tian
  • , Qi Qian
  • , Ji Zhang
  • , Qin Jin
  • , Liang He
  • , Xin Lin
  • , Fei Huang
  • *此作品的通讯作者
  • East China Normal University
  • Alibaba Group Holding Ltd.
  • Renmin University of China

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Text is ubiquitous in our visual world, conveying crucial information, such as in documents, websites, and everyday photographs. In this work, we propose UReader, a first exploration of universal OCR-free visually-situated language understanding based on the Multimodal Large Language Model (MLLM). By leveraging the shallow text recognition ability of the MLLM, we only finetuned 1.2% parameters and the training cost is much lower than previous work following domain-specific pretraining and finetuning paradigms. Concretely, UReader is jointly finetuned on a wide range of Visually-situated Language Understanding tasks via a unified instruction format. To enhance the visual text and semantic understanding, we further apply two auxiliary tasks with the same format, namely text reading and key points generation tasks. We design a shape-adaptive cropping module before the encoder-decoder architecture of MLLM to leverage the frozen low-resolution vision encoder for processing high-resolution images. Without downstream finetuning, our single model achieves state-of-the-art ocr-free performance in 8 out of 10 visually-situated language understanding tasks, across 5 domains: documents, tables, charts, natural images, and webpage screenshots. Codes and instruction-tuning datasets are released at https://github.com/LukeForeverYoung/UReader.

源语言英语
主期刊名Findings of the Association for Computational Linguistics
主期刊副标题EMNLP 2023
出版商Association for Computational Linguistics (ACL)
2841-2858
页数18
ISBN(电子版)9798891760615
DOI
出版状态已出版 - 2023
活动2023 Findings of the Association for Computational Linguistics: EMNLP 2023 - Hybrid, 新加坡
期限: 6 12月 202310 12月 2023

出版系列

姓名Findings of the Association for Computational Linguistics: EMNLP 2023

会议

会议2023 Findings of the Association for Computational Linguistics: EMNLP 2023
国家/地区新加坡
Hybrid
时期6/12/2310/12/23

指纹

探究 'UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model' 的科研主题。它们共同构成独一无二的指纹。

引用此