跳到主要导航 跳到搜索 跳到主要内容

Visual Graph Reasoning Network

  • East China Normal University
  • Zhejiang University

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Visual question answering (VQA) is a fundamental and challenging cross-modal task. This task requires the model to fully understand the image's content and reason out the answer based on the question. Existing VQA models understand visual content mainly based on bottom-up or grid features. However, both types of vision features have some drawbacks. The discreteness and independence of bottom-up features pre-vent models from adequately performing relational reasoning. Image segmentation by grid features leads to the fragmentation of meaningful visual regions, limiting the cross-modal alignment capability of the model. Therefore, we proposed a more flexible method called Visual Graph. It can connect different patches according to semantic similarity and spatial relevance to model the potential relationships and cluster the adjacent homologous patches. Based on the Visual Graph, we designed a Visual Graph Reasoning Network for VQA. We evaluated our model on GQA and VQA-v2. The experimental results show that our models can achieve excellent performance between single models.

源语言英语
主期刊名ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing, Proceedings
出版商Institute of Electrical and Electronics Engineers Inc.
ISBN(电子版)9781728163277
DOI
出版状态已出版 - 2023
活动48th IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2023 - Rhodes Island, 希腊
期限: 4 6月 202310 6月 2023

出版系列

姓名ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings
2023-June
ISSN(印刷版)1520-6149

会议

会议48th IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2023
国家/地区希腊
Rhodes Island
时期4/06/2310/06/23

学术指纹

探究 'Visual Graph Reasoning Network' 的科研主题。它们共同构成独一无二的学术指纹。

引用此