Skip to main navigation Skip to search Skip to main content

Evidence-chain-driven multimodal retrieval question answering

  • East China Normal University
  • Donghua University

Research output: Contribution to journalArticlepeer-review

Abstract

Multimodal retrieval question answering is a critical task in Artificial General Intelligence. Existing methods typically retrieve key evidence by measuring semantic similarity between questions and retrieval information. However, they often fail to effectively distinguish between direct and indirect evidence in retrieving information, resulting in the oversight of some indirect evidence. Furthermore, current methods overlook the logical relationships between evidence, thus hindering the formation of complete evidence chains and impacting the accuracy of answer inference. To address this, we propose an Evidence-Chain-Driven Retrieval Question Answering Framework to effectively differentiate between direct and indirect evidence and form complete evidence chains through logical associations. Our framework consists of three modules. The question-evidence matching module identifies direct evidence from candidates and filters out irrelevant ones using a dual-encoder structure. It enhances discrimination against irrelevant evidence through semi-supervised contrastive learning. The indirect evidence retrieval module is utilized to further retrieve indirect evidence based on the question and the retrieved direct evidence. This module adopts a multi-layer transformer structure, with its core focusing on leveraging cross-attention mechanisms to effectively capture contextual relationships among the question, direct evidence, and indirect evidence. The logical correlation module is responsible for synthesizing the retrieved evidence into a coherent chain, which is then used to infer the final answer. We validated our proposed approach on two publicly available datasets, and the results demonstrate that our method outperforms existing approaches, achieving state-of-the-art (SOTA) performance. Extensive comparisons with contemporary LLM-based RAG baselines further show that our lightweight retrieval achieves superior precision and efficiency over LLM-guided evidence selection. Additionally, we conducted a series of ablation experiments to further validate the significance of the proposed method.

Original languageEnglish
Article number116358
JournalKnowledge-Based Systems
Volume348
DOIs
StatePublished - 3 Aug 2026

Keywords

  • Contrastive learning
  • Deep learning
  • Multimodal retrieval question answering
  • Transformer
  • Visual question answering

Fingerprint

Dive into the research topics of 'Evidence-chain-driven multimodal retrieval question answering'. Together they form a unique fingerprint.

Cite this