TY - JOUR
T1 - Evidence-chain-driven multimodal retrieval question answering
AU - Yang, Shuwen
AU - Wu, Anran
AU - Wu, Xingjiao
AU - Zhao, Jiabao
AU - Wang, Linlin
AU - He, Liang
N1 - Publisher Copyright:
© 2026 Elsevier B.V.
PY - 2026/8/3
Y1 - 2026/8/3
N2 - Multimodal retrieval question answering is a critical task in Artificial General Intelligence. Existing methods typically retrieve key evidence by measuring semantic similarity between questions and retrieval information. However, they often fail to effectively distinguish between direct and indirect evidence in retrieving information, resulting in the oversight of some indirect evidence. Furthermore, current methods overlook the logical relationships between evidence, thus hindering the formation of complete evidence chains and impacting the accuracy of answer inference. To address this, we propose an Evidence-Chain-Driven Retrieval Question Answering Framework to effectively differentiate between direct and indirect evidence and form complete evidence chains through logical associations. Our framework consists of three modules. The question-evidence matching module identifies direct evidence from candidates and filters out irrelevant ones using a dual-encoder structure. It enhances discrimination against irrelevant evidence through semi-supervised contrastive learning. The indirect evidence retrieval module is utilized to further retrieve indirect evidence based on the question and the retrieved direct evidence. This module adopts a multi-layer transformer structure, with its core focusing on leveraging cross-attention mechanisms to effectively capture contextual relationships among the question, direct evidence, and indirect evidence. The logical correlation module is responsible for synthesizing the retrieved evidence into a coherent chain, which is then used to infer the final answer. We validated our proposed approach on two publicly available datasets, and the results demonstrate that our method outperforms existing approaches, achieving state-of-the-art (SOTA) performance. Extensive comparisons with contemporary LLM-based RAG baselines further show that our lightweight retrieval achieves superior precision and efficiency over LLM-guided evidence selection. Additionally, we conducted a series of ablation experiments to further validate the significance of the proposed method.
AB - Multimodal retrieval question answering is a critical task in Artificial General Intelligence. Existing methods typically retrieve key evidence by measuring semantic similarity between questions and retrieval information. However, they often fail to effectively distinguish between direct and indirect evidence in retrieving information, resulting in the oversight of some indirect evidence. Furthermore, current methods overlook the logical relationships between evidence, thus hindering the formation of complete evidence chains and impacting the accuracy of answer inference. To address this, we propose an Evidence-Chain-Driven Retrieval Question Answering Framework to effectively differentiate between direct and indirect evidence and form complete evidence chains through logical associations. Our framework consists of three modules. The question-evidence matching module identifies direct evidence from candidates and filters out irrelevant ones using a dual-encoder structure. It enhances discrimination against irrelevant evidence through semi-supervised contrastive learning. The indirect evidence retrieval module is utilized to further retrieve indirect evidence based on the question and the retrieved direct evidence. This module adopts a multi-layer transformer structure, with its core focusing on leveraging cross-attention mechanisms to effectively capture contextual relationships among the question, direct evidence, and indirect evidence. The logical correlation module is responsible for synthesizing the retrieved evidence into a coherent chain, which is then used to infer the final answer. We validated our proposed approach on two publicly available datasets, and the results demonstrate that our method outperforms existing approaches, achieving state-of-the-art (SOTA) performance. Extensive comparisons with contemporary LLM-based RAG baselines further show that our lightweight retrieval achieves superior precision and efficiency over LLM-guided evidence selection. Additionally, we conducted a series of ablation experiments to further validate the significance of the proposed method.
KW - Contrastive learning
KW - Deep learning
KW - Multimodal retrieval question answering
KW - Transformer
KW - Visual question answering
UR - https://www.scopus.com/pages/publications/105041056722
U2 - 10.1016/j.knosys.2026.116358
DO - 10.1016/j.knosys.2026.116358
M3 - 文章
AN - SCOPUS:105041056722
SN - 0950-7051
VL - 348
JO - Knowledge-Based Systems
JF - Knowledge-Based Systems
M1 - 116358
ER -