TY - JOUR
T1 - Human-like Single Image Reflection Removal via Deep Unfolding with Large Kernels
AU - Zhang, Haoyang
AU - Zhang, Junkang
AU - Fang, Faming
AU - Wang, Tingting
AU - Zhang, Guixu
AU - Song, Haichuan
N1 - Publisher Copyright:
© 1991-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - Single image reflection removal aims to remove the reflection part from a reflection image. Existing approaches predominantly rely on either mathematical modeling or deep learning techniques. However, their reflection removal methods often neglect the global understanding inherent to human visual perception when observing reflected scenes. The human visual system employs a ”global-to-local” hierarchical processing: first rapidly comprehending global scene information, then progressively focusing on local details to effectively distinguish real scenes from reflection images. Inspired by this observation, we developed LKCDU-Net, a deep unfolding model that innovatively apply large kernel convolution. This allows our model to effectively capture and utilize the global context of reflection information, while combining flexibility and interpretability. Specifically, we have meticulously crafted an initialization module for our model. This module employs large kernel convolutions to emulate the global perception capabilities of the human visual system, thereby capturing a more holistic view of the image. We then build an optimization model for local details refinement, improving the performance of our model through the deep unfolding technology. Benefiting from the initialization module, our deep unfolding model is capable of providing powerful performance with a small number of parameters. Extensive experiments on real-world reflection images show that LKCDU-Net outperforms SOTA models, effectively removing reflections while preserving details, which demonstrates its superior performance.
AB - Single image reflection removal aims to remove the reflection part from a reflection image. Existing approaches predominantly rely on either mathematical modeling or deep learning techniques. However, their reflection removal methods often neglect the global understanding inherent to human visual perception when observing reflected scenes. The human visual system employs a ”global-to-local” hierarchical processing: first rapidly comprehending global scene information, then progressively focusing on local details to effectively distinguish real scenes from reflection images. Inspired by this observation, we developed LKCDU-Net, a deep unfolding model that innovatively apply large kernel convolution. This allows our model to effectively capture and utilize the global context of reflection information, while combining flexibility and interpretability. Specifically, we have meticulously crafted an initialization module for our model. This module employs large kernel convolutions to emulate the global perception capabilities of the human visual system, thereby capturing a more holistic view of the image. We then build an optimization model for local details refinement, improving the performance of our model through the deep unfolding technology. Benefiting from the initialization module, our deep unfolding model is capable of providing powerful performance with a small number of parameters. Extensive experiments on real-world reflection images show that LKCDU-Net outperforms SOTA models, effectively removing reflections while preserving details, which demonstrates its superior performance.
KW - computer vision
KW - low-level image processing
KW - Single image reflection removal
UR - https://www.scopus.com/pages/publications/105042923696
U2 - 10.1109/TCSVT.2026.3704623
DO - 10.1109/TCSVT.2026.3704623
M3 - 文章
AN - SCOPUS:105042923696
SN - 1051-8215
JO - IEEE Transactions on Circuits and Systems for Video Technology
JF - IEEE Transactions on Circuits and Systems for Video Technology
ER -