TY - JOUR
T1 - Unleashing Diffusion and State Space Models for Medical Image Segmentation
AU - Wu, Rong
AU - Chen, Ziqi
AU - Zhong, Liming
AU - Li, Heng
AU - Shu, Hai
N1 - Publisher Copyright:
© The Author(s) under exclusive licence to Society for Imaging Informatics in Medicine 2026.
PY - 2026
Y1 - 2026
N2 - Existing segmentation models trained on a single medical imaging dataset often lack robustness when encountering unseen organs or tumors. Developing a robust model capable of identifying rare or novel tumor categories not present during training is crucial for advancing medical imaging applications. We propose DSM, a novel framework that leverages diffusion and state space models to segment unseen tumor categories beyond the training data. DSM utilizes two sets of object queries trained within modified attention decoders to enhance classification accuracy. Initially, the model learns organ queries using an object-aware feature grouping strategy to capture organ-level visual features. It then refines tumor queries by focusing on diffusion-based visual prompts, enabling precise segmentation of previously unseen tumors. Furthermore, we incorporate diffusion-guided feature fusion to improve semantic segmentation performance. By integrating CLIP text embeddings, DSM captures category-sensitive classes to improve linguistic transfer knowledge, thereby enhancing the model’s robustness across diverse scenarios and multi-label tasks. DSM consistently outperforms state-of-the-art out-of-distribution detection methods, achieving improvements of 0.1962 in mean AUROC, 0.2675 in mean FPR95, and 0.1736 in mean DSC. Extensive experiments demonstrate the superior performance of DSM in various tumor segmentation tasks.
AB - Existing segmentation models trained on a single medical imaging dataset often lack robustness when encountering unseen organs or tumors. Developing a robust model capable of identifying rare or novel tumor categories not present during training is crucial for advancing medical imaging applications. We propose DSM, a novel framework that leverages diffusion and state space models to segment unseen tumor categories beyond the training data. DSM utilizes two sets of object queries trained within modified attention decoders to enhance classification accuracy. Initially, the model learns organ queries using an object-aware feature grouping strategy to capture organ-level visual features. It then refines tumor queries by focusing on diffusion-based visual prompts, enabling precise segmentation of previously unseen tumors. Furthermore, we incorporate diffusion-guided feature fusion to improve semantic segmentation performance. By integrating CLIP text embeddings, DSM captures category-sensitive classes to improve linguistic transfer knowledge, thereby enhancing the model’s robustness across diverse scenarios and multi-label tasks. DSM consistently outperforms state-of-the-art out-of-distribution detection methods, achieving improvements of 0.1962 in mean AUROC, 0.2675 in mean FPR95, and 0.1736 in mean DSC. Extensive experiments demonstrate the superior performance of DSM in various tumor segmentation tasks.
KW - Diffusion process
KW - Medical image segmentation
KW - State space model
KW - Text embeddings
UR - https://www.scopus.com/pages/publications/105034979499
U2 - 10.1007/s10278-026-01926-y
DO - 10.1007/s10278-026-01926-y
M3 - 文章
AN - SCOPUS:105034979499
SN - 2948-2933
JO - Journal of Imaging Informatics in Medicine
JF - Journal of Imaging Informatics in Medicine
ER -