基于视觉感兴趣区域学习概念文本提示的医学图像分割方法

Zhu He ,  Haoran Zhang ,  Wentao Zhang ,  Shen Zhao ,  Qiqi Liu ,  Xiaohu Wu ,  Qicheng Lao

工程(英文) ›› 2026, Vol. 62 ›› Issue (7) : 198 -213.

工程(英文) ›› 2026, Vol. 62 ›› Issue (7) : 198 -213. DOI: 10.1016/j.eng.2026.04.006
研究论文

基于视觉感兴趣区域学习概念文本提示的医学图像分割方法

作者信息 +

Learning Conceptual Text Prompts from Visual Regions of Interest for Medical Image Segmentation

Author information +
文章历史 +

摘要

视觉语言分割模型在医学图像分割任务中表现出良好的性能。然而,这类模型的一个主要局限在于其依赖人工设计的文本输入。已有研究采用视觉问答技术半自动生成文本信息,但此类方法仍面临误差累积等问题。为此,本研究提出一种直接从视觉感兴趣区域(region of interestROI)学习概念文本提示的方法,以促进医学图像分割。首先,利用大型多模态模型从ROI中提取文本概念属性,生成粗粒度真实文本提示;之后,设计文本潜在空间变换模块,以ROI图像为输入生成细粒度伪文本提示,用于弥补上述真实文本提示在图像细节感知方面的不足。随后,将两类提示编码为统一的文本嵌入。在此基础上,提出一种自加噪知识蒸馏方法,将文本嵌入中的知识迁移至图像编码器的类别词元,从而在测试阶段实现无需文本输入的直接文本引导推理,并有效降低误差累积。该方法通过结合显式离散文本提示与隐式连续文本提示,有效引导视觉分割,同时显著减少了人工设计提示的需求。在13个医学图像分割数据集上的大量实验结果表明,所提出的方法优于当前最先进的视觉语言分割模型以及纯视觉分割模型,展现出更优的分割精度。


Abstract

Vision–language segmentation models (VLSMs) are effective in medical image segmentation tasks. However, a major limitation of these models is their dependence on manually crafted textual inputs. Studies have used visual question answering to semiautomatically generate textual information. However, these methods encounter challenges such as error accumulation. Herein, we propose a method to learn conceptual text prompts directly from visual regions of interest (ROIs) for facilitating medical image segmentation. We extracted textual conceptual attributes from ROIs using a large multimodal model to derive coarse real-text prompts. A text latent space transformation module accepted the ROI images as input for generating fine-grained pseudo-text prompts to compensate for the lack of image detail perception in the abovementioned real-text prompts. These prompts were encoded into a unified text embedding. Thereafter, we applied a self-adding noise knowledge distillation method to transfer the knowledge from text embedding to the class token of the image encoder, enabling direct text-guided inference during testing while reducing error accumulation. Our approach minimized the need for man- ual prompt design by leveraging explicit discrete and implicit continuous text prompts to effectively guide visual segmentation. Extensive evaluation across 13 medical image segmentation datasets demon- strated that our model outperformed the state-of-the-art VLSMs and vision-based segmentation models, exhibiting superior segmentation accuracy.

关键词

概念文本 / 提示学习 / 知识蒸馏 / 医学图像分割

Key words

Conceptual text Prompt learning / Knowledge distillation / Medical image segmentation

引用本文

引用格式 ▾
Zhu He,Haoran Zhang,Wentao Zhang,Shen Zhao,Qiqi Liu,Xiaohu Wu,Qicheng Lao. 基于视觉感兴趣区域学习概念文本提示的医学图像分割方法[J]. 工程(英文), 2026, 62(7): 198-213 DOI:10.1016/j.eng.2026.04.006

登录浏览全文

4963

注册一个新账户 忘记密码

参考文献

AI Summary AI Mindmap

516

访问

0

被引

详细

导航
相关文章

AI思维导图

/