[1]孟祥福,杨雨卓,张霄雁,等.基于分组注意力并行编码的医学图像分割网络[J].智能系统学报,2026,21(5):1194-1210.[doi:10.11992/tis.202511010]
MENG Xiangfu,YANG Yuzhuo,ZHANG Xiaoyan,et al.Medical image segmentation network based on group attention parallel encoding[J].CAAI transactions on intelligent systems,2026,21(5):1194-1210.[doi:10.11992/tis.202511010]
点击复制
《智能系统学报》[ISSN 1673-4785/CN 23-1538/TP] 卷:
21
期数:
2026年第5期
页码:
1194-1210
栏目:
学术论文—机器学习
出版日期:
2026-09-05
- Title:
-
Medical image segmentation network based on group attention parallel encoding
- 作者:
-
孟祥福1, 杨雨卓1, 张霄雁1, 李帅2
-
1. 辽宁工程技术大学 电子与信息工程学院, 辽宁 葫芦岛 125105;
2. 辽宁省健康产业集团阜新矿总医院, 辽宁 阜新 123000
- Author(s):
-
MENG Xiangfu1, YANG Yuzhuo1, ZHANG Xiaoyan1, LI Shuai2
-
1. School of Electronic and Information Engineering, Liaoning Technical University, Huludao 125105, China;
2. Liaoning Health Industry Group Fuxin Mine General Hospital, Fuxin 123000, China
-
- 关键词:
-
医学图像分割; 双分支编码器; 分组注意力; 特征提取; Transformer; 卷积神经网络; 并行融合; 空间-通道解码
- Keywords:
-
medical image segmentation; dual-branch encoder; group attention; feature extraction; Transformer; CNN; parallel fusion; spatial-channel decoding
- 分类号:
-
TP391.4
- DOI:
-
10.11992/tis.202511010
- 摘要:
-
针对医学图像分割任务,传统单分支卷积神经网络(convolutional neural network, CNN)架构因感受野限制难以有效整合局部细节与全局语义,导致多尺度结构建模不足及模态泛化能力弱。为此提出了一种分组注意力并行融合编码的分割网络,采用Transformer和CNN并行双分支编码器,分别提取图像的全局语义信息与局部细节信息。编码器后设计分组注意力并行融合模块,解决了局部与全局语义割裂及多尺度结构冲突问题;解码器中提出的空间-通道双注意力门控模块,能够有效解决医学图像分割中的特征干扰问题,提升分割精度。在Synapse、ACDC和AVT等数据集上开展了大量实验,Dice相似系数分别为84.86%、91.66%和88.79%,95%豪斯多夫距离分别为13.54、1.20和4.02 mm。实验结果表明,与现有主流医学图像分割模型相比,所提方法在分割任务中表现出更高的分割精度和鲁棒性,尤其在处理复杂解剖结构的器官数据时优势显著。
- Abstract:
-
For the task of medical image segmentation, traditional single-branch convolutional neural network (CNN) architectures, limited by their receptive fields, struggle to effectively integrate local details with global semantics. This results in insufficient modeling of multi-scale structures and weak generalization across modalities. To address these limitations, we propose a novel segmentation network with group attention parallel fusion encoding. It employs a parallel dual-branch encoder combining Transformer and CNN to extract global semantic information and local details from images, respectively. The group attention parallel fusion module following the encoder resolves the disconnection between local and global semantics as well as conflicts in multi-scale structures. Additionally, the spatial-channel dual-attention gating module in the decoder effectively tackles feature interference in medical image segmentation, thereby enhancing segmentation accuracy. Extensive experiments were conducted on the Synapse, ACDC and AVT datasets, yielding dice similarity coefficient (DSC) of 84.86%, 91.66%, and 88.79%, and 95% hausdorff distance (HD95) of 13.54, 1.20, and 4.02 mm, respectively. The results demonstrate that compared with existing major medical image segmentation models, our approach achieves higher segmentation accuracy and robustness, especially when dealing with complex anatomical structures of organ data.
更新日期/Last Update:
2026-09-05