LW-CTrans: A lightweight hybrid network of CNN and Transformer for 3D medical image segmentation

IF 10.7 1区医学 Q1 COMPUTER SCIENCE, ARTIFICIAL INTELLIGENCE

Medical image analysis Pub Date : 2025-03-17 DOI:10.1016/j.media.2025.103545

Hulin Kuang , Yahui Wang , Xianzhen Tan , Jialin Yang , Jiarui Sun , Jin Liu , Wu Qiu , Jingyang Zhang , Jiulou Zhang , Chunfeng Yang , Jianxin Wang , Yang Chen

{"title":"LW-CTrans: A lightweight hybrid network of CNN and Transformer for 3D medical image segmentation","authors":"Hulin Kuang , Yahui Wang , Xianzhen Tan , Jialin Yang , Jiarui Sun , Jin Liu , Wu Qiu , Jingyang Zhang , Jiulou Zhang , Chunfeng Yang , Jianxin Wang , Yang Chen","doi":"10.1016/j.media.2025.103545","DOIUrl":null,"url":null,"abstract":"<div><div>Recent models based on convolutional neural network (CNN) and Transformer have achieved the promising performance for 3D medical image segmentation. However, these methods cannot segment small targets well even when equipping large parameters. Therefore, We design a novel lightweight hybrid network that combines the strengths of CNN and Transformers (LW-CTrans) and can boost the global and local representation capability at different stages. Specifically, we first design a dynamic stem that can accommodate images of various resolutions. In the first stage of the hybrid encoder, to capture local features with fewer parameters, we propose a multi-path convolution (MPConv) block. In the middle stages of the hybrid encoder, to learn global and local features meantime, we propose a multi-view pooling based Transformer (MVPFormer) which projects the 3D feature map onto three 2D subspaces to deal with small objects, and use the MPConv block for enhancing local representation learning. In the final stage, to mostly capture global features, only the proposed MVPFormer is used. Finally, to reduce the parameters of the decoder, we propose a multi-stage feature fusion module. Extensive experiments on 3 public datasets for three tasks: stroke lesion segmentation, pancreas cancer segmentation and brain tumor segmentation, show that the proposed LW-CTrans achieves Dices of 62.35±19.51%, 64.69±20.58% and 83.75±15.77% on the 3 datasets, respectively, outperforming 16 state-of-the-art methods, and the numbers of parameters (2.08M, 2.14M and 2.21M on 3 datasets, respectively) are smaller than the non-lightweight 3D methods and close to the lightweight methods. Besides, LW-CTrans also achieves the best performance for small lesion segmentation.</div></div>","PeriodicalId":18328,"journal":{"name":"Medical image analysis","volume":"102 ","pages":"Article 103545"},"PeriodicalIF":10.7000,"publicationDate":"2025-03-17","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Medical image analysis","FirstCategoryId":"5","ListUrlMain":"https://www.sciencedirect.com/science/article/pii/S1361841525000921","RegionNum":1,"RegionCategory":"医学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q1","JCRName":"COMPUTER SCIENCE, ARTIFICIAL INTELLIGENCE","Score":null,"Total":0}

引用次数: 0

Abstract

Recent models based on convolutional neural network (CNN) and Transformer have achieved the promising performance for 3D medical image segmentation. However, these methods cannot segment small targets well even when equipping large parameters. Therefore, We design a novel lightweight hybrid network that combines the strengths of CNN and Transformers (LW-CTrans) and can boost the global and local representation capability at different stages. Specifically, we first design a dynamic stem that can accommodate images of various resolutions. In the first stage of the hybrid encoder, to capture local features with fewer parameters, we propose a multi-path convolution (MPConv) block. In the middle stages of the hybrid encoder, to learn global and local features meantime, we propose a multi-view pooling based Transformer (MVPFormer) which projects the 3D feature map onto three 2D subspaces to deal with small objects, and use the MPConv block for enhancing local representation learning. In the final stage, to mostly capture global features, only the proposed MVPFormer is used. Finally, to reduce the parameters of the decoder, we propose a multi-stage feature fusion module. Extensive experiments on 3 public datasets for three tasks: stroke lesion segmentation, pancreas cancer segmentation and brain tumor segmentation, show that the proposed LW-CTrans achieves Dices of 62.35±19.51%, 64.69±20.58% and 83.75±15.77% on the 3 datasets, respectively, outperforming 16 state-of-the-art methods, and the numbers of parameters (2.08M, 2.14M and 2.21M on 3 datasets, respectively) are smaller than the non-lightweight 3D methods and close to the lightweight methods. Besides, LW-CTrans also achieves the best performance for small lesion segmentation.

查看原文本刊更多论文

LW-CTrans：用于3D医学图像分割的CNN和Transformer的轻量级混合网络

最近基于卷积神经网络（CNN）和Transformer的模型在三维医学图像分割中取得了很好的效果。然而，这些方法即使在配置大参数的情况下，也不能很好地分割小目标。因此，我们设计了一种新型的轻量级混合网络，它结合了CNN和transformer的优点（LW-CTrans），可以提高不同阶段的全局和局部表示能力。具体来说，我们首先设计了一个动态系统，可以容纳不同分辨率的图像。在混合编码器的第一阶段，为了捕获具有较少参数的局部特征，我们提出了一个多路径卷积（MPConv）块。在混合编码器的中间阶段，为了同时学习全局和局部特征，我们提出了一种基于多视图池的变压器（MVPFormer），该变压器将3D特征映射投影到三个二维子空间来处理小对象，并使用MPConv块来增强局部表征学习。在最后阶段，为了捕获全局特征，只使用建议的MVPFormer。最后，为了减少解码器的参数，我们提出了一个多级特征融合模块。在脑卒中病灶分割、胰腺癌分割和脑肿瘤分割3个公开数据集上进行的大量实验表明，所提出的LW-CTrans在3个数据集上的准确率分别为62.35±19.51%、64.69±20.58%和83.75±15.77%，优于16种最先进的方法，且参数个数（分别为2.08M、2.14M和2.21M）小于非轻量化3D方法，接近轻量化方法。此外，LW-CTrans在小病灶分割方面也取得了最好的效果。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

Medical image analysis 工程技术-工程：生物医学

CiteScore

22.10

自引率

6.40%

发文量

309

审稿时长

6.6 months

期刊介绍： Medical Image Analysis serves as a platform for sharing new research findings in the realm of medical and biological image analysis, with a focus on applications of computer vision, virtual reality, and robotics to biomedical imaging challenges. The journal prioritizes the publication of high-quality, original papers contributing to the fundamental science of processing, analyzing, and utilizing medical and biological images. It welcomes approaches utilizing biomedical image datasets across all spatial scales, from molecular/cellular imaging to tissue/organ imaging.