Jian Zhang , Ze Ji , Changdong Zhao , Meng Huang , Xufei Hu , Qilin Li , Ming Li , Heng Zhang
{"title":"PVT-MSFF: A Pyramid Vision Transformer with Multi-Scale Feature Fusion for polyp segmentation in endoscopic images","authors":"Jian Zhang , Ze Ji , Changdong Zhao , Meng Huang , Xufei Hu , Qilin Li , Ming Li , Heng Zhang","doi":"10.1016/j.eij.2026.101001","DOIUrl":null,"url":null,"abstract":"<div><div>Colorectal cancer (CRC) is one of the most fatal malignancies worldwide, and its early diagnosis relies on the accurate detection and segmentation of polyps in endoscopic images. However, existing methods are often challenged by the diverse morphology of polyps, ambiguous boundaries, and differences in data centers, which can lead to missed detections and limited generalization. Here, we propose a novel Pyramid Vision Transformer with Multi-Scale Feature Fusion network (PVT-MSFF), which combines a Pyramid Vision Transformer encoder with a cooperative multi-module strategy. Our method introduces a Feature Enhancement Module (FEM) with cross-attention mechanism, a Multi-Scale Fusion Module (MSFM) for hierarchical feature expression, and a Global Context Sensing (GCS) module to enhance boundary sensitivity and semantic integration. We evaluate PVT-MSFF on five public polyp segmentation benchmark datasets, including Kvasir-SEG, ClinicDB, ColonDB, ETIS, and CVC-300. Experimental results demonstrate that our method achieves highly competitive segmentation performance, with mDice scores of 0.921, 0.949, 0.813, 0.809, and 0.878 on the respective datasets. Although SAM2-UNet attains 0.928 mDice on Kvasir-SEG and ASPS reaches 0.950 mDice on CVC-ClinicDB, showing results closely comparable to ours, our model exhibits significant advantages in computational efficiency, with parameter count (25.16M) and computational cost (10.12 GFLOPs) substantially lower than these large-scale model-based approaches. This excellent balance between high accuracy and efficiency makes PVT-MSFF particularly valuable for practical clinical applications. Moreover, validation on a clinical dataset (G-endoscope) further confirms robust generalization and accuracy, highlighting the potential of PVT-MSFF for intelligent endoscopy systems and computer-assisted diagnosis. Our code will be available at: <span><span>https://github.com/jize123457/PVT-MSFF</span><svg><path></path></svg></span>.</div></div>","PeriodicalId":56010,"journal":{"name":"Egyptian Informatics Journal","volume":"34 ","pages":"Article 101001"},"PeriodicalIF":4.2000,"publicationDate":"2026-06-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Egyptian Informatics Journal","FirstCategoryId":"94","ListUrlMain":"https://www.sciencedirect.com/science/article/pii/S1110866526001180","RegionNum":3,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"2026/6/11 0:00:00","PubModel":"Epub","JCR":"Q1","JCRName":"COMPUTER SCIENCE, ARTIFICIAL INTELLIGENCE","Score":null,"Total":0}
引用次数: 0
Abstract
Colorectal cancer (CRC) is one of the most fatal malignancies worldwide, and its early diagnosis relies on the accurate detection and segmentation of polyps in endoscopic images. However, existing methods are often challenged by the diverse morphology of polyps, ambiguous boundaries, and differences in data centers, which can lead to missed detections and limited generalization. Here, we propose a novel Pyramid Vision Transformer with Multi-Scale Feature Fusion network (PVT-MSFF), which combines a Pyramid Vision Transformer encoder with a cooperative multi-module strategy. Our method introduces a Feature Enhancement Module (FEM) with cross-attention mechanism, a Multi-Scale Fusion Module (MSFM) for hierarchical feature expression, and a Global Context Sensing (GCS) module to enhance boundary sensitivity and semantic integration. We evaluate PVT-MSFF on five public polyp segmentation benchmark datasets, including Kvasir-SEG, ClinicDB, ColonDB, ETIS, and CVC-300. Experimental results demonstrate that our method achieves highly competitive segmentation performance, with mDice scores of 0.921, 0.949, 0.813, 0.809, and 0.878 on the respective datasets. Although SAM2-UNet attains 0.928 mDice on Kvasir-SEG and ASPS reaches 0.950 mDice on CVC-ClinicDB, showing results closely comparable to ours, our model exhibits significant advantages in computational efficiency, with parameter count (25.16M) and computational cost (10.12 GFLOPs) substantially lower than these large-scale model-based approaches. This excellent balance between high accuracy and efficiency makes PVT-MSFF particularly valuable for practical clinical applications. Moreover, validation on a clinical dataset (G-endoscope) further confirms robust generalization and accuracy, highlighting the potential of PVT-MSFF for intelligent endoscopy systems and computer-assisted diagnosis. Our code will be available at: https://github.com/jize123457/PVT-MSFF.
期刊介绍:
The Egyptian Informatics Journal is published by the Faculty of Computers and Artificial Intelligence, Cairo University. This Journal provides a forum for the state-of-the-art research and development in the fields of computing, including computer sciences, information technologies, information systems, operations research and decision support. Innovative and not-previously-published work in subjects covered by the Journal is encouraged to be submitted, whether from academic, research or commercial sources.