IET Image Processing最新文献

筛选
英文 中文
TCAFNet: Transformer-Guided Cross-Scale Attention and Deep Semantic Fusion for Brain Tumour Classification From MRI TCAFNet:变压器引导的跨尺度注意和深度语义融合用于脑肿瘤MRI分类
IF 2.2 4区 计算机科学
IET Image Processing Pub Date : 2026-03-03 DOI: 10.1049/ipr2.70303
Maedeh Mohammadi, Majid Ziaratban
{"title":"TCAFNet: Transformer-Guided Cross-Scale Attention and Deep Semantic Fusion for Brain Tumour Classification From MRI","authors":"Maedeh Mohammadi,&nbsp;Majid Ziaratban","doi":"10.1049/ipr2.70303","DOIUrl":"https://doi.org/10.1049/ipr2.70303","url":null,"abstract":"<p>Accurate brain tumour classification from MRI is essential for computer-aided diagnosis and treatment planning. This paper proposes TCAFNet, a lightweight deep learning framework that integrates multi-scale attention mechanisms and transformer-based refinement to enhance both local feature discrimination and global contextual reasoning. The proposed model is built upon a pre-trained EfficientNetB0 backbone and employs an early semantic fusion strategy, in which deep features are injected into intermediate layers to guide feature learning from early stages. Cross-scale dependencies are explicitly modelled using pairwise cross-attention modules operating on multi-level feature maps. Each attention-enhanced representation is further refined using a customized adaptive convolutional block attention module and an adaptive feature pyramid attention block. The refined features are then integrated through an attention-based fusion mechanism and processed by a lightweight Transformer block to capture long-range spatial relationships. Experimental results show that the proposed model achieves an accuracy of 99.35% on the test set of the Figshare dataset, and demonstrates strong cross-dataset robustness with accuracies of 99.29% and 99.67% on the test sets of Nickparvar and Br35H datasets, respectively. These results demonstrate the effectiveness of combining early feature guidance, cross-scale attention and transformer-based modelling for robust multi-class brain tumour classification.</p>","PeriodicalId":56303,"journal":{"name":"IET Image Processing","volume":"20 1","pages":""},"PeriodicalIF":2.2,"publicationDate":"2026-03-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://ietresearch.onlinelibrary.wiley.com/doi/epdf/10.1049/ipr2.70303","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147653283","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
Advancing Traffic Safety in Dense Urban Settings: A Robust Real-Time Wrong-Way Driving Detection System with Hough Transform and YOLOv5 推进城市交通安全:基于Hough变换和YOLOv5的鲁棒实时错路驾驶检测系统
IF 2.2 4区 计算机科学
IET Image Processing Pub Date : 2026-03-03 DOI: 10.1049/ipr2.70322
Anik Majumder, Joya Mallick, Dhrubajyoti Das, Avishek Chowdhury, Shrabanti Debi
{"title":"Advancing Traffic Safety in Dense Urban Settings: A Robust Real-Time Wrong-Way Driving Detection System with Hough Transform and YOLOv5","authors":"Anik Majumder,&nbsp;Joya Mallick,&nbsp;Dhrubajyoti Das,&nbsp;Avishek Chowdhury,&nbsp;Shrabanti Debi","doi":"10.1049/ipr2.70322","DOIUrl":"https://doi.org/10.1049/ipr2.70322","url":null,"abstract":"<p>In Bangladesh, where a burgeoning population strains limited road networks, wrong-way driving (WWD) exacerbates traffic congestion and elevates accident risks, posing a critical public safety challenge. We present a pioneering real-time WWD detection system that integrates cutting-edge image processing and deep learning to transform traffic monitoring. Our methodology combines the Hough line transform for precise road boundary extraction with the state-of-the-art YOLOv5 algorithm for high-fidelity, real-time vehicle detection. Vehicle direction is inferred from single-frame orientation (distinguishing front/rear views) without multi-frame tracking, ensuring lightweight operation and scalable deployment. Tested on a novel, region-specific dataset from Chattogram, the system achieves an impressive 96% vehicle detection accuracy ([email protected]) during training and validation. Real-world evaluations on 20 diverse test videos yield an overall 60.6% recall for wrong-way vehicles, with a representative subset of 10 videos reaching 75% recall (and 41% precision). This research advances computer vision applications in traffic safety, providing an effective and deployable framework for reducing WWD incidents in high congestion urban regions and paving the way for enhanced intelligent transportation systems.</p>","PeriodicalId":56303,"journal":{"name":"IET Image Processing","volume":"20 1","pages":""},"PeriodicalIF":2.2,"publicationDate":"2026-03-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://ietresearch.onlinelibrary.wiley.com/doi/epdf/10.1049/ipr2.70322","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147653282","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
MSTNet: Multi-Scale Contextual Analysis Network for Semantic Segmentation of Remote Sensing Images 遥感图像语义分割的多尺度上下文分析网络
IF 2.2 4区 计算机科学
IET Image Processing Pub Date : 2026-03-02 DOI: 10.1049/ipr2.70326
Longbao Wang, Mingxuan Wang, Xiaoliang Luo, Lvchun Wang, Mu He, Chong Long, Meng Ding
{"title":"MSTNet: Multi-Scale Contextual Analysis Network for Semantic Segmentation of Remote Sensing Images","authors":"Longbao Wang,&nbsp;Mingxuan Wang,&nbsp;Xiaoliang Luo,&nbsp;Lvchun Wang,&nbsp;Mu He,&nbsp;Chong Long,&nbsp;Meng Ding","doi":"10.1049/ipr2.70326","DOIUrl":"https://doi.org/10.1049/ipr2.70326","url":null,"abstract":"<p>With the rapid development of deep learning, research on semantic segmentation of remote sensing images has made significant progress. However, there are common problems in remote sensing images, such as large-scale differences between different types of objects and unbalanced sample numbers, which leads to poor semantic segmentation results, especially for small targets and rare types of objects. To address these challenges, a remote sensing image semantic segmentation method based on multi-scale contextual information analysis named MSTNet is innovatively proposed. Its core design includes the semantic information enhancement module (SIE) of feature adaptive clustering, which strengthens the feature expression of different categories through adaptive clustering to alleviate sample imbalance; the weighted feature fusion module (WFF) adaptively aggregates cross-level features, cooperates with the multi-scale context enhancement module (MSCE), combines convolution and Transformer operations, and deeply mines local and global contexts to cope with scale changes; in addition, the network also contains a pixel space feature optimisation module (SFEM) to enhance spatial details. Experiments on the UAVid, LoveDA, Potsdam and Vaihingen datasets show that MSTNet significantly improves the ability to handle scale changes and imbalance problems and reaches advanced levels in key indicators such as OA, mIoU, and mF1, proving that MSTNet achieves competitive or even better performance.</p>","PeriodicalId":56303,"journal":{"name":"IET Image Processing","volume":"20 1","pages":""},"PeriodicalIF":2.2,"publicationDate":"2026-03-02","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://ietresearch.onlinelibrary.wiley.com/doi/epdf/10.1049/ipr2.70326","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147650367","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
Detection of Adversarial Examples Through Chaotic Features Extracted From Ordinal Patterns 从有序模式中提取混沌特征的对抗样例检测
IF 2.2 4区 计算机科学
IET Image Processing Pub Date : 2026-02-26 DOI: 10.1049/ipr2.70323
Harbinder Singh, Oscar Deniz, Anibal Pedraza, Simrandeep Singh, Gloria Bueno
{"title":"Detection of Adversarial Examples Through Chaotic Features Extracted From Ordinal Patterns","authors":"Harbinder Singh,&nbsp;Oscar Deniz,&nbsp;Anibal Pedraza,&nbsp;Simrandeep Singh,&nbsp;Gloria Bueno","doi":"10.1049/ipr2.70323","DOIUrl":"https://doi.org/10.1049/ipr2.70323","url":null,"abstract":"<p>Deep learning (DL) has significantly transformed computer vision, demonstrating remarkable achievements and extensive real-world applications. However, recent studies have highlighted a critical vulnerability of DL models to adversarial examples (AE), where slight perturbations in input data can lead to erroneous outputs. We observe that the behaviour of the AE is similar to a chaotic system, where a minor change in the input leads to a significantly different output. In response, we propose a novel approach for detecting and categorizing adversarial inputs encountered by classification neural networks. The proposed approach focuses on extracting statistical profiles, termed as chaotic feature vectors (CFVs), from a collection of features derived from ordinal patterns (OP). In this work, the proposed AE detection method is tested on seven attack methods and three image datasets including MNIST, FMNIST and CIFAR10. The results indicate that CFVs exhibit promising capabilities in discerning AE against various types of adversarial attacks on different datasets. This advancement lays the foundation for devising attack mitigation strategies, thereby enhancing the robustness and security of DL models in the face of adversarial threats.</p>","PeriodicalId":56303,"journal":{"name":"IET Image Processing","volume":"20 1","pages":""},"PeriodicalIF":2.2,"publicationDate":"2026-02-26","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://ietresearch.onlinelibrary.wiley.com/doi/epdf/10.1049/ipr2.70323","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147653360","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
Reversible Data Hiding in Encrypted Images Based on Bit-Plane Classification and Adaptive Group Coding 基于位平面分类和自适应分组编码的加密图像可逆数据隐藏
IF 2.2 4区 计算机科学
IET Image Processing Pub Date : 2026-02-24 DOI: 10.1049/ipr2.70310
Guoyan Zhou, Nianqiao Li, Chunqiang Yu, Xianquan Zhang, Zhenjun Tang
{"title":"Reversible Data Hiding in Encrypted Images Based on Bit-Plane Classification and Adaptive Group Coding","authors":"Guoyan Zhou,&nbsp;Nianqiao Li,&nbsp;Chunqiang Yu,&nbsp;Xianquan Zhang,&nbsp;Zhenjun Tang","doi":"10.1049/ipr2.70310","DOIUrl":"https://doi.org/10.1049/ipr2.70310","url":null,"abstract":"<p>Reversible data hiding in encrypted images (RDH-EI) is a key technique for secure communication, secure cloud storage and privacy protection. Most existing RDH-EI algorithms rely on a single and static bit-plane encoding strategy, failing to fully exploit the structural differences of bit-planes from the block level to the sequence level, thereby limiting further improvements on embedding capacity. To address this issue, this paper proposes a novel RDH-EI algorithm based on Bit-plane Classification and Adaptive Group Coding (hereafter BCAGC algorithm). The core contributions are as follows: (1) A bit-plane classification mechanism is proposed. It categorises bit-planes into simple or complex types according to the number of non-all-zero 4-bit sequence (NAZ-4BS) patterns. (2) An adaptive group coding scheme is proposed. It dynamically selects between fixed-length coding and Huffman coding based on bit-plane types and the frequency distribution of NAZ-4BS patterns, thereby achieving compact coding tables and efficient bit-plane compression. Experimental results demonstrate that the proposed BCAGC algorithm achieves average embedding rates of 4.1472, 4.0565 and 3.4465 bpp on the BOSSbase, BOWS-2 and UCID datasets, respectively, outperforming several state-of-the-art RDH-EI methods.</p>","PeriodicalId":56303,"journal":{"name":"IET Image Processing","volume":"20 1","pages":""},"PeriodicalIF":2.2,"publicationDate":"2026-02-24","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://ietresearch.onlinelibrary.wiley.com/doi/epdf/10.1049/ipr2.70310","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147653312","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
Enhancing Diffusion Models Towards Anomaly-Aware Reconstructions for Medical Image Anomaly Detection 面向异常感知重建的扩散模型增强医学图像异常检测
IF 2.2 4区 计算机科学
IET Image Processing Pub Date : 2026-02-24 DOI: 10.1049/ipr2.70311
Wanying Wu, Xiaofeng Qu, Fenghang Zhang, Mengjiao Zhang, Xizhan Gao, Sijie Niu
{"title":"Enhancing Diffusion Models Towards Anomaly-Aware Reconstructions for Medical Image Anomaly Detection","authors":"Wanying Wu,&nbsp;Xiaofeng Qu,&nbsp;Fenghang Zhang,&nbsp;Mengjiao Zhang,&nbsp;Xizhan Gao,&nbsp;Sijie Niu","doi":"10.1049/ipr2.70311","DOIUrl":"https://doi.org/10.1049/ipr2.70311","url":null,"abstract":"<p>Diffusion-based unsupervised anomaly detection in medical images has emerged as an effective paradigm, leveraging unlabelled healthy data to precisely characterize the distribution of normal anatomy and identify a wide range of pathological abnormalities. The method reconstructs a pseudo-healthy image from a potentially anomalous input and identifies anomalies by measuring pixel-wise reconstruction errors. However, existing approaches often preserve anomalous regions in the reconstruction, resulting in less prominent anomaly segmentation. Additionally, their inability to accurately restore normal areas can lead to increased false positives. In this work, we propose CS-Unet to advance this paradigm by realizing the concept of anomaly-aware reconstruction, defined as reconstructions that are consciously devoid of anomalies while faithfully restoring normal regions. Firstly, we propose a compression-expansion DenseNet (CompExDenseNet), which performs a dense cascade of nonlinear dimension transformations to extract compact feature representations, suppressing the reconstruction of anomalous patterns. Secondly, we design an attention gate (AG) unit to control the flow of low-frequency information, mitigating the leakage of anomalous information. Finally, we propose a frequency-domain adaptive residual convolution (FreAR) module that selectively enhances the most relevant frequency components to facilitate high-fidelity restoration of normal regions. Experimental results demonstrate that CS-Unet achieves outstanding performance in unsupervised anomaly detection, confirming its effectiveness.</p>","PeriodicalId":56303,"journal":{"name":"IET Image Processing","volume":"20 1","pages":""},"PeriodicalIF":2.2,"publicationDate":"2026-02-24","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://ietresearch.onlinelibrary.wiley.com/doi/epdf/10.1049/ipr2.70311","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147653313","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
SSGA-YOLO: A Lightweight Sonar Image Object Detection Network With Efficient Convolution and Acoustic-Aware Attention for Embedded Systems SSGA-YOLO:一种具有高效卷积和声感知关注的嵌入式系统声纳图像目标检测网络
IF 2.2 4区 计算机科学
IET Image Processing Pub Date : 2026-02-23 DOI: 10.1049/ipr2.70313
Yan Liu, Gan Yan, Tong Chen, Guanying Huo
{"title":"SSGA-YOLO: A Lightweight Sonar Image Object Detection Network With Efficient Convolution and Acoustic-Aware Attention for Embedded Systems","authors":"Yan Liu,&nbsp;Gan Yan,&nbsp;Tong Chen,&nbsp;Guanying Huo","doi":"10.1049/ipr2.70313","DOIUrl":"https://doi.org/10.1049/ipr2.70313","url":null,"abstract":"<p>To address the problems of high computational complexity and limited deploy ability in underwater sonar image detection, we propose SSGA-YOLO, a lightweight and efficient object detection algorithm optimised for sonar imagery and successfully deployed on the Ascend AI embedded platform. SSGA-YOLO focuses on three key aspects to achieve a favourable balance between accuracy and efficiency. The S-Net backbone employs depthwise separable convolutions and a lightweight attention mechanism to enhance the extraction of weak echo and shadow features in sonar images while reducing redundancy for efficient deployment. The Efficient Group Shuffle Convolution (EGSConv) enhances cross-channel feature interaction to improve the detection of small, low-contrast sonar targets and the Lightweight Shuffle-Aware Group Attention (LSGA) refines key acoustic and spatial cues in the presence of strong noise. Furthermore, SSGA-YOLO significantly reduces model parameters and computational complexity: compared to YOLOv8n, it achieves reductions of 79.82% and 74.07% in parameter count and GFLOPs, respectively. To evaluate model performance across diverse environments, embedded deployment experiments were conducted on three datasets representing distinct scenarios: MDFD for controlled artificial tanks, UATD for complex natural waters and MOTfish for dynamic video sequences. SSGA-YOLO consistently achieves high detection accuracy, with an mAP50 exceeding 0.930 on all datasets and peaking at 0.983 on MDFD. In terms of inference efficiency, the model demonstrates exceptional real-time capability, reaching a frame rate of 65.77 FPS on MOTfish. These results outperform other lightweight detectors, confirming the model's effectiveness for practical underwater applications.</p>","PeriodicalId":56303,"journal":{"name":"IET Image Processing","volume":"20 1","pages":""},"PeriodicalIF":2.2,"publicationDate":"2026-02-23","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://ietresearch.onlinelibrary.wiley.com/doi/epdf/10.1049/ipr2.70313","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147288502","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
The Power of Modality: Improving Polyp Segmentation With Multimodal Information 模态的力量:利用多模态信息改进息肉分割
IF 2.2 4区 计算机科学
IET Image Processing Pub Date : 2026-02-19 DOI: 10.1049/ipr2.70305
Fang Wang, Pu Wang, Meng Zhao, Chenggang Shan, Zhen Yang
{"title":"The Power of Modality: Improving Polyp Segmentation With Multimodal Information","authors":"Fang Wang,&nbsp;Pu Wang,&nbsp;Meng Zhao,&nbsp;Chenggang Shan,&nbsp;Zhen Yang","doi":"10.1049/ipr2.70305","DOIUrl":"10.1049/ipr2.70305","url":null,"abstract":"<p>Accurate polyp segmentation from a single color image remains a significant challenge due to the complex appearance of lesions and the lack of diverse contextual priors. Existing methods usually rely on a limited image prior, which leads to suboptimal results. We propose a novel approach that leverages rich contextual information from multiple modalities (including depth, higher order semantics, and edge features) to produce precise segmentation results. By integrating the Depth Anything model, Large Language Models, and Edge Extractor, our method effectively fuses these diverse priors to overcome the limitations of single modality approaches. In addition, we introduce a diffusion modeling framework to bring powerful generative priors for endoscopic image segmentation. This is a flexible deep network architecture that efficiently fuses multimodal information and can accommodate any number of multimodal inputs. Extensive experimental results demonstrate that our approach achieves promising performance on several benchmarks.</p>","PeriodicalId":56303,"journal":{"name":"IET Image Processing","volume":"20 1","pages":""},"PeriodicalIF":2.2,"publicationDate":"2026-02-19","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://ietresearch.onlinelibrary.wiley.com/doi/epdf/10.1049/ipr2.70305","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147299942","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
End-to-End Multi-Entity Customization 端到端多实体定制
IF 2.2 4区 计算机科学
IET Image Processing Pub Date : 2026-02-16 DOI: 10.1049/ipr2.70306
Wonhark Park, Jaehyun Lee, Wonsik Shin, Junhoo Lee, Nojun Kwak
{"title":"End-to-End Multi-Entity Customization","authors":"Wonhark Park,&nbsp;Jaehyun Lee,&nbsp;Wonsik Shin,&nbsp;Junhoo Lee,&nbsp;Nojun Kwak","doi":"10.1049/ipr2.70306","DOIUrl":"10.1049/ipr2.70306","url":null,"abstract":"<p>Recent advancements in text-to-image (T2I) models have enabled the synthesis of personalized images that align closely with user-specified prompts, especially through the use of modifiers. However, generating multiple detailed objects with distinct modifiers in a single image remains challenging due to concept-mixing, resulting from the difficulty of capturing interactions among text tokens. This paper proposes a modifier-based approach to mitigate concept-mixing by addressing the interaction among text tokens. Our method enables practical multi-personalization while preserving the original T2I model's straightforward inference pipeline. Without structural guidance, it ensures seamless object interaction with enhanced consistency. Through a loss-based finetuning approach, our method is adaptable to various concept-learning algorithms, enabling plug-and-play functionality. Through both qualitative and quantitative evaluations, we demonstrate that our method effectively resolves concept-mixing issues to better preserve concepts' identities and outperforms recent baselines in both quantitative and qualitative results. Our code will be publicly available.</p>","PeriodicalId":56303,"journal":{"name":"IET Image Processing","volume":"20 1","pages":""},"PeriodicalIF":2.2,"publicationDate":"2026-02-16","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://ietresearch.onlinelibrary.wiley.com/doi/epdf/10.1049/ipr2.70306","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"146217270","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
Realtime Data Augmentation for Breast Cancer Dataset: Dynamic Fine-Tuning Bounding Box Coordinates and Segmentation Mask 乳腺癌数据集的实时数据增强:动态微调边界框坐标和分割掩码
IF 2.2 4区 计算机科学
IET Image Processing Pub Date : 2026-02-12 DOI: 10.1049/ipr2.70308
Hassan Mahichi, Vahid Ghods, Mohammad Karim Sohrabi, Arash Sabbaghi
{"title":"Realtime Data Augmentation for Breast Cancer Dataset: Dynamic Fine-Tuning Bounding Box Coordinates and Segmentation Mask","authors":"Hassan Mahichi,&nbsp;Vahid Ghods,&nbsp;Mohammad Karim Sohrabi,&nbsp;Arash Sabbaghi","doi":"10.1049/ipr2.70308","DOIUrl":"10.1049/ipr2.70308","url":null,"abstract":"<p>Data augmentation is crucial for training deep learning models in breast cancer detection and segmentation, but conventional methods can cause annotation errors and visual artefacts, which harm instance-level localisation, generalisation and clinical reliability. This study aims to develop an annotation-aware, real-time data augmentation framework that preserves spatial and clinical integrity during geometric transformations to improve model robustness and performance. A real-time augmentation framework dynamically recalculates bounding box coordinates and segmentation masks during cropping and rotation. Instance-level annotation correction is performed on-the-fly within the training pipeline, without offline preprocessing, additional storage, or generative modelling and is evaluated on the DDSM, INbreast and BUSI datasets. Experimental results show consistent performance gains across all datasets and models, with detection metrics improving by an average of 4.2–4.4% and segmentation accuracy increasing by up to 9.3%. The real-time implementation achieves low preprocessing latency (≈0.12 s per batch, 3.75 ms per image), enabling high-throughput training without added computational overhead. By preserving annotation integrity during geometric transformations, the proposed framework provides a computationally efficient and easily integrable solution for breast cancer imaging, with broader applicability to other medical image analysis tasks requiring precise spatial annotations.</p>","PeriodicalId":56303,"journal":{"name":"IET Image Processing","volume":"20 1","pages":""},"PeriodicalIF":2.2,"publicationDate":"2026-02-12","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://ietresearch.onlinelibrary.wiley.com/doi/epdf/10.1049/ipr2.70308","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"146217041","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
0
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
确定
请完成安全验证×
相关产品
×
本文献相关产品
联系我们:info@booksci.cn Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。 Copyright © 2023 布克学术 All rights reserved.
京ICP备2023020795号-1
ghs 京公网安备 11010802042870号
Book学术文献互助
Book学术文献互助群
群 号:604180095
Book学术官方微信
小红书