{"title":"TCAFNet: Transformer-Guided Cross-Scale Attention and Deep Semantic Fusion for Brain Tumour Classification From MRI","authors":"Maedeh Mohammadi, Majid Ziaratban","doi":"10.1049/ipr2.70303","DOIUrl":"https://doi.org/10.1049/ipr2.70303","url":null,"abstract":"<p>Accurate brain tumour classification from MRI is essential for computer-aided diagnosis and treatment planning. This paper proposes TCAFNet, a lightweight deep learning framework that integrates multi-scale attention mechanisms and transformer-based refinement to enhance both local feature discrimination and global contextual reasoning. The proposed model is built upon a pre-trained EfficientNetB0 backbone and employs an early semantic fusion strategy, in which deep features are injected into intermediate layers to guide feature learning from early stages. Cross-scale dependencies are explicitly modelled using pairwise cross-attention modules operating on multi-level feature maps. Each attention-enhanced representation is further refined using a customized adaptive convolutional block attention module and an adaptive feature pyramid attention block. The refined features are then integrated through an attention-based fusion mechanism and processed by a lightweight Transformer block to capture long-range spatial relationships. Experimental results show that the proposed model achieves an accuracy of 99.35% on the test set of the Figshare dataset, and demonstrates strong cross-dataset robustness with accuracies of 99.29% and 99.67% on the test sets of Nickparvar and Br35H datasets, respectively. These results demonstrate the effectiveness of combining early feature guidance, cross-scale attention and transformer-based modelling for robust multi-class brain tumour classification.</p>","PeriodicalId":56303,"journal":{"name":"IET Image Processing","volume":"20 1","pages":""},"PeriodicalIF":2.2,"publicationDate":"2026-03-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://ietresearch.onlinelibrary.wiley.com/doi/epdf/10.1049/ipr2.70303","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147653283","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Advancing Traffic Safety in Dense Urban Settings: A Robust Real-Time Wrong-Way Driving Detection System with Hough Transform and YOLOv5","authors":"Anik Majumder, Joya Mallick, Dhrubajyoti Das, Avishek Chowdhury, Shrabanti Debi","doi":"10.1049/ipr2.70322","DOIUrl":"https://doi.org/10.1049/ipr2.70322","url":null,"abstract":"<p>In Bangladesh, where a burgeoning population strains limited road networks, wrong-way driving (WWD) exacerbates traffic congestion and elevates accident risks, posing a critical public safety challenge. We present a pioneering real-time WWD detection system that integrates cutting-edge image processing and deep learning to transform traffic monitoring. Our methodology combines the Hough line transform for precise road boundary extraction with the state-of-the-art YOLOv5 algorithm for high-fidelity, real-time vehicle detection. Vehicle direction is inferred from single-frame orientation (distinguishing front/rear views) without multi-frame tracking, ensuring lightweight operation and scalable deployment. Tested on a novel, region-specific dataset from Chattogram, the system achieves an impressive 96% vehicle detection accuracy ([email protected]) during training and validation. Real-world evaluations on 20 diverse test videos yield an overall 60.6% recall for wrong-way vehicles, with a representative subset of 10 videos reaching 75% recall (and 41% precision). This research advances computer vision applications in traffic safety, providing an effective and deployable framework for reducing WWD incidents in high congestion urban regions and paving the way for enhanced intelligent transportation systems.</p>","PeriodicalId":56303,"journal":{"name":"IET Image Processing","volume":"20 1","pages":""},"PeriodicalIF":2.2,"publicationDate":"2026-03-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://ietresearch.onlinelibrary.wiley.com/doi/epdf/10.1049/ipr2.70322","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147653282","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"MSTNet: Multi-Scale Contextual Analysis Network for Semantic Segmentation of Remote Sensing Images","authors":"Longbao Wang, Mingxuan Wang, Xiaoliang Luo, Lvchun Wang, Mu He, Chong Long, Meng Ding","doi":"10.1049/ipr2.70326","DOIUrl":"https://doi.org/10.1049/ipr2.70326","url":null,"abstract":"<p>With the rapid development of deep learning, research on semantic segmentation of remote sensing images has made significant progress. However, there are common problems in remote sensing images, such as large-scale differences between different types of objects and unbalanced sample numbers, which leads to poor semantic segmentation results, especially for small targets and rare types of objects. To address these challenges, a remote sensing image semantic segmentation method based on multi-scale contextual information analysis named MSTNet is innovatively proposed. Its core design includes the semantic information enhancement module (SIE) of feature adaptive clustering, which strengthens the feature expression of different categories through adaptive clustering to alleviate sample imbalance; the weighted feature fusion module (WFF) adaptively aggregates cross-level features, cooperates with the multi-scale context enhancement module (MSCE), combines convolution and Transformer operations, and deeply mines local and global contexts to cope with scale changes; in addition, the network also contains a pixel space feature optimisation module (SFEM) to enhance spatial details. Experiments on the UAVid, LoveDA, Potsdam and Vaihingen datasets show that MSTNet significantly improves the ability to handle scale changes and imbalance problems and reaches advanced levels in key indicators such as OA, mIoU, and mF1, proving that MSTNet achieves competitive or even better performance.</p>","PeriodicalId":56303,"journal":{"name":"IET Image Processing","volume":"20 1","pages":""},"PeriodicalIF":2.2,"publicationDate":"2026-03-02","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://ietresearch.onlinelibrary.wiley.com/doi/epdf/10.1049/ipr2.70326","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147650367","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Harbinder Singh, Oscar Deniz, Anibal Pedraza, Simrandeep Singh, Gloria Bueno
{"title":"Detection of Adversarial Examples Through Chaotic Features Extracted From Ordinal Patterns","authors":"Harbinder Singh, Oscar Deniz, Anibal Pedraza, Simrandeep Singh, Gloria Bueno","doi":"10.1049/ipr2.70323","DOIUrl":"https://doi.org/10.1049/ipr2.70323","url":null,"abstract":"<p>Deep learning (DL) has significantly transformed computer vision, demonstrating remarkable achievements and extensive real-world applications. However, recent studies have highlighted a critical vulnerability of DL models to adversarial examples (AE), where slight perturbations in input data can lead to erroneous outputs. We observe that the behaviour of the AE is similar to a chaotic system, where a minor change in the input leads to a significantly different output. In response, we propose a novel approach for detecting and categorizing adversarial inputs encountered by classification neural networks. The proposed approach focuses on extracting statistical profiles, termed as chaotic feature vectors (CFVs), from a collection of features derived from ordinal patterns (OP). In this work, the proposed AE detection method is tested on seven attack methods and three image datasets including MNIST, FMNIST and CIFAR10. The results indicate that CFVs exhibit promising capabilities in discerning AE against various types of adversarial attacks on different datasets. This advancement lays the foundation for devising attack mitigation strategies, thereby enhancing the robustness and security of DL models in the face of adversarial threats.</p>","PeriodicalId":56303,"journal":{"name":"IET Image Processing","volume":"20 1","pages":""},"PeriodicalIF":2.2,"publicationDate":"2026-02-26","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://ietresearch.onlinelibrary.wiley.com/doi/epdf/10.1049/ipr2.70323","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147653360","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Reversible Data Hiding in Encrypted Images Based on Bit-Plane Classification and Adaptive Group Coding","authors":"Guoyan Zhou, Nianqiao Li, Chunqiang Yu, Xianquan Zhang, Zhenjun Tang","doi":"10.1049/ipr2.70310","DOIUrl":"https://doi.org/10.1049/ipr2.70310","url":null,"abstract":"<p>Reversible data hiding in encrypted images (RDH-EI) is a key technique for secure communication, secure cloud storage and privacy protection. Most existing RDH-EI algorithms rely on a single and static bit-plane encoding strategy, failing to fully exploit the structural differences of bit-planes from the block level to the sequence level, thereby limiting further improvements on embedding capacity. To address this issue, this paper proposes a novel RDH-EI algorithm based on Bit-plane Classification and Adaptive Group Coding (hereafter BCAGC algorithm). The core contributions are as follows: (1) A bit-plane classification mechanism is proposed. It categorises bit-planes into simple or complex types according to the number of non-all-zero 4-bit sequence (NAZ-4BS) patterns. (2) An adaptive group coding scheme is proposed. It dynamically selects between fixed-length coding and Huffman coding based on bit-plane types and the frequency distribution of NAZ-4BS patterns, thereby achieving compact coding tables and efficient bit-plane compression. Experimental results demonstrate that the proposed BCAGC algorithm achieves average embedding rates of 4.1472, 4.0565 and 3.4465 bpp on the BOSSbase, BOWS-2 and UCID datasets, respectively, outperforming several state-of-the-art RDH-EI methods.</p>","PeriodicalId":56303,"journal":{"name":"IET Image Processing","volume":"20 1","pages":""},"PeriodicalIF":2.2,"publicationDate":"2026-02-24","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://ietresearch.onlinelibrary.wiley.com/doi/epdf/10.1049/ipr2.70310","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147653312","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Enhancing Diffusion Models Towards Anomaly-Aware Reconstructions for Medical Image Anomaly Detection","authors":"Wanying Wu, Xiaofeng Qu, Fenghang Zhang, Mengjiao Zhang, Xizhan Gao, Sijie Niu","doi":"10.1049/ipr2.70311","DOIUrl":"https://doi.org/10.1049/ipr2.70311","url":null,"abstract":"<p>Diffusion-based unsupervised anomaly detection in medical images has emerged as an effective paradigm, leveraging unlabelled healthy data to precisely characterize the distribution of normal anatomy and identify a wide range of pathological abnormalities. The method reconstructs a pseudo-healthy image from a potentially anomalous input and identifies anomalies by measuring pixel-wise reconstruction errors. However, existing approaches often preserve anomalous regions in the reconstruction, resulting in less prominent anomaly segmentation. Additionally, their inability to accurately restore normal areas can lead to increased false positives. In this work, we propose CS-Unet to advance this paradigm by realizing the concept of anomaly-aware reconstruction, defined as reconstructions that are consciously devoid of anomalies while faithfully restoring normal regions. Firstly, we propose a compression-expansion DenseNet (CompExDenseNet), which performs a dense cascade of nonlinear dimension transformations to extract compact feature representations, suppressing the reconstruction of anomalous patterns. Secondly, we design an attention gate (AG) unit to control the flow of low-frequency information, mitigating the leakage of anomalous information. Finally, we propose a frequency-domain adaptive residual convolution (FreAR) module that selectively enhances the most relevant frequency components to facilitate high-fidelity restoration of normal regions. Experimental results demonstrate that CS-Unet achieves outstanding performance in unsupervised anomaly detection, confirming its effectiveness.</p>","PeriodicalId":56303,"journal":{"name":"IET Image Processing","volume":"20 1","pages":""},"PeriodicalIF":2.2,"publicationDate":"2026-02-24","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://ietresearch.onlinelibrary.wiley.com/doi/epdf/10.1049/ipr2.70311","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147653313","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"SSGA-YOLO: A Lightweight Sonar Image Object Detection Network With Efficient Convolution and Acoustic-Aware Attention for Embedded Systems","authors":"Yan Liu, Gan Yan, Tong Chen, Guanying Huo","doi":"10.1049/ipr2.70313","DOIUrl":"https://doi.org/10.1049/ipr2.70313","url":null,"abstract":"<p>To address the problems of high computational complexity and limited deploy ability in underwater sonar image detection, we propose SSGA-YOLO, a lightweight and efficient object detection algorithm optimised for sonar imagery and successfully deployed on the Ascend AI embedded platform. SSGA-YOLO focuses on three key aspects to achieve a favourable balance between accuracy and efficiency. The S-Net backbone employs depthwise separable convolutions and a lightweight attention mechanism to enhance the extraction of weak echo and shadow features in sonar images while reducing redundancy for efficient deployment. The Efficient Group Shuffle Convolution (EGSConv) enhances cross-channel feature interaction to improve the detection of small, low-contrast sonar targets and the Lightweight Shuffle-Aware Group Attention (LSGA) refines key acoustic and spatial cues in the presence of strong noise. Furthermore, SSGA-YOLO significantly reduces model parameters and computational complexity: compared to YOLOv8n, it achieves reductions of 79.82% and 74.07% in parameter count and GFLOPs, respectively. To evaluate model performance across diverse environments, embedded deployment experiments were conducted on three datasets representing distinct scenarios: MDFD for controlled artificial tanks, UATD for complex natural waters and MOTfish for dynamic video sequences. SSGA-YOLO consistently achieves high detection accuracy, with an mAP50 exceeding 0.930 on all datasets and peaking at 0.983 on MDFD. In terms of inference efficiency, the model demonstrates exceptional real-time capability, reaching a frame rate of 65.77 FPS on MOTfish. These results outperform other lightweight detectors, confirming the model's effectiveness for practical underwater applications.</p>","PeriodicalId":56303,"journal":{"name":"IET Image Processing","volume":"20 1","pages":""},"PeriodicalIF":2.2,"publicationDate":"2026-02-23","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://ietresearch.onlinelibrary.wiley.com/doi/epdf/10.1049/ipr2.70313","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147288502","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Fang Wang, Pu Wang, Meng Zhao, Chenggang Shan, Zhen Yang
{"title":"The Power of Modality: Improving Polyp Segmentation With Multimodal Information","authors":"Fang Wang, Pu Wang, Meng Zhao, Chenggang Shan, Zhen Yang","doi":"10.1049/ipr2.70305","DOIUrl":"10.1049/ipr2.70305","url":null,"abstract":"<p>Accurate polyp segmentation from a single color image remains a significant challenge due to the complex appearance of lesions and the lack of diverse contextual priors. Existing methods usually rely on a limited image prior, which leads to suboptimal results. We propose a novel approach that leverages rich contextual information from multiple modalities (including depth, higher order semantics, and edge features) to produce precise segmentation results. By integrating the Depth Anything model, Large Language Models, and Edge Extractor, our method effectively fuses these diverse priors to overcome the limitations of single modality approaches. In addition, we introduce a diffusion modeling framework to bring powerful generative priors for endoscopic image segmentation. This is a flexible deep network architecture that efficiently fuses multimodal information and can accommodate any number of multimodal inputs. Extensive experimental results demonstrate that our approach achieves promising performance on several benchmarks.</p>","PeriodicalId":56303,"journal":{"name":"IET Image Processing","volume":"20 1","pages":""},"PeriodicalIF":2.2,"publicationDate":"2026-02-19","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://ietresearch.onlinelibrary.wiley.com/doi/epdf/10.1049/ipr2.70305","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147299942","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Wonhark Park, Jaehyun Lee, Wonsik Shin, Junhoo Lee, Nojun Kwak
{"title":"End-to-End Multi-Entity Customization","authors":"Wonhark Park, Jaehyun Lee, Wonsik Shin, Junhoo Lee, Nojun Kwak","doi":"10.1049/ipr2.70306","DOIUrl":"10.1049/ipr2.70306","url":null,"abstract":"<p>Recent advancements in text-to-image (T2I) models have enabled the synthesis of personalized images that align closely with user-specified prompts, especially through the use of modifiers. However, generating multiple detailed objects with distinct modifiers in a single image remains challenging due to concept-mixing, resulting from the difficulty of capturing interactions among text tokens. This paper proposes a modifier-based approach to mitigate concept-mixing by addressing the interaction among text tokens. Our method enables practical multi-personalization while preserving the original T2I model's straightforward inference pipeline. Without structural guidance, it ensures seamless object interaction with enhanced consistency. Through a loss-based finetuning approach, our method is adaptable to various concept-learning algorithms, enabling plug-and-play functionality. Through both qualitative and quantitative evaluations, we demonstrate that our method effectively resolves concept-mixing issues to better preserve concepts' identities and outperforms recent baselines in both quantitative and qualitative results. Our code will be publicly available.</p>","PeriodicalId":56303,"journal":{"name":"IET Image Processing","volume":"20 1","pages":""},"PeriodicalIF":2.2,"publicationDate":"2026-02-16","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://ietresearch.onlinelibrary.wiley.com/doi/epdf/10.1049/ipr2.70306","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"146217270","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Hassan Mahichi, Vahid Ghods, Mohammad Karim Sohrabi, Arash Sabbaghi
{"title":"Realtime Data Augmentation for Breast Cancer Dataset: Dynamic Fine-Tuning Bounding Box Coordinates and Segmentation Mask","authors":"Hassan Mahichi, Vahid Ghods, Mohammad Karim Sohrabi, Arash Sabbaghi","doi":"10.1049/ipr2.70308","DOIUrl":"10.1049/ipr2.70308","url":null,"abstract":"<p>Data augmentation is crucial for training deep learning models in breast cancer detection and segmentation, but conventional methods can cause annotation errors and visual artefacts, which harm instance-level localisation, generalisation and clinical reliability. This study aims to develop an annotation-aware, real-time data augmentation framework that preserves spatial and clinical integrity during geometric transformations to improve model robustness and performance. A real-time augmentation framework dynamically recalculates bounding box coordinates and segmentation masks during cropping and rotation. Instance-level annotation correction is performed on-the-fly within the training pipeline, without offline preprocessing, additional storage, or generative modelling and is evaluated on the DDSM, INbreast and BUSI datasets. Experimental results show consistent performance gains across all datasets and models, with detection metrics improving by an average of 4.2–4.4% and segmentation accuracy increasing by up to 9.3%. The real-time implementation achieves low preprocessing latency (≈0.12 s per batch, 3.75 ms per image), enabling high-throughput training without added computational overhead. By preserving annotation integrity during geometric transformations, the proposed framework provides a computationally efficient and easily integrable solution for breast cancer imaging, with broader applicability to other medical image analysis tasks requiring precise spatial annotations.</p>","PeriodicalId":56303,"journal":{"name":"IET Image Processing","volume":"20 1","pages":""},"PeriodicalIF":2.2,"publicationDate":"2026-02-12","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://ietresearch.onlinelibrary.wiley.com/doi/epdf/10.1049/ipr2.70308","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"146217041","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}