{"title":"KNN improved Transformer for 3D object detection","authors":"Chen Jiang, Shuxia Lu, Xianghu Zhou, Tingting Ma, Junhai Zhai","doi":"10.1016/j.image.2026.117488","DOIUrl":"10.1016/j.image.2026.117488","url":null,"abstract":"<div><div>In recent years, 3D object detection in autonomous driving perception has gained significant attention in the industry. Due to its characteristics, LiDAR has become the most commonly used and essential sensor. However, voxel-based networks often lose context information during the voxelization process, which negatively impacts the detection of small objects. In this paper, we address the challenge of low accuracy in LiDAR-based detection, especially for small object categories, by proposing an improved Transformer structure. Transformers are a type of deep learning model known for their ability to capture long-range dependencies and contextual relationships in data. In our approach, we incorporate a k-Nearest Neighbors (KNN) algorithm, which is a method for identifying the closest points in space, to enhance the spatial relationships between point clouds. This combination allows the model to better capture context information, strengthen feature extraction, and significantly reduce both missed and false detections. Our method is designed to be plug-and-play, allowing it to be directly applied to existing point cloud detectors. We evaluate our approach on the public KITTI and Astyx datasets. Experimental results show significant improvements, especially in detecting small object categories, even in challenging conditions.</div></div>","PeriodicalId":49521,"journal":{"name":"Signal Processing-Image Communication","volume":"142 ","pages":"Article 117488"},"PeriodicalIF":2.7,"publicationDate":"2026-03-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"146023542","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":3,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Cryptospace image steganography for cloud security via cycle-consistent GAN","authors":"Shuying Xu , Chin-Chen Chang , Ji-Hwei Horng , Ching-Chun Chang","doi":"10.1016/j.image.2026.117487","DOIUrl":"10.1016/j.image.2026.117487","url":null,"abstract":"<div><div>Cryptospace steganography has attracted increasing attention as an effective approach for enhancing data security in cloud environments. This paper proposes a hybrid framework that integrates cycle-consistent generative adversarial networks (CycleGAN) with the difference expansion (DE) technique to provide both image encryption and data hiding services. In the proposed framework, an image encryption network and a data encryption network are designed to encrypt the digital image and secret data, respectively, enabling a key-free architecture. A dropout-driven strategy is further introduced to support secure and isolated access control for multiple user groups on a shared cloud platform. Experimental results show that the proposed method achieves an embedding rate above 0.47 bpp and a secret data extraction accuracy exceeding 91%, demonstrating superior performance compared with state-of-the-art methods.</div></div>","PeriodicalId":49521,"journal":{"name":"Signal Processing-Image Communication","volume":"142 ","pages":"Article 117487"},"PeriodicalIF":2.7,"publicationDate":"2026-03-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"146023541","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":3,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"DPMMN: A dual performer-multi-modal network for emotion recognition","authors":"Shivanand S. Gornale , Shivakumara Palaiahnakote , Amruta Unki , Sunil Vadera","doi":"10.1016/j.image.2025.117464","DOIUrl":"10.1016/j.image.2025.117464","url":null,"abstract":"<div><h3>Background and Objective</h3><div>Although emotion recognition systems have been widely advocated, their accuracy can be affected when a person’s normal facial features overlap with their expressions when in a particular emotional state. This study, therefore, explores how heatmaps of electroencephalography (EEG) signals can be integrated with facial information to improve the accuracy of emotion recognition systems.</div></div><div><h3>Method</h3><div>The key idea of the proposed work is to fuse EEG signal heatmaps and Facial information for recognizing eight different emotions. For implementing this new idea, we propose a Dual Performer Multi-Modal Network (DPMMN). For each modality, the proposed work integrates modified Vision Transformer (ViT) and Long Short-Term Memory (LSTM). The integration is achieved by concatenating the features extracted from each modality and using them to classify the different emotions. In contrast to a baseline ViT, which uses self-attention layers, the proposed work replaces self-attention layers with the Performer layers through a kernelized attention approach. This results in extracting distinct visual features from EEG signal heatmaps and facial images. Similarly, for capturing temporal features from EEG heatmaps and Facial videos, the proposed LSTM replaces a traditional feed-forward network with a recurrent structure. This step helps to learn sequential dependencies across the patches.</div></div><div><h3>Results</h3><div>A comprehensive evaluation of DPMMN with respect to current state-of-the-art systems shows favorable results, with DPMN achieving 97.02 % in identifying eight distinct emotions on the DEAP benchmark dataset.</div></div><div><h3>Conclusion</h3><div>The proposed work shows that the use of EEG signal heatmap with facial information is better than EEG signal and facial information alone. Similarly, integrating performer layers with ViT and LSTM is better than existing models for extracting distinct features to classify eight emotions.</div></div>","PeriodicalId":49521,"journal":{"name":"Signal Processing-Image Communication","volume":"142 ","pages":"Article 117464"},"PeriodicalIF":2.7,"publicationDate":"2026-03-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"145842306","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":3,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"HOICNet: Low-Dose CT image denoising network based on higher-order feature attention mechanism and irregular convolution","authors":"Aimin Huang , Lina Jia , Beibei Jia , Zhiguo Gui , Jianan Liang","doi":"10.1016/j.image.2025.117457","DOIUrl":"10.1016/j.image.2025.117457","url":null,"abstract":"<div><div>Convolution Neural Networks (CNNs) with attention mechanisms show great potential for improving low-dose computed tomography (LDCT) image quality. However, most of these methods use first-order statistics for channel or space processing, ignoring the higher-order statistics of the channel or space features. In addition, the conventional convolution has limited receptive field and a poor performance on the edge of LDCT images. In this study, we aim to develop a CNN model incorporating higher-order feature attention mechanism that both enlarges the receptive field and clearly recovers edges and details. We propose an LDCT image denoising network named as HOICNet based on a higher-order feature attention mechanism and irregular convolution. Specifically, we first propose a new higher-order feature attention mechanism that utilizes higher-order feature statistics to enhance features in different channels and spatial regions. Second, we propose a new irregular convolutional feature extraction module (ICFE) that contains self-calibrating convolution (SC) and side window convolution (SWC). SC is used to enlarge receptive fields, and SWC is used to improve the edge information in denoised images. Finally, we introduce the contrast regularization mechanism (CRM) with positive and negative samples to bring the denoised image closer and closer to the positive samples while moving away from the negative samples to alleviate the problem of over-smoothing of the denoised images. Our experimental results show that the peak signal-to-noise ratio (PSNR), the structural similarity (SSIM), the root mean square error (RMSE) and the visual information fidelity (VIF) values achieved significant improvements in both the AAPM dataset and the piglet dataset.</div></div>","PeriodicalId":49521,"journal":{"name":"Signal Processing-Image Communication","volume":"142 ","pages":"Article 117457"},"PeriodicalIF":2.7,"publicationDate":"2026-03-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"145885409","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":3,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Hao Zhai, Anyu Li, Yan Wei, Huashan Tan, Yiyang Ru
{"title":"SACIFuse: Adaptive enhancement of salient features and cross-modal attention interaction for infrared and visible image fusion","authors":"Hao Zhai, Anyu Li, Yan Wei, Huashan Tan, Yiyang Ru","doi":"10.1016/j.image.2025.117467","DOIUrl":"10.1016/j.image.2025.117467","url":null,"abstract":"<div><div>The goal of infrared and visible light image fusion is to create images that highlight infrared thermal targets while preserving texture information under challenging lighting conditions. However, in extreme environments like heavy fog or overexposure, visible light images often contain redundant information, negatively affecting fusion results. To better emphasize salient targets in infrared images and reduce interference from redundant information, this paper proposes an adaptive salient enhancement fusion method for infrared and visible light images, called SACIFuse. First, we designed a Salient Feature Prediction Enhancement Module (SFEM), which extracts image gradients through edge operators and generates a mask quantifying the probability of redundant information. This mask is used to adaptively weight the source image, thereby suppressing redundant visible light information while enhancing infrared targets. Additionally, we introduced a Salient Feature Interaction Attention Module (SFIM), capable of employing residual attention combined with spatial and channel attention mechanisms to guide the interaction between the enhanced salient features and the source image features, ensuring that the fusion results highlight infrared targets while preserving visible light texture. Finally, our proposed loss function constructs a binary mask of the fused image to impose constraints on salient targets, effectively preventing adverse effects of redundant information on key regions. Extensive testing on public datasets shows that SACIfuse outperforms existing state-of-the-art methods in both qualitative and quantitative evaluations. Moreover, generalization experiments conducted on other datasets demonstrate that the proposed model exhibits strong generalization capabilities.</div></div>","PeriodicalId":49521,"journal":{"name":"Signal Processing-Image Communication","volume":"142 ","pages":"Article 117467"},"PeriodicalIF":2.7,"publicationDate":"2026-03-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"145927549","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":3,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Learned lossless medical image compression via dual transform and subimage-wise auto-regression","authors":"Tiantian Li , Yue Li , Ruixiao Guo , Gaobo Yang","doi":"10.1016/j.image.2025.117455","DOIUrl":"10.1016/j.image.2025.117455","url":null,"abstract":"<div><div>While learning-based compression has achieved significant success for natural images, its application to medical imaging requires specialized approaches to ensure zero information loss. In this work, we propose a novel lossless compression framework for 2D medical images, concentrated on two objectives: entropy reduction and precise probability estimation. We design a dual-transform module, namely DCT and a reversible linear map decomposition, which transforms image pixels into low-entropy representations — structural and textural components in integer DCT domain. A novel row-column grouping strategy is designed to decompose textural components into subimages. Unlike the existing parametric models, DCT-LSBNet models the non-parametric probability estimation for each subimage in an auto-regression way, avoiding the bias hypothesis of probability distribution and balancing well between compression rate and encoding latency. Extensive experimental results show that compared with the existing traditional and learned lossless compression, our method achieves state-of-the-art performance for X-ray images and ultrasound images.</div></div>","PeriodicalId":49521,"journal":{"name":"Signal Processing-Image Communication","volume":"142 ","pages":"Article 117455"},"PeriodicalIF":2.7,"publicationDate":"2026-03-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"145927658","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":3,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Mahamat Issa Choueb , Praveen Kumar Sekharamantry , Giulia Martinelli , Francesco De Natale , Nicola Conci
{"title":"NFlowAD: A normalizing flow model for anomaly detection in human motion animations","authors":"Mahamat Issa Choueb , Praveen Kumar Sekharamantry , Giulia Martinelli , Francesco De Natale , Nicola Conci","doi":"10.1016/j.image.2025.117469","DOIUrl":"10.1016/j.image.2025.117469","url":null,"abstract":"<div><div>Anomaly detection has been extensively investigated in numerous application areas. Hand-crafted rules have gradually given way to supervised classification techniques, which frequently rely on a small number of anomaly labels and related architectures. When it comes to human motion, abnormalities emerge at a fine-grained temporal or joint level rather than over a whole video sequence.</div><div>This study introduces NFlowAD, a self-supervised system that analyzes body joints to detect irregularities in human motion. It blends normalizing flows with masked motion modeling to describe normal motion data without the need for anomaly labels. Inference uses both reconstruction mistakes and flow-based likelihoods to detect anomalies. The validation pipeline on various state-of-the-art datasets demonstrates NFlowAD’s efficiency in recognizing, locating, and analyzing anomalous motion sequences, while maintaining robust detection and interpretability.</div></div>","PeriodicalId":49521,"journal":{"name":"Signal Processing-Image Communication","volume":"142 ","pages":"Article 117469"},"PeriodicalIF":2.7,"publicationDate":"2026-03-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"145842321","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":3,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Walter Brescia , Pedro Gomes , Laura Toni , Saverio Mascolo , Luca De Cicco
{"title":"GT-MilliNoise: Graph transformer for point-wise denoising of indoor millimetre-wave point clouds","authors":"Walter Brescia , Pedro Gomes , Laura Toni , Saverio Mascolo , Luca De Cicco","doi":"10.1016/j.image.2025.117453","DOIUrl":"10.1016/j.image.2025.117453","url":null,"abstract":"<div><div>Millimetre-wave (mmWave) radars are gaining popularity thanks to their low cost and robustness in low-visibility conditions. However, the 3D point clouds they produce are sparser and noisier than those from LiDARs and depth cameras. These differences create challenges when applying existing methods, originally designed for dense point clouds, to mmWave data. Specifically, there is a gap in point-level precision tasks, such as full point cloud denoising for mmWave data, partly due to the lack of fully annotated datasets. In this work, we employ the MilliNoise dataset, a fully annotated indoor mmWave point clouds dataset, to advance the understanding of mmWave point clouds denoising via two main steps: (i) we carry out an experimental analysis of the most common point cloud processing approaches and show their limitations in exploring the local-to-global structures in sparse and noisy point clouds; (ii) in light of the identified limitations, we propose a graph-based transformer architecture, denoted as GT-<em>MilliNoise</em>, composed of two main blocks to effectively leverage both the temporal and geometric structures of the data: a <em>Temporal</em> block leverages the sparsity of data to learn the dynamic behaviour of the points; a <em>Geometric</em> block, uses a point-wise attention mechanism to form representative neighbourhoods for feature extraction. The experimental results obtained in the MilliNoise dataset show that our proposed GT-<em>MilliNoise</em> architecture outperforms the state-of-the-art both qualitatively and quantitatively. Specifically, it achieves 75% accuracy (5% gain compared to the state-of-the-art), and a significantly low Earth Mover’s distance value of 0.193.</div></div>","PeriodicalId":49521,"journal":{"name":"Signal Processing-Image Communication","volume":"142 ","pages":"Article 117453"},"PeriodicalIF":2.7,"publicationDate":"2026-03-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"146023540","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":3,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Wan Li , Hengji Xie , Bin Yao , Xiaolin Zhang , Rongrong Fei
{"title":"Low-light image enhancement via boundary constraints and non-local similarity","authors":"Wan Li , Hengji Xie , Bin Yao , Xiaolin Zhang , Rongrong Fei","doi":"10.1016/j.image.2025.117459","DOIUrl":"10.1016/j.image.2025.117459","url":null,"abstract":"<div><div>Low light often leads to poor image visibility, which can easily affect the performance of computer vision algorithms. The traditional enhancement methods focus excessively on illumination map restoration while neglecting the non-local similarity in natural images. In this paper, we propose an effective low-light image enhancement method based on boundary constraints and non-local similarity. First, a fast and effective boundary constraints method is proposed to estimate illumination maps. Then, a combined optimization model with low-rank and context constraints was presented to improve the enhancing results. Among them, low-rank constraints are used to capture the non-local similarity of the reflectance image, and context constraints are used to improve the accuracy of the illumination map. Finally, alternating iterative optimization is employed for solving non-independent constraints between the illumination and reflectance maps. Experimental results demonstrate that the proposed algorithm enhances images efficiently in terms of both objective quality and subjective quality.</div></div>","PeriodicalId":49521,"journal":{"name":"Signal Processing-Image Communication","volume":"142 ","pages":"Article 117459"},"PeriodicalIF":2.7,"publicationDate":"2026-03-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"145885406","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":3,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Jing Li , Tao Chen , Xiangyu Han , Xilin Luan , Jintao Li
{"title":"Multi-Frame Adaptive Image Enhancement Algorithm for vehicle-mounted dynamic scenes","authors":"Jing Li , Tao Chen , Xiangyu Han , Xilin Luan , Jintao Li","doi":"10.1016/j.image.2025.117458","DOIUrl":"10.1016/j.image.2025.117458","url":null,"abstract":"<div><div>To address the issue of image blurring caused by high-speed vehicle motion and complex road conditions in autonomous driving scenarios, this paper proposes a lightweight Multi-frame Adaptive Image Enhancement Network (MAIE-Net). The network innovatively introduces a hybrid motion compensation mechanism that integrates optical flow alignment and deformable convolution, effectively solving the non-rigid motion alignment problem in complex dynamic scenes. Additionally, a temporal feature enhancement module is constructed, leveraging 3D convolution and attention mechanisms to achieve adaptive fusion of multi-frame information. In terms of architecture design, an edge-guided U-Net structure is employed for multi-scale feature extraction and reconstruction. The framework incorporates edge feature extraction and attention mechanisms within the encoder–decoder to balance feature representation and computational efficiency. The overall lightweight design enables the model to adapt to in-vehicle computing platforms. Experimental results demonstrate that the proposed method significantly improves image quality while maintaining efficient real-time processing capabilities, effectively enhancing the environmental perception performance of in-vehicle vision systems across various driving scenarios, thereby providing a reliable visual enhancement solution for autonomous driving.</div></div>","PeriodicalId":49521,"journal":{"name":"Signal Processing-Image Communication","volume":"142 ","pages":"Article 117458"},"PeriodicalIF":2.7,"publicationDate":"2026-03-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"145977991","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":3,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}