Neural NetworksPub Date : 2026-09-02DOI: 10.1016/j.neunet.2026.109583
Yuping Zhang, Yan Liu
{"title":"SynNeura: Event-driven liquid-spiking dynamics for weakly supervised multimodal temporal alignment.","authors":"Yuping Zhang, Yan Liu","doi":"10.1016/j.neunet.2026.109583","DOIUrl":"https://doi.org/10.1016/j.neunet.2026.109583","url":null,"abstract":"<p><p>Multimodal temporal alignment is a critical task for applications such as audiovisual understanding, lip reading, and instruction following. However, in weakly supervised settings, challenges like asynchronous sampling, irregular event triggers, and coarse labels hinder precise cross-modal alignment. Frame-based methods rely on fixed time grids, leading to redundant computation in sparse-event scenarios and reducing event-level precision. Differentiable time-warping methods typically require high-resolution inputs, resulting in high computational costs and sensitivity to numerical parameters. To address these challenges, we propose SynNeura, an event-driven continuous-time liquid-spiking neural framework for fine-grained alignment under weak supervision. SynNeura models alignment as a continuous-time latent-state process, analytically propagating states between events and updating only when spikes occur. This results in computational complexity that scales with the number of events rather than the sequence length. SynNeura introduces a piecewise-analytic update using matrix exponentials and trace variables, along with a hierarchical contrastive alignment objective at spike, trajectory, and state levels, enhancing robustness and consistency. Experiments with the AVE, LRS2, and YouCook2 datasets show SynNeura consistently outperforms frame-based and continuous-time baselines in alignment accuracy, temporal consistency, and efficiency. SynNeura achieves 0.291 MAE and 0.713 CAS on AVE, 0.648 TC on LRS2, and 0.829/0.794 EP/ER on YouCook2, with an overall score of 0.738. These results show SynNeura is a scalable, interpretable, and efficient solution for event-driven multimodal temporal alignment.</p>","PeriodicalId":49763,"journal":{"name":"Neural Networks","volume":"205 Pt C","pages":"109583"},"PeriodicalIF":7.2,"publicationDate":"2026-09-02","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148898268","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":1,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Neural NetworksPub Date : 2026-09-02DOI: 10.1016/j.neunet.2026.109581
Xi Yang, Hexun Zhou, Haiyang Zhu, Nannan Wang
{"title":"Multimodal-guided self-distillation for unified person search.","authors":"Xi Yang, Hexun Zhou, Haiyang Zhu, Nannan Wang","doi":"10.1016/j.neunet.2026.109581","DOIUrl":"https://doi.org/10.1016/j.neunet.2026.109581","url":null,"abstract":"<p><p>Person search is challenging due to limitations in identity representation. Existing methods rely on one-hot encoding, ignoring semantic relationships among pedestrians. This leads to a fragmented feature space and reduces generalization ability, especially in large-scale scenarios with a significant proportion of unlabeled identities. For instance, in the CUHK-SYSU dataset, 72.7% of pedestrians lack identity annotations, limiting the effectiveness of supervised learning. To address these issues, we propose a novel Multimodal-Guided Self-Distillation (MGSD) method for Unified Person Search that leverages multimodal textual descriptions and self-distillation to enhance pedestrian representation learning. Specifically, we introduce three key innovations: (1) Multimodal LLM-Assisted Text Generation (MLTG) to provide fine-grained semantic context beyond discrete identity labels, enabling the model to capture inter-person relationships based on clothing attributes, appearance features, and environmental cues; (2) Semantic Structural Consistency Constraint (SSCC) to impose global structural constraints on the feature space, ensuring that distinct identities remain separable while preserving semantic similarities among visually similar individuals; and (3) Multimodal-Aware Self-Distillation Framework (MSDF), where the learnable visual encoder is progressively aligned with the pre-trained CLIP multimodal encoder, improving robustness to variations in illumination, occlusion, and background clutter. Extensive experiments demonstrate that our method significantly enhances retrieval accuracy and generalization, achieving state-of-the-art performance with an mAP of 56.1% on the PRW dataset while maintaining computational efficiency for large-scale real-world applications.</p>","PeriodicalId":49763,"journal":{"name":"Neural Networks","volume":"205 Pt C","pages":"109581"},"PeriodicalIF":7.2,"publicationDate":"2026-09-02","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148898284","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":1,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Neural NetworksPub Date : 2026-09-01DOI: 10.1016/j.neunet.2026.109569
Shengbing Tang, Chen Xiao, Bin He, Na Li
{"title":"Gaussian processes with prior-model-informed kernel for dynamical system modeling.","authors":"Shengbing Tang, Chen Xiao, Bin He, Na Li","doi":"10.1016/j.neunet.2026.109569","DOIUrl":"https://doi.org/10.1016/j.neunet.2026.109569","url":null,"abstract":"<p><p>Gaussian processes (GPs) are widely used for modeling dynamical systems due to their ability to provide principled uncertainty estimation and to incorporate prior knowledge via the mean or kernel function. However, GPs typically exhibit poor extrapolation performance outside the training region. To address this limitation while preserving the favorable uncertainty behavior of standard GPs, we propose a novel kernel design that integrates prior models-either analytical physical models or imperfect simulators-into the kernel function of the GP. Our approach is motivated by the assumption that similarity in the prior model implies similarity in the true system dynamics. Specifically, we transform the input space using the prior model and apply a base kernel (e.g., Radial Basis Function or Matérn) to construct a prior-model-informed kernel that reflects this assumption. To further enhance modeling flexibility, we add a standard residual kernel to correct discrepancies between the prior model and the true system. This yields our final model: a GP with prior-model-informed kernel (GP-PI-K). Through extensive experiments on benchmark dynamical systems, we demonstrate that GP-PI-K consistently outperforms existing baselines, including standard GPs, GPs with physics-informed or neural network mean functions, and deep kernel learning models, in terms of one-step and multi-step prediction accuracy, uncertainty estimation, and uncertainty calibration. Moreover, GP-PI-K achieves superior performance in downstream applications such as active learning and model-based reinforcement learning.</p>","PeriodicalId":49763,"journal":{"name":"Neural Networks","volume":"205 Pt C","pages":"109569"},"PeriodicalIF":7.2,"publicationDate":"2026-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148898293","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":1,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Corrigendum to \"BAED: a New Paradigm for Few-shot Graph Learning with Explanation in the Loop\" [Neural Networks, 199 (2026), 108673].","authors":"Chao Chen, Xujia Li, Dongsheng Hong, Shanshan Lin, Xiangwen Liao, Chuanyi Liu, Lei Chen","doi":"10.1016/j.neunet.2026.109116","DOIUrl":"10.1016/j.neunet.2026.109116","url":null,"abstract":"","PeriodicalId":49763,"journal":{"name":"Neural Networks","volume":"201 ","pages":"109116"},"PeriodicalIF":7.2,"publicationDate":"2026-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148007239","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":1,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Neural NetworksPub Date : 2026-09-01DOI: 10.1016/j.neunet.2026.109560
Yinfeng Zeng, Xuesong Wang, Yuhu Cheng
{"title":"Similarity-guided state attention for visual reinforcement learning.","authors":"Yinfeng Zeng, Xuesong Wang, Yuhu Cheng","doi":"10.1016/j.neunet.2026.109560","DOIUrl":"https://doi.org/10.1016/j.neunet.2026.109560","url":null,"abstract":"<p><p>Visual reinforcement learning (VRL) aims to extract effective visual information from high-dimensional observations to optimize decision-making policies. While existing VRL methods have achieved significant progress in various control tasks through data augmentation and auxiliary tasks, agents remain susceptible to distractions from redundant information and irrelevant factors, leading to overfitting and degraded generalization in unseen environments. To address this challenge, we propose a similarity-guided state attention (SSA) for VRL. In this method, a similarity guidance module (SGM) is designed to leverage the similarity between state embeddings from the original and augmented observations to guide the encoder to focus on task-relevant regions in the original observation, thereby producing a corresponding state attention map. Meanwhile, the state attention map of augmented observation is obtained by decoding its state embedding. Furthermore, cosine similarity is introduced to measure the global similarity between state attention maps of original and augmented observations, which is incorporated into self-supervised learning objective together with binary cross-entropy loss to encourage alignment of state attention maps in representation space. The combination of SGM and cosine similarity-based alignment of state attention maps facilitates self-supervised learning to obtain more robust state representations for downstream reinforcement learning, thereby enabling the agent to learn optimal policies and improve its generalization ability in unseen environments. Experimental results on the DeepMind Control Generalization Benchmark (DMControl-GB) demonstrate that SSA achieves superior robustness and generalization performance compared with representative VRL baselines.</p>","PeriodicalId":49763,"journal":{"name":"Neural Networks","volume":"205 Pt C","pages":"109560"},"PeriodicalIF":7.2,"publicationDate":"2026-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148892746","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":1,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Enhancing knowledge tracing with multi-level individualized perception and teacher-student semantic distillation.","authors":"Zhenqiang Yu, Luyao Huang, Xingbing Li, Yuncheng Jiang","doi":"10.1016/j.neunet.2026.109576","DOIUrl":"https://doi.org/10.1016/j.neunet.2026.109576","url":null,"abstract":"<p><p>Knowledge tracing aims to model students' dynamic knowledge states based on their historical learning interactions and predict future learning performance. Existing sequence modeling methods often overlook individual differences and have limited semantic modeling capability. To address these issues, this paper proposes a knowledge tracing model that combines personalized modeling with knowledge distillation. The model introduces three personalized modules: (1) a personalized question understanding module that captures individual differences in students' understanding of the same question; (2) a personalized question-knowledge association module that models relationships between questions and relevant knowledge concepts; and (3) a personalized knowledge state forgetting module that simulates students' memory decay patterns. These modules allow for more accurate modeling of students' dynamic knowledge states. Furthermore, to overcome the semantic limitations of lightweight models, a large language model (LLM) is used as the teacher, and its semantic modeling capability is transferred to an LSTM-based student model via knowledge distillation. Experiments show that the proposed method consistently improves prediction performance on two benchmark datasets, demonstrating its effectiveness in modeling personalized learning and enhancing semantic representation.</p>","PeriodicalId":49763,"journal":{"name":"Neural Networks","volume":"205 Pt C","pages":"109576"},"PeriodicalIF":7.2,"publicationDate":"2026-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148898329","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":1,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Patient-independent seizure onset zone localization with generalizable feature learning and multi-task supervision.","authors":"Jinjie Guo, Tao Feng, Yiping Wang, Yanfeng Yang, Guixia Kang, Guoguang Zhao","doi":"10.1016/j.neunet.2026.109538","DOIUrl":"https://doi.org/10.1016/j.neunet.2026.109538","url":null,"abstract":"<p><p>Most existing ictal stereoelectroencephalography (SEEG)-based seizure onset zone (SOZ) localization methods rely on patient-specific training, limiting their clinical applicability due to the scarcity of seizure recordings and substantial inter-patient variability. Consequently, robust patient-independent SOZ localization remains a major challenge. In this work, we propose a deep learning approach for patient-independent SOZ localization using ictal SEEG recordings, aiming to improve cross-patient generalization while preserving seizure-related temporal characteristics. To mitigate domain shifts across subjects, we introduce a clinically guided feature learning strategy that combines a cross-frequency coupling (CFC) mechanism to capture SOZ-related abnormal interactions across frequency bands with a self-comparison (SC) mechanism to emphasize seizure-onset evolution patterns within SEEG channels. We further incorporate seizure detection as an auxiliary task within a multi-task learning framework to provide seizure-onset-related temporal supervision, thereby improving the temporal awareness and generalizability of SOZ localization. Experiments on the public OpenNeuro HUP dataset demonstrate substantial improvements over existing methods, while additional evaluations on a private clinical dataset further validate the robustness and cross-patient generalization capability of the proposed method. Moreover, comparisons between the learned CFC representations and clinically established phase-amplitude coupling (PAC) metrics reveal consistent physiological patterns, supporting the interpretability of the learned representations.</p>","PeriodicalId":49763,"journal":{"name":"Neural Networks","volume":"205 Pt C","pages":"109538"},"PeriodicalIF":7.2,"publicationDate":"2026-08-31","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148867574","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":1,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Neural NetworksPub Date : 2026-08-31DOI: 10.1016/j.neunet.2026.109571
Xinyi Guo, Yiwei Lu, Tao Yan
{"title":"Robust webly supervised fine-grained recognition via decoupled global-local fusion and geometric-semantic consensus.","authors":"Xinyi Guo, Yiwei Lu, Tao Yan","doi":"10.1016/j.neunet.2026.109571","DOIUrl":"https://doi.org/10.1016/j.neunet.2026.109571","url":null,"abstract":"<p><p>Webly Supervised Fine-Grained Recognition demands robust learning mechanisms to mitigate the impact of noisy labels and misleading annotations. To address this, current methods typically rely on prediction probabilities to filter noisy labels. However, such paradigms often struggle with mislabeled yet high-confidence samples, leading to confirmation bias. In this paper, we propose a novel framework named Decoupled Global-local Consensus Learning (DGCL) that identifies and leverages noisy-labeled samples through multi-perspective analysis, enhancing both robustness and representation power. Motivated by the high-quality dense representations of foundation models, we first introduce the Decoupled Global-Local Fusion (DGLF) framework, which leverages LoRA to calibrate the encoder's attention, enabling precise identification of discriminative regions. Subsequently, a dual-stream bilinear mechanism is designed to facilitate global-local interactions for enhanced detail capture. To mitigate confirmation bias, we introduce the Geometric-Semantic Consensus (GSC) strategy, which integrates geometric consistency and semantic confidence to partition samples into three mutually exclusive subsets. Finally, a Noise-Aware Supervised Contrastive (NASC) loss is introduced to convert filtered noise into repulsion signals, effectively compressing intra-class variance. Extensive experiments demonstrate that DGCL achieves a new state-of-the-art average accuracy of 91.73% on three web-supervised benchmark datasets, surpassing the previous best method by 1.63 percentage points. Our source code will be made publicly available at: https://github.com/YT3DVision/DGLF.</p>","PeriodicalId":49763,"journal":{"name":"Neural Networks","volume":"205 Pt C","pages":"109571"},"PeriodicalIF":7.2,"publicationDate":"2026-08-31","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148892795","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":1,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Neural NetworksPub Date : 2026-08-31DOI: 10.1016/j.neunet.2026.109578
Anzhi Wang, Jintao Wu, Yun Liu
{"title":"Lightweight attention-aware fusion network based on state-space model for V-D-T salient object detection.","authors":"Anzhi Wang, Jintao Wu, Yun Liu","doi":"10.1016/j.neunet.2026.109578","DOIUrl":"https://doi.org/10.1016/j.neunet.2026.109578","url":null,"abstract":"<p><p>With the popularization of various sensors, multimodal Salient Object Detection (SOD) methods including Visual-Depth-Thermal Salient Object Detection (V-D-T SOD) have made remarkable developments. However, most existing V-D-T SOD methods usually achieve accurate detection with expensive computational cost, which limits their development at the edge application. To address this issue, we develop a Lightweight Attention-Aware Fusion Network (LAANet) based on the State Space Model (SSM). Specifically, inspired by the cross-attention mechanism, the Cross Mamba Fusion Module (CMFM) is proposed to realize attention perception among shallow high-resolution multimodal features by exploiting SSM with linear complexity. For low-resolution deep features, Attention Perception Fusion Module (APFM) is proposed based on the self-attention mechanism to mine semantic cues and complement the potential information lost in SSM amnestic memory. In addition, to response the problem of excessive decoding parameters and efficiently achieve feature reconstruction,we design a lightweight decoder (LDB). It is consist of simple dilated convolutions and linear operations with only 0.37M parameter. Extensive experiments on the VDT2048 dataset show that our method achieves performance close to that of SOTA, while having smaller number of parameters (6.78M), lower model complexity (5.23G), and faster inference speed (23.7FPS when the input size is 384*384). The code is available at https://github.com/GZNU-WJT/LAANet.</p>","PeriodicalId":49763,"journal":{"name":"Neural Networks","volume":"205 Pt C","pages":"109578"},"PeriodicalIF":7.2,"publicationDate":"2026-08-31","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148892824","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":1,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Neural NetworksPub Date : 2026-08-31DOI: 10.1016/j.neunet.2026.109570
Jinchang Xu, Xiangji Guo, Guifan Zhang, Fei Xie, Ming Ming
{"title":"SFMambaSR: A spatial-frequency enhanced Mamba network for wafer image super-resolution.","authors":"Jinchang Xu, Xiangji Guo, Guifan Zhang, Fei Xie, Ming Ming","doi":"10.1016/j.neunet.2026.109570","DOIUrl":"https://doi.org/10.1016/j.neunet.2026.109570","url":null,"abstract":"<p><p>Wafer defect inspection is crucial for yield and reliability, but shrinking defect sizes demand higher imaging resolution. While high-magnification optics provide resolution, their narrow field of view limits inspection efficiency. To balance precision and throughput, we propose a solution that reconstructs high-resolution wafer images from large-field low-magnification captures via a super-resolution algorithm. This method can improve detection efficiency without affecting accuracy. We design a dual-domain fusion lightweight SR network (SFMambaSR) specifically for wafer microscopy images. In the spatial domain, a Visual State Space Model (VSSM) and Multi-Scale Feature Extraction (MSFE) module jointly fuse global and local representations, while in the frequency domain, a wavelet-based Frequency-Domain Transformation (FDT) module enhances high-frequency defect details. Experiments on our large-scale wafer microscopy dataset demonstrate that SFMambaSR achieves the best PSNR and competitive or best SSIM across 2 × , 3 × , and 4 × upscaling factors, while using only 807K parameters.</p>","PeriodicalId":49763,"journal":{"name":"Neural Networks","volume":"205 Pt C","pages":"109570"},"PeriodicalIF":7.2,"publicationDate":"2026-08-31","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148898273","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":1,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}