Zhouwei Lin, Yudong Xia, Pingping Wu, Han Lin, Qingsong Xu, Zheng Fan, Juan Tu, Jun Wu
{"title":"A Lightweight Hybrid Network for Medical Image Segmentation With Adaptive Feature Selection","authors":"Zhouwei Lin, Yudong Xia, Pingping Wu, Han Lin, Qingsong Xu, Zheng Fan, Juan Tu, Jun Wu","doi":"10.1049/cit2.70168","DOIUrl":"https://doi.org/10.1049/cit2.70168","url":null,"abstract":"<p>Accurate medical image segmentation with low model complexity remains difficult because lesions are often small in scale and boundary cues are easily corrupted by noise. Although recent segmentation methods have achieved strong performance, many of them rely on increasingly complex architectures with high computational costs, limiting their applicability in resource-constrained settings. To alleviate this issue, we propose PLNet, a lightweight hybrid segmentation framework for medical image analysis. The network improves representation learning by integrating fine-grained local-structure perception with long-range contextual modelling, thereby enhancing the representation of complex medical images. In addition, a feature selection mechanism is employed to emphasise informative feature responses and suppress redundant activations while preserving discriminative representations. Extensive experiments on multiple datasets demonstrate that PLNet achieves competitive IoU and Dice scores with relatively low model complexity. Overall, these results highlight the potential of PLNet as an effective lightweight framework that balances segmentation performance and computational complexity.</p>","PeriodicalId":46211,"journal":{"name":"CAAI Transactions on Intelligence Technology","volume":"11 4","pages":"1047-1061"},"PeriodicalIF":5.1,"publicationDate":"2026-08-31","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://onlinelibrary.wiley.com/doi/epdf/10.1049/cit2.70168","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148859964","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Edge-Channel Aggregation Network and Two-Stage Fine Tuning Scheme for Handwritten Dongba Character Recognition","authors":"Xiali Li, Yang Xiao, Xiaojun Bi, Yanlong Luo","doi":"10.1049/cit2.70161","DOIUrl":"https://doi.org/10.1049/cit2.70161","url":null,"abstract":"<p>Handwritten Dongba Character Recognition (HDCR) contains a large number of visually similar characters with subtle and fragile edge cues, posing severe challenges to feature learning. To address this issue, an Edge Channel Aggregation Network (EdgeCANet) model is proposed. The core of EdgeCANet is the Edge Channel Attention Aggregation (ECAA) block, which integrates the Edge-Guided Enhancement Module (EGEM) for edge-aware spatial refinement and the Channel Texture Recalibration Module (CTRM) for channel-wise texture recalibration. Besides, the two-stage misclassified sample weight redistribution fine-tuning scheme based on Exponential Moving Average (EMA) algorithm is proposed to stabilise parameter updates, and to emphasise difficult and highly similar classes. The recognition accuracy of EdgeCANet significantly outperforms state-of-the-art models across the constructed high-similarity datasets HS-DB and HS-S20 K, as well as the public DB1404, OBC306, Sketch-20 K, and ImageNet-Sketch datasets. The results validate that EdgeCANet effectively enhances recognition accuracy for highly similar handwritten characters and ancient minority-script characters, offering more promising applicability for related recognition tasks.</p>","PeriodicalId":46211,"journal":{"name":"CAAI Transactions on Intelligence Technology","volume":"11 4","pages":"1254-1273"},"PeriodicalIF":5.1,"publicationDate":"2026-08-31","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://onlinelibrary.wiley.com/doi/epdf/10.1049/cit2.70161","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148859979","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Brain-Mimetic Mapless Navigation Framework Integrating Visual Streams and Entorhinal-Hippocampal-Prefrontal Circuits","authors":"Yishen Liao, Naigong Yu, Hejie Yu, Shufei Fu","doi":"10.1049/cit2.70167","DOIUrl":"https://doi.org/10.1049/cit2.70167","url":null,"abstract":"<p>Efficient autonomous navigation in complex and unstructured environments without pre-existing maps remains a significant challenge for mobile robotics. Drawing inspiration from rodent neural architectures, this study proposes a brain-mimetic mapless navigation framework for mobile robots. It integrates three key components to achieve autonomous navigation: first, a visual stream-based localisation model utilising object-vector cells to compensate for cumulative path integration errors; second, a navigation-guiding model where the hippocampal-prefrontal circuitry outputs initial paths through exploration and learning, which are then dynamically optimised for obstacle avoidance through a self-organising strategy of hippocampal CA1 place cells incorporating boundary-vector cells' inputs; and finally, the consolidation of these optimised paths into stable navigation habits via spatial cells' theoretical firing rate and spike-timing-dependent plasticity learning rule. Extensive 2D and 3D simulation experiments demonstrate that the proposed framework not only outperforms various baseline algorithms in terms of convergence speed and path efficiency but also exhibits robust error correction under motion noise, rapid habit reshaping during task changes and strong invariance to the initial heading direction. Moreover, real-world experiments on a mobile robot in complex indoor environments further confirm its practical feasibility and effectiveness, which offers a biologically plausible and computationally efficient navigation solution.</p>","PeriodicalId":46211,"journal":{"name":"CAAI Transactions on Intelligence Technology","volume":"11 4","pages":"1231-1253"},"PeriodicalIF":5.1,"publicationDate":"2026-08-31","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://onlinelibrary.wiley.com/doi/epdf/10.1049/cit2.70167","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148859953","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Jiaye Wei, Xin Xu, Yixing Lan, Tenglong Liu, Yueying Wang
{"title":"AGT: Efficient Offline Reinforcement Learning With Advantage-Guided Transformer","authors":"Jiaye Wei, Xin Xu, Yixing Lan, Tenglong Liu, Yueying Wang","doi":"10.1049/cit2.70094","DOIUrl":"https://doi.org/10.1049/cit2.70094","url":null,"abstract":"<p>Offline reinforcement learning (RL) is a paradigm that seeks to train policies directly based on fixed datasets derived from previous interactions with the environment. However, offline RL faces critical challenges in environments characterised by sparse rewards and datasets dominated by suboptimal trajectories. In this article, we propose a simple and efficient offline RL framework via advantage-guided transformer (AGT). Different from previous works, AGT promotes an advantage-guided policy constraint into the loss function during policy training that adaptively learns the features of trajectories with high advantage values. Furthermore, AGT samples actions with high advantage values during action generation, replacing manual designs with advantage values to update the return-to-go (RTG) as conditional inputs, to ensure action consistency with expected returns. Theoretical analysis shows that AGT's policy improvement guarantees convergence to near-optimal policies in environments with low stochasticity. Extensive experiments on D4RL benchmarks, including robotic locomotion control, dexterous manipulation and maze navigation tasks, show AGT outperforms state-of-the-art (SOTA) offline RL methods.</p>","PeriodicalId":46211,"journal":{"name":"CAAI Transactions on Intelligence Technology","volume":"11 4","pages":"1127-1143"},"PeriodicalIF":5.1,"publicationDate":"2026-08-31","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://onlinelibrary.wiley.com/doi/epdf/10.1049/cit2.70094","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148859974","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"CDFNet: Cross-Modal Deep Fusion for Monocular 3D Semantic Scene Completion","authors":"Xianjing Cheng, Lintai Wu, Jie Wen, Tianhao Peng, Yong Xu, Zairong Wei","doi":"10.1049/cit2.70124","DOIUrl":"https://doi.org/10.1049/cit2.70124","url":null,"abstract":"<p>Semantic scene completion (SSC) aims to predict the semantic occupancy and geometry of 3D scenes. Recently, most studies focus on camera-based approaches due to the rich visual cues of images and the cost-effectiveness of cameras. However, these methods usually lack efficient fusion and fine-grained processing of cross-modal semantic information, resulting in sub-optimal performance. To address these issues, we propose a novel cross-modal semantic deep fusion framework. Unlike previous approaches, our method effectively integrates 2D textural, 2D spatial and 3D geometric knowledge from three different modalities to reconstruct complete 3D scenes. Specifically, we employ two encoders to extract 2D textural and 2D spatial features from RGB images and depth maps, which are then fused and lifted into 3D space via our tailored cross-modal semantic fusion module. In contrast to previous methods that encompass extensive redundant 3D voxel features, we design a lightweight voxel feature filter to efficiently eliminate these redundancies. Furthermore, 3D geometric features are extracted from the point cloud derived from the depth map. The 3D features from multiple modalities are deeply fused and further refined by a sparse-to-dense voxel completion module, which effectively enriches the semantic information. Besides, we propose a new evaluation metric more suitable for assessing the SSC task with class imbalance issues in the dataset. Extensive experiments show that our method achieves the state of the art in camera-based semantic scene completion. We will release the source code publicly.</p>","PeriodicalId":46211,"journal":{"name":"CAAI Transactions on Intelligence Technology","volume":"11 4","pages":"1161-1178"},"PeriodicalIF":5.1,"publicationDate":"2026-08-31","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://onlinelibrary.wiley.com/doi/epdf/10.1049/cit2.70124","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148860212","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Multi-Strategy Synergy-Based Differential Evolution for Large Wind Farm Layout Optimization","authors":"Haotian Li, Yifei Yang, Qiong Fu, Zhenyu Lei, Zhe Xu, Shangce Gao","doi":"10.1049/cit2.70150","DOIUrl":"https://doi.org/10.1049/cit2.70150","url":null,"abstract":"<p>Exploring clean energy alternatives to fossil fuels has become a major research focus worldwide. Wind energy is regarded as one of the most promising renewable energy sources due to its clean and sustainable nature. However, wake interactions among turbines significantly reduce power conversion efficiency, making wind farm layout optimization (WFLO) crucial for maximising energy output. As the number of turbines increases, wake interactions become more pronounced, further deteriorating overall efficiency. Metaheuristic algorithms have been widely adopted to address the complex constraints and design objectives of WFLO. Nevertheless, traditional heuristic methods often suffer from poor solution quality and premature convergence when dealing with large-scale WFLO problems under complex wind conditions. To overcome these limitations, this study proposes a multi-strategy synergy-based differential evolution algorithm (LSDE) for large-scale WFLO under complex wind scenarios. The proposed method is evaluated against nine representative WFLO algorithms under four complex wind scenarios (4, 5, 6 and 7 wind directions) and three turbine scales (30, 50 and 100 turbines). Experimental results demonstrate that LSDE achieves superior performance, stability and robustness. Specifically, LSDE improves power conversion efficiency by 3.40%, 3.51%, 3.10% and 3.36% under the four wind scenarios, respectively, compared with other mainstream algorithms.</p>","PeriodicalId":46211,"journal":{"name":"CAAI Transactions on Intelligence Technology","volume":"11 4","pages":"994-1018"},"PeriodicalIF":5.1,"publicationDate":"2026-08-31","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://onlinelibrary.wiley.com/doi/epdf/10.1049/cit2.70150","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148860135","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Yongji Guan, Shunbao Zhang, Longbang Wang, Erchao Li
{"title":"Enhancing Generalisation via Cascaded Inertia SGD With Learnt Hyperparameters","authors":"Yongji Guan, Shunbao Zhang, Longbang Wang, Erchao Li","doi":"10.1049/cit2.70111","DOIUrl":"https://doi.org/10.1049/cit2.70111","url":null,"abstract":"<p>A central challenge in deep learning lies in achieving strong model generalisation, an area in which conventional optimisers such as stochastic gradient descent (SGD) often exhibit limitations, even though they ensure convergence. This paper introduces cascaded inertia SGD (CISGD), a novel optimisation algorithm specifically designed to address this challenge. The core mechanism of CISGD involves a hierarchical aggregation of historical gradients: it first accumulates past gradients and then compounds these accumulations to form multi-stage inertial terms. This deep integration of gradient history allows the optimiser produce more stable optimisation trajectories, thereby encouraging convergence towards flatter minima that are empirically associated with improved generalisation. To enhance practicality and robustness, we further eliminate manual hyperparameter tuning through a lightweight feed-forward neural network that automatically learns the optimal coefficients of the inertial terms, ensuring adaptability across diverse architectures and tasks. Extensive experiments on benchmark image classification and remote sensing change detection tasks demonstrate that CISGD consistently achieves superior generalisation performance compared with existing optimisers, demonstrating that CISGD achieves consistent, though moderate, improvements over standard optimisers while maintaining computational efficiency. This work offers a practical and theoretically grounded framework for automated and generalisable optimisation, promoting broader applications of adaptive deep learning in complex real-world tasks.</p>","PeriodicalId":46211,"journal":{"name":"CAAI Transactions on Intelligence Technology","volume":"11 4","pages":"1019-1027"},"PeriodicalIF":5.1,"publicationDate":"2026-08-31","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://onlinelibrary.wiley.com/doi/epdf/10.1049/cit2.70111","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148860180","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"APTNet: A Condition-Sensitive Modulation Framework for Surface Pressure Prediction on Supersonic Aircraft","authors":"Hongbin Xu, Yin Long, Junlin Wu","doi":"10.1049/cit2.70149","DOIUrl":"https://doi.org/10.1049/cit2.70149","url":null,"abstract":"<p>Accurate surface-pressure prediction over broad operating envelopes is critical for supersonic aerodynamic analysis and design. To overcome the bottlenecks of traditional computational fluid dynamics (CFD) in real-time performance and computational efficiency, data-driven deep learning methods have emerged. However, existing data-driven models face two fundamental physical challenges: first, the pressure response to Mach number, angle of attack and altitude is strongly spatially heterogeneous, which cannot be adequately captured by traditional global condition fusion methods; second, the inherent strong anisotropy of supersonic flows renders standard isotropic graph construction based on Euclidean distance prone to spurious cross-shock connections, resulting in nonphysical smoothing of predictions. To address these issues, we propose aircraft pressure transformer network (APTNet), a physics-informed deep learning framework that explicitly encodes flow–geometry interactions and streamwise physical priors. Specifically, we first propose a condition-sensitive local modulation (CSLM) module, which utilises a surface-normal–freestream alignment term to distinguish windward and leeward regions, and learns per-point sensitivity fields to implement physically interpretable local pressure modulation via a feature-wise linear modulation (FiLM) mechanism. Furthermore, we introduce a flow-aligned anisotropic graph construction strategy, which reshapes the local receptive fields of graph nodes into ellipsoids elongated along streamlines, enforcing information to propagate preferentially along streamlines and suppressing cross-shock feature mixing. Experiments on diverse complex aircraft configurations demonstrate that APTNet consistently outperforms existing baseline models in terms of both accuracy and efficiency. Beyond quantitative improvements, the model provides interpretable intermediate signals, such as attention maps and feature responses, which can be directly integrated into the aerodynamic design workflow.</p>","PeriodicalId":46211,"journal":{"name":"CAAI Transactions on Intelligence Technology","volume":"11 4","pages":"1302-1319"},"PeriodicalIF":5.1,"publicationDate":"2026-08-31","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://onlinelibrary.wiley.com/doi/epdf/10.1049/cit2.70149","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148860183","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"TriCrackNet: Trilateral Segmentation Network for Real-Time Crack Segmentation","authors":"Haixin Jia, Han Wang, Xiaoqiang Shi, Guoying Zhang, Hanqi Jiang, Jianbo Peng","doi":"10.1049/cit2.70129","DOIUrl":"https://doi.org/10.1049/cit2.70129","url":null,"abstract":"<p>To achieve high-precision real-time crack segmentation, we propose TriCrackNet, an efficient network based on a tri-branch collaborative architecture incorporating boundary constraints, semantic parsing, and spatial refinement. In the semantic branch, efficient atrous spatial pyramid pooling (EASPP) is integrated. The EASPP performs feature grouping and efficiently utilises each group's features through hierarchical residual connections. Simultaneously, it employs progressive dilated convolutions with incrementally expanding receptive fields to comprehensively capture multi-scale crack features. Meanwhile, an efficient feature enhancement module (EFEM) was inserted between the semantic branch and the other two branches. This module selectively utilises global features from the semantic branch to provide semantic guidance for the other branches, while enabling self-adaptive enhancement of local features in the target branches based on their own feature characteristics. Finally, the features of the last three branches are efficiently fused through an efficient feature fusion module that performs adaptive spatial weight weighting. Compared with other state-of-the-art methods, TriCrackNet achieves the best trade-off between inference speed and accuracy on CrackForest, Crack500, and DeepCrack datasets.</p>","PeriodicalId":46211,"journal":{"name":"CAAI Transactions on Intelligence Technology","volume":"11 4","pages":"1144-1160"},"PeriodicalIF":5.1,"publicationDate":"2026-08-31","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://onlinelibrary.wiley.com/doi/epdf/10.1049/cit2.70129","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148859956","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Training-Efficient Personalised Knee Joint Moment Estimation via an Incremental Broad Learning System","authors":"Guoyu Zuo, Minghui Zhang, Jingwei Ren, Qifei Wu, Jiaxin Li, Jianyu Chen, Chengyu Qiao, Shuangyue Yu","doi":"10.1049/cit2.70165","DOIUrl":"https://doi.org/10.1049/cit2.70165","url":null,"abstract":"<p>Knee joint moment estimation is a critical component in biomechanical analysis, with profound implications for rehabilitation assessment and the development of assistive devices like exoskeletons. However, existing data-driven approaches often rely on highly complex network architectures in pursuit of high accuracy, resulting in inherently inefficient training procedures. Furthermore, incorporating new subject-specific data typically requires retraining the entire model, which substantially increases computational cost. These limitations make it challenging for current data-driven methods to achieve personalised knee joint moment estimation. To address this challenge, this paper proposes a novel framework based on an Incremental Broad Learning System (IBLS) that can achieve efficient model training through a flattened network architecture and an analytical ridge regression solver. More importantly, it integrates an incremental learning algorithm that enables rapid model updates using only newly acquired subject-specific data, eliminating the need for costly full retraining. These capabilities collectively make efficient personalised knee joint moment estimation feasible. Extensive experiments on Dataset A and Dataset B demonstrate that our method achieves higher prediction accuracy compared to standard baselines (ANN and LSTM). Furthermore, our approach exhibits competitive performance when compared with advanced deep learning architectures (Transformer and TCN-LSTM). Notably, the training efficiency is significantly enhanced: the training time required by our method is approximately 30% of that for ANN, 7.5% for LSTM, and only 3.5% for both the Transformer and TCN-LSTM. Further experimental evidence indicates that incorporating incremental learning provides an additional 60% reduction in training time relative to full retraining, without sacrificing prediction accuracy. In addition, we systematically identify optimal Inertial Measurement Unit (IMU) configurations that balance accuracy with practical wearability, providing actionable guidelines for implementation. This work provides an efficient, accurate, and personalised solution for knee joint moment estimation, paving the way for the development of adaptive and efficient wearable robotic systems.</p>","PeriodicalId":46211,"journal":{"name":"CAAI Transactions on Intelligence Technology","volume":"11 4","pages":"1179-1193"},"PeriodicalIF":5.1,"publicationDate":"2026-08-31","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://onlinelibrary.wiley.com/doi/epdf/10.1049/cit2.70165","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148860030","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}