Pattern Recognition最新文献

筛选
英文 中文
Semantic change detection of roads and bridges: A fine-grained dataset and multimodal frequency-driven detector 道路和桥梁的语义变化检测:一个细粒度数据集和多模态频率驱动检测器
IF 7.6 1区 计算机科学
Pattern Recognition Pub Date : 2026-08-01 Epub Date: 2026-01-29 DOI: 10.1016/j.patcog.2026.113191
Qing-Ling Shu , Si-Bao Chen , Xiao Wang , Zhi-Hui You , Wei Lu , Jin Tang , Bin Luo
{"title":"Semantic change detection of roads and bridges: A fine-grained dataset and multimodal frequency-driven detector","authors":"Qing-Ling Shu ,&nbsp;Si-Bao Chen ,&nbsp;Xiao Wang ,&nbsp;Zhi-Hui You ,&nbsp;Wei Lu ,&nbsp;Jin Tang ,&nbsp;Bin Luo","doi":"10.1016/j.patcog.2026.113191","DOIUrl":"10.1016/j.patcog.2026.113191","url":null,"abstract":"<div><div>Accurate detection of road and bridge changes is crucial for urban planning and transportation management, yet presents unique challenges for general change detection (CD). Key difficulties arise from maintaining the continuity of roads and bridges as linear structures and disambiguating visually similar land covers (e.g., road construction vs. bare land). Existing spatial-domain models struggle with these issues, further hindered by the lack of specialized, semantically rich datasets. To fill these gaps, we introduce the Road and Bridge Semantic Change Detection (RB-SCD) dataset. Unlike existing benchmarks that primarily focus on general land cover changes, RB-SCD is the first to systematically target 11 specific semantic change transition types (e.g., water → bridge) anchored to traffic infrastructure. This enables a detailed analysis of traffic infrastructure evolution. Building on this, we propose a novel framework, the Multimodal Frequency-Driven Change Detector (MFDCD). MFDCD integrates multimodal features in the frequency domain through two key components: (1) the Dynamic Frequency Coupler (DFC), which leverages wavelet transform to decompose visual features, enabling it to robustly model the continuity of linear transitions; and (2) the Textual Frequency Filter (TFF), which encodes semantic priors into frequency-domain graphs and applies filter banks to align them with visual features, resolving semantic ambiguities. Experiments demonstrate the state-of-the-art performance of MFDCD on RB-SCD and three public CD datasets. The code will be available at <span><span>https://github.com/DaGuangDaGuang/RB-SCD</span><svg><path></path></svg></span>.</div></div>","PeriodicalId":49713,"journal":{"name":"Pattern Recognition","volume":"176 ","pages":"Article 113191"},"PeriodicalIF":7.6,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"146174257","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":1,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
DAFS: A distribution-aware hierarchical feature selection method for long-tailed classification DAFS:一种分布感知的长尾分类分层特征选择方法
IF 7.6 1区 计算机科学
Pattern Recognition Pub Date : 2026-08-01 Epub Date: 2026-02-06 DOI: 10.1016/j.patcog.2026.113218
Yang Zhang , Jie Shi , Yanfang Liu , Hong Zhao
{"title":"DAFS: A distribution-aware hierarchical feature selection method for long-tailed classification","authors":"Yang Zhang ,&nbsp;Jie Shi ,&nbsp;Yanfang Liu ,&nbsp;Hong Zhao","doi":"10.1016/j.patcog.2026.113218","DOIUrl":"10.1016/j.patcog.2026.113218","url":null,"abstract":"<div><div>Feature selection for long-tailed data has become a research hotspot due to high-dimensional features and imbalanced distributions in real-world data. Although some of them effectively balance the data, correctly classifying tail classes and distinguishing easy-confused classes in long-tailed data are still two significant challenges. To address these issues, we propose a distribution-aware hierarchical feature selection method for long-tailed classification (DAFS). First, we embed sample distribution-based punishment coefficients into loss and regularization terms to balance feature weights for head and tail classes, which enhances the accuracy of classifying tail classes. Then, we use multi-granularity knowledge and similarities among classes to design feature differentiation regularization terms for improving the distinguishability of easy-confused classes. Finally, extensive experimental results demonstrate that DAFS outperforms the other ten traditional and advanced feature selection methods on different datasets.</div></div>","PeriodicalId":49713,"journal":{"name":"Pattern Recognition","volume":"176 ","pages":"Article 113218"},"PeriodicalIF":7.6,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"146174265","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":1,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
CD-DPC: Centrifugal degree based density peaks clustering algorithm CD-DPC:基于离心度的密度峰聚类算法
IF 7.6 1区 计算机科学
Pattern Recognition Pub Date : 2026-08-01 Epub Date: 2026-02-06 DOI: 10.1016/j.patcog.2026.113223
Linlin Ma , Hui Li , Xincheng Liu , Huihui Chu , Yue Guan , Yuzhen Zhao , Yawen Chen , Da Wang , Wenke Zang
{"title":"CD-DPC: Centrifugal degree based density peaks clustering algorithm","authors":"Linlin Ma ,&nbsp;Hui Li ,&nbsp;Xincheng Liu ,&nbsp;Huihui Chu ,&nbsp;Yue Guan ,&nbsp;Yuzhen Zhao ,&nbsp;Yawen Chen ,&nbsp;Da Wang ,&nbsp;Wenke Zang","doi":"10.1016/j.patcog.2026.113223","DOIUrl":"10.1016/j.patcog.2026.113223","url":null,"abstract":"<div><div>The Density Peak Clustering (DPC) algorithm is simple and efficient. But DPC and its variants identify clusters only by identifying the centers of single or multiple sparse clusters without considering the coherence of the clustering structure, which tends to result in clusters that cannot be accurately captured. In addition, relative distance and density are only used to identify the centers of clusters and do not provide a description of the relative positions of the remaining sample points. To address these issues, this paper proposes an adaptive density peak clustering algorithm based on centrifugal degree (CD-DPC). The centrifugal degree reflects the relative position of the sample points in the cluster. The CD-DPC categorizes sample points into support, structural, coherent and decoration points based on centrifugal degree. Based on this, the number of clusters is automatically obtained by using different association methods for sample points with different centrifugal degrees, which greatly reduces the influence of human factors. Finally, the clustering results are further improved by introducing shared nearest neighbors for the final association of decorated points. Extensive experiments on synthetic and UCI datasets show that this algorithm outperforms other comparative algorithms.</div></div>","PeriodicalId":49713,"journal":{"name":"Pattern Recognition","volume":"176 ","pages":"Article 113223"},"PeriodicalIF":7.6,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"146174446","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":1,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
MonoTDF: Temporal deep feature learning for generalizable monocular 3D object detection 用于泛化单目3D物体检测的时间深度特征学习
IF 7.6 1区 计算机科学
Pattern Recognition Pub Date : 2026-08-01 Epub Date: 2026-01-28 DOI: 10.1016/j.patcog.2026.113184
Xiu-Zhi Chen , Yi-Kai Chiu , Chih-Sheng Huang , Yen-Lin Chen
{"title":"MonoTDF: Temporal deep feature learning for generalizable monocular 3D object detection","authors":"Xiu-Zhi Chen ,&nbsp;Yi-Kai Chiu ,&nbsp;Chih-Sheng Huang ,&nbsp;Yen-Lin Chen","doi":"10.1016/j.patcog.2026.113184","DOIUrl":"10.1016/j.patcog.2026.113184","url":null,"abstract":"<div><div>Monocular 3D object detection has gained significant attention due to its cost-effectiveness and practicality in real-world applications. However, existing monocular methods often struggle with depth estimation and spatial consistency, limiting their accuracy in complex environments. In this work, we introduce a Temporal Deep Feature Learning framework, which enhances monocular 3D object detection by integrating temporal features across sequential frames. Our approach leverages a novel deep feature auxiliary module based on convolutional recurrent structures, effectively capturing spatiotemporal information to improve depth perception and detection robustness. The proposed module is model-agnostic and can be seamlessly integrated into various existing monocular detection frameworks. Extensive experiments across multiple state-of-the-art monocular 3D object detection models demonstrate consistent performance improvements, particularly in detecting small or partially occluded objects. Our results highlight the effectiveness and generalizability of the proposed approach, making it a promising solution for real-world autonomous perception systems. The source code of this work is at: <span><span>https://github.com/Shuray36/MonoTDF-Temporal-Deep-Feature-Learning-for-Generalizable-Monocular-3D-Object-Detection</span><svg><path></path></svg></span>.</div></div>","PeriodicalId":49713,"journal":{"name":"Pattern Recognition","volume":"176 ","pages":"Article 113184"},"PeriodicalIF":7.6,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"146174372","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":1,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
Multimodal behavioral analysis for autism spectrum disorder assessment 自闭症谱系障碍评估的多模态行为分析
IF 7.6 1区 计算机科学
Pattern Recognition Pub Date : 2026-08-01 Epub Date: 2026-02-07 DOI: 10.1016/j.patcog.2026.113155
Yunxiu Zhao , Shigang Wang , Feiyong Jia , Honghua Li , Jinyang Wu , Jian Wei , Yan Zhao
{"title":"Multimodal behavioral analysis for autism spectrum disorder assessment","authors":"Yunxiu Zhao ,&nbsp;Shigang Wang ,&nbsp;Feiyong Jia ,&nbsp;Honghua Li ,&nbsp;Jinyang Wu ,&nbsp;Jian Wei ,&nbsp;Yan Zhao","doi":"10.1016/j.patcog.2026.113155","DOIUrl":"10.1016/j.patcog.2026.113155","url":null,"abstract":"<div><div>Scale-dependent approaches have shown great potential in diagnosing autism spectrum disorder (ASD). However, such methods often involve lengthy evaluation procedures and require substantial resources, including trained professionals and specialized equipment, which significantly limit their scalability and feasibility for large-scale or routine clinical assessments. In this paper, we propose a novel multimodal behavioral signal analysis (MBSA) approach for the intelligent assessment of ASD. Specifically, we first leverage speech and visual cues to identify the Target Movement Area (TMA), thereby enhancing recognition efficiency. Then, an adaptive fine-tuning strategy is employed to improve the generalization and efficiency of pre-trained models in small-sample action recognition tasks. An attention-based detection method is further incorporated to strengthen the semantic understanding of observed behavioral patterns. To enable effective ASD classification, we develop a behavioral quantification scoring method that structurally models the relationship between behavioral features and disease indicators. We collected a multimodal behavioral database of 160 participants in a real clinical setting and assessed ASD using this data. Extensive experiments demonstrate that the proposed MBSA approach significantly outperforms many state-of-the-art methods. With competitive performance and a solid theoretical foundation, MBSA provides a practical and scalable solution for ASD screening and holds promise for broader applications in the intelligent diagnosis of other neurodevelopmental disorders.</div></div>","PeriodicalId":49713,"journal":{"name":"Pattern Recognition","volume":"176 ","pages":"Article 113155"},"PeriodicalIF":7.6,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"146174545","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":1,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
Continual relation extraction with wake-sleep memory consolidation 连续关系提取与清醒-睡眠记忆巩固
IF 7.6 1区 计算机科学
Pattern Recognition Pub Date : 2026-08-01 Epub Date: 2026-01-29 DOI: 10.1016/j.patcog.2026.113192
Tingting Hang , Ya Guo , Jun Huang , Yirui Wu , Umapada Pal , Shivakumara Palaiahnakote
{"title":"Continual relation extraction with wake-sleep memory consolidation","authors":"Tingting Hang ,&nbsp;Ya Guo ,&nbsp;Jun Huang ,&nbsp;Yirui Wu ,&nbsp;Umapada Pal ,&nbsp;Shivakumara Palaiahnakote","doi":"10.1016/j.patcog.2026.113192","DOIUrl":"10.1016/j.patcog.2026.113192","url":null,"abstract":"<div><div>Continual Relation Extraction (CRE) has achieved significant success due to its ability to adapt to new relations without frequent retraining. However, existing methods still face challenges such as overfitting and representation bias. Inspired by the wake-sleep memory consolidation process of the human brain, this paper proposes a <strong>W</strong>ake-<strong>S</strong>leep <strong>M</strong>emory <strong>C</strong>onsolidation (WSMC) framework to address these issues systematically. During the wake phase, the model simulates the brain’s information processing mechanism, quickly encoding new relations and storing them in short-term memory. We also introduce the Experience Iterative Learning (EIL) approach, which dynamically adjusts the distribution of relation samples. This approach corrects the model’s representation bias and enhances memory stability through experience replay. During the sleep phase, the model consolidates existing knowledge by replaying long-term memory. Moreover, the framework generates diverse dream data from existing memory sets, thereby increasing the diversity of the training data and improving the model’s generalization capability. Experimental results show that WSMC significantly outperforms other CRE baseline methods on FewRel and TACRED datasets, demonstrating its superior performance compared to baseline methods. Our source code is available at <span><span>https://github.com/Gyanis9/WSMC.git</span><svg><path></path></svg></span>.</div></div>","PeriodicalId":49713,"journal":{"name":"Pattern Recognition","volume":"176 ","pages":"Article 113192"},"PeriodicalIF":7.6,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"146174535","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":1,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
FPMT: Fast and precise high-resolution makeup transfer via Laplacian pyramid FPMT:通过拉普拉斯金字塔快速精确的高分辨率化妆转移
IF 7.6 1区 计算机科学
Pattern Recognition Pub Date : 2026-08-01 Epub Date: 2026-02-02 DOI: 10.1016/j.patcog.2026.113221
Zhaoyang Sun , Shengwu Xiong , Yi Rong
{"title":"FPMT: Fast and precise high-resolution makeup transfer via Laplacian pyramid","authors":"Zhaoyang Sun ,&nbsp;Shengwu Xiong ,&nbsp;Yi Rong","doi":"10.1016/j.patcog.2026.113221","DOIUrl":"10.1016/j.patcog.2026.113221","url":null,"abstract":"<div><div>In this paper, we focus on accelerating high-resolution makeup transfer process without compromising generative performance. To this end, we propose a Fast and Precise Makeup Transfer (FPMT) framework based on Laplacian pyramid. In FPMT, we reveal that most makeup changes are concentrated in the low-frequency component, while a small amount of color- and texture-related details are included in the high-frequency components. Leveraging this insight, FPMT employs a lightweight encoder-decoder network to perform makeup transfer on the low-frequency component of inputs, thus improving efficiency. For each high-frequency component, FPMT implements a tiny refinement network that progressively predicts a mask and adaptively refines the makeup details to ensure transfer quality. By stacking the computationally efficient refinement network, FPMT can process higher-resolution images, demonstrating its flexibility and scalability. Using a single GTX 1660Ti GPU, FPMT can achieve an inference speed of about 42 FPS for input images with 1024 × 1024 resolution, which is much faster than the state-of-the-art methods. Extensive quantitative and qualitative analyses validate the efficiency and effectiveness of the proposed FPMT framework. The source code is available at: <span><span>https://github.com/Snowfallingplum/FPMT</span><svg><path></path></svg></span>.</div></div>","PeriodicalId":49713,"journal":{"name":"Pattern Recognition","volume":"176 ","pages":"Article 113221"},"PeriodicalIF":7.6,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"146174538","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":1,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
MC-MVSNet: When multi-view stereo meets monocular cues 当多视角立体遇到单目线索时
IF 7.6 1区 计算机科学
Pattern Recognition Pub Date : 2026-08-01 Epub Date: 2026-01-27 DOI: 10.1016/j.patcog.2026.113166
Xincheng Tang , Mengqi Rong , Bin Fan , Hongmin Liu , Shuhan Shen
{"title":"MC-MVSNet: When multi-view stereo meets monocular cues","authors":"Xincheng Tang ,&nbsp;Mengqi Rong ,&nbsp;Bin Fan ,&nbsp;Hongmin Liu ,&nbsp;Shuhan Shen","doi":"10.1016/j.patcog.2026.113166","DOIUrl":"10.1016/j.patcog.2026.113166","url":null,"abstract":"<div><div>Learning-based Multi-View Stereo (MVS) has become a key technique for reconstructing dense 3D point clouds from multiple calibrated images. However, real-world challenges such as occlusions and textureless regions often hinder accurate depth estimation. Recent advances in monocular Vision Foundation Models (VFMs) have demonstrated strong generalization capabilities in scene understanding, offering new opportunities to enhance the robustness of MVS. In this paper, we present MC-MVSNet, a novel MVS framework that integrates diverse monocular cues to improve depth estimation under challenging conditions. During feature extraction, we fuse conventional CNN features with VFM-derived representations through a hybrid feature fusion module, effectively combining local details and global context for more discriminative feature matching. We also propose a cost volume filtering module that enforces cross-view geometric consistency on monocular depth predictions, pruning redundant depth hypotheses to reduce the depth search space and mitigate matching ambiguity. Additionally, we leverage monocular surface normals to construct a curved patch cost aggregation module that aggregates costs over geometry-aligned curved patches, which improves depth estimation accuracy in curved and textureless regions. Extensive experiments on the DTU, Tanks and Temples, and ETH3D benchmarks demonstrate that MC-MVSNet achieves state-of-the-art performance and exhibits strong generalization capabilities, validating the effectiveness and robustness of the proposed method.</div></div>","PeriodicalId":49713,"journal":{"name":"Pattern Recognition","volume":"176 ","pages":"Article 113166"},"PeriodicalIF":7.6,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"146081411","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":1,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
Attribute graph adjusted trace ratio linear discriminant analysis for feature extraction 属性图调整迹比线性判别分析特征提取
IF 7.6 1区 计算机科学
Pattern Recognition Pub Date : 2026-08-01 Epub Date: 2026-01-22 DOI: 10.1016/j.patcog.2026.113136
Quan Wang , Hao Lei , Fei Wang , Xinpei Wen , Zhiping Lin , Feiping Nie
{"title":"Attribute graph adjusted trace ratio linear discriminant analysis for feature extraction","authors":"Quan Wang ,&nbsp;Hao Lei ,&nbsp;Fei Wang ,&nbsp;Xinpei Wen ,&nbsp;Zhiping Lin ,&nbsp;Feiping Nie","doi":"10.1016/j.patcog.2026.113136","DOIUrl":"10.1016/j.patcog.2026.113136","url":null,"abstract":"<div><div>Trace Ratio Linear Discriminant Analysis (TRLDA) is an appealing supervised feature extraction method because it explicitly reflects the Euclidean distances between and within classes of projected samples while preserving data similarity through its orthogonal constraint. However, TRLDA fails to account for inter-attribute correlations, which may limit its discriminant capability. To overcome this limitation, we propose Attribute Graph Adjusted Trace Ratio Linear Discriminant Analysis (AGATRLDA), a novel method that incorporates attribute-level relationships into the discriminant projection matrix. In our approach, each attribute is represented as a point formed by the values of that attribute across all samples. An attribute graph is then constructed by connecting these attribute points with edges weighted according to their pairwise similarity. By integrating the Laplacian matrix of this attribute graph into the optimization framework, AGATRLDA adjusts the discriminant projection matrix to account for inter-attribute correlations. This adjustment encourages attributes with higher similarity to have more aligned coefficients in the projection matrix, thereby improving discriminative performance. Experimental results demonstrate that AGATRLDA consistently outperforms the original TRLDA method as well as several state-of-the-art feature extraction techniques, validating the benefit of incorporating inter-attribute correlations in the discriminant learning process.</div></div>","PeriodicalId":49713,"journal":{"name":"Pattern Recognition","volume":"176 ","pages":"Article 113136"},"PeriodicalIF":7.6,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"146081398","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":1,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
A two-stage learning framework with a beam image dataset for automatic laser resonator alignment 基于光束图像数据集的激光谐振器自动对准两阶段学习框架
IF 7.6 1区 计算机科学
Pattern Recognition Pub Date : 2026-08-01 Epub Date: 2026-01-27 DOI: 10.1016/j.patcog.2026.113145
Shaoxiang Guo , Donald Risbridger , David A. Robb , Xianwen Kong , M. J. Daniel Esser , Michael J. Chantler , Richard M. Carter , Mustafa Suphi Erden
{"title":"A two-stage learning framework with a beam image dataset for automatic laser resonator alignment","authors":"Shaoxiang Guo ,&nbsp;Donald Risbridger ,&nbsp;David A. Robb ,&nbsp;Xianwen Kong ,&nbsp;M. J. Daniel Esser ,&nbsp;Michael J. Chantler ,&nbsp;Richard M. Carter ,&nbsp;Mustafa Suphi Erden","doi":"10.1016/j.patcog.2026.113145","DOIUrl":"10.1016/j.patcog.2026.113145","url":null,"abstract":"<div><div>Accurate alignment of a laser resonator is essential for upscaling industrial laser manufacturing and precision processing. However, traditional manual or semi-automatic methods depend heavily on operator expertise, and struggle with the interdependence among multiple alignment parameters. To tackle this, we introduce the first real-world image dataset for automatic laser resonator alignment, collected on a laboratory-built resonator setup. It comprises over 6000 beam profiler images annotated with four key alignment parameters (intracavity iris aperture diameter, output coupler pitch and yaw actuator displacements, and axial position of the output coupler), with over 500,000 paired samples for data-driven alignment. Given a pair of beam profiler images exhibiting distinct beam patterns under different configurations, the system predicts the control-parameter changes required to realign the resonator. Leveraging this dataset, we propose a novel two-stage deep learning framework for automatic resonator alignment. In Stage 1, a multi-scale CNN augmented with cross-attention and correlation-difference modules, extracts features and outputs an initial coarse prediction of alignment parameters. In Stage 2, a feature-difference map is computed by subtracting the paired feature representations and fed into an iterative refinement module to correct residual misalignments. The final prediction combines coarse and refined estimates, integrating global context with fine-grained corrections for accurate inference. Experiments on our dataset and a different instance of the same physical system from which the CNN was trained suggest superior accuracy and practicality to manual alignment.</div></div>","PeriodicalId":49713,"journal":{"name":"Pattern Recognition","volume":"176 ","pages":"Article 113145"},"PeriodicalIF":7.6,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"146081394","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":1,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
0
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
确定
请完成安全验证×
相关产品
×
本文献相关产品
联系我们:info@booksci.cn Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。 Copyright © 2023 布克学术 All rights reserved.
京ICP备2023020795号-1
ghs 京公网安备 11010802042870号
Book学术文献互助
Book学术文献互助群
群 号:604180095
Book学术官方微信
小红书