Maria Dhakal, Chia-Yi Su, Robert Wallace, Chris Fakhimi, Aakash Bansal, Toby Li, Yu Huang, Collin McMillan
{"title":"A Grounded Theory Study to Guide AI-Driven Code Comment Improvement","authors":"Maria Dhakal, Chia-Yi Su, Robert Wallace, Chris Fakhimi, Aakash Bansal, Toby Li, Yu Huang, Collin McMillan","doi":"10.1002/smr.70157","DOIUrl":"https://doi.org/10.1002/smr.70157","url":null,"abstract":"<p>Current automated code summarization research focuses primarily on mimicking existing human-written comments, ignoring the fact that many baseline comments are incomplete or misaligned with developer needs. This paper presents a novel approach to intentionally improve code comments along different quality axes by rewriting those comments with customized artificial intelligence (AI)-based tools. We first conduct an empirical study with developers, followed by grounded theory to develop a theoretical construct we define as comment refocusing. This theory identifies seven distinct quality axes grouped into two dimensions: Internalization (anchoring comments in implementation logic) and Externalization (projecting comments into system context). Based on this theoretical foundation, we propose a two-stage alignment procedure using reinforcement learning from AI feedback (RLAIF) followed by human feedback (RLHF) to steer large language models toward these quality axes. We implement this procedure using a large-scale teacher model (GPT-4o) and distill the knowledge into smaller, privacy-preserving models (e.g., jam, CodeLlama-13B) capable of in-house deployment. Our evaluation demonstrates that our procedure improves code comments along the quality axes, satisfying specific developer-defined quality goals. We release all data and source code in an online repository for reproducibility.</p>","PeriodicalId":48898,"journal":{"name":"Journal of Software-Evolution and Process","volume":"38 7","pages":""},"PeriodicalIF":1.6,"publicationDate":"2026-07-26","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://onlinelibrary.wiley.com/doi/epdf/10.1002/smr.70157","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148615820","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Elizangela de Freitas Ximenes, José Siqueira Cerqueira, Pekka Abrahamsson, Edna Dias Canedo
{"title":"ERC4AI: A BERT-Based Model for Ethical Requirements Classification in AI Systems","authors":"Elizangela de Freitas Ximenes, José Siqueira Cerqueira, Pekka Abrahamsson, Edna Dias Canedo","doi":"10.1002/smr.70156","DOIUrl":"https://doi.org/10.1002/smr.70156","url":null,"abstract":"<div>\u0000 \u0000 <p>The operationalization of AI ethical principles during the requirements engineering phase is paramount for creating ethically aligned AI systems. However, ethical requirements classification remains a significant challenge due to the abstract nature of ethical principles and the lack of practical tools to support developers. We introduce the EthicalRequirements4AI, a dataset comprising 1091 annotated requirements, including ethical and non-ethical requirements. We also provide the Ethical Requirements Classification for AI (ERC4AI): a BERT-based multi-label classification model for classifying ethical requirements in AI, fine-tuned on our labeled dataset. Passenger Flow and PROMISE datasets were used to construct the EthicalRequirements4AI dataset, encompassing 385 requirements aligned with 11 AI ethical principles and 706 non-ethical requirements. Three transformer-based models—XLM-RoBERTa, BERT, and DistilBERT—were evaluated for multi-label classification task. BERT demonstrated the most stable overall performance across seeds while achieving the highest macro-average performance among the evaluated models (XLM-RoBERTa, BERT, and DistilBERT), achieving a macro F1 score of 0.76 and a weighted F1 score of 0.83 on the test set. Although BERT showed the best aggregate results, performance varied across different ethical principles. ERC4AI offers a practical means for researchers and developers to address ethical concerns early in AI development. The dataset and model are publicly available to foster further research and advancements in practical AI ethics.</p>\u0000 </div>","PeriodicalId":48898,"journal":{"name":"Journal of Software-Evolution and Process","volume":"38 7","pages":""},"PeriodicalIF":1.6,"publicationDate":"2026-07-21","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148615591","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"On the Diffuseness and the Impact of Multi-Language Design Smells","authors":"Md Shahrukh Ansari, Salman Abdul Moiz","doi":"10.1002/smr.70153","DOIUrl":"https://doi.org/10.1002/smr.70153","url":null,"abstract":"<div>\u0000 \u0000 <p>Multi-language software systems integrate components written in different programming languages, offering flexibility but also introduces design challenges at language boundaries. These challenges often manifest as multi-language design smells-suboptimal design practices that can affect maintainability and reliability. Previous research proposed a detector for these smells and analyzed their quality impact; however, its accuracy and generalizability have not been independently validated. This study evaluates and refines the existing multi-language design smell detector, reassesses smell prevalence using the improved tool, and reanalyzes their relationship with software quality attributes such as fault- and change-proneness. The detector was first evaluated on 60 open-source projects to test its generalizability, after which rule-level inconsistencies were identified and corrected. The refined detector was then applied to multiple releases of nine Java Native Interface (JNI) systems for updated prevalence and impact analysis. The refined detector achieves perfect alignment with formal definitions and manual annotations. The reanalysis reveals revised distributions of smells and a much weaker statistical relationship with software faults than previously reported, whereas new evidence shows only limited associations with change-proneness.</p>\u0000 </div>","PeriodicalId":48898,"journal":{"name":"Journal of Software-Evolution and Process","volume":"38 7","pages":""},"PeriodicalIF":1.6,"publicationDate":"2026-07-21","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148615588","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Predictive Modeling of Software Defect Density Using BAHA-WOA–Based Feature Selection","authors":"Jasmeet Kaur, Arvinder Kaur, Kamaldeep Kaur","doi":"10.1002/smr.70154","DOIUrl":"https://doi.org/10.1002/smr.70154","url":null,"abstract":"<div>\u0000 \u0000 <p>Accurate prediction of software defect density is crucial for identifying defect-prone modules early, improving software quality, optimizing testing efforts, and enabling efficient resource allocation; however, software defect datasets often contain redundant or irrelevant features, and complex nonlinear relationships make predictive modeling challenging. This study aims to develop an efficient and robust feature selection approach to enhance the accuracy and reliability of software defect density prediction. To address these challenges, a novel binary hybrid feature selection algorithm, BAHA-WOA, is proposed by integrating the artificial hummingbird algorithm (AHA) for global exploration with the whale optimization algorithm (WOA) for local exploitation, incorporating a crossover mechanism to enhance population diversity and employing linear regression (LR) as a wrapper-based fitness evaluator; continuous solutions are transformed into binary form using an S-shaped transfer function, and the approach is evaluated on eight JIRA-based Apache project datasets against 10 state-of-the-art binary metaheuristic algorithms. The findings demonstrate that BAHA-WOA (1) identifies the most relevant features, reducing dimensionality and improving model generalization; (2) achieves superior predictive accuracy with lower RMSE and MAE; (3) maintains computational efficiency through LR-guided evaluation; and (4) produces stable and statistically significant results validated using the Friedman test. Overall, BAHA-WOA provides a robust, efficient, and effective solution for feature selection in software defect density prediction, enhancing software quality assessment and reliability.</p>\u0000 </div>","PeriodicalId":48898,"journal":{"name":"Journal of Software-Evolution and Process","volume":"38 7","pages":""},"PeriodicalIF":1.6,"publicationDate":"2026-07-20","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148615584","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Enhancing Feature Quality in JIT-SDP: An Empirical Study of Branch Dependency Mining","authors":"Zuowei Chen, Liyan Song","doi":"10.1002/smr.70155","DOIUrl":"https://doi.org/10.1002/smr.70155","url":null,"abstract":"<div>\u0000 \u0000 <p>Just-in-time software defect prediction (JIT-SDP) aims to predict defect-inducing software changes in a timely manner, thereby enhancing development efficiency and product quality, especially during the software maintenance phase. Previous studies have primarily focused on proposing novel methods to achieve stronger predictive performance, often evaluated using open-source projects produced through data extraction tools. This extraction process typically operates under the implicit assumption that the extracted data features are exempt from noise. However, limited attention has been given to data quality, particularly regarding feature quality. We have identified that common feature extraction methods, such as those implemented by verb—Commit Guru—usually overlook branch dependency information, which can compromise feature quality and lead to unreliable or invalid conclusions. This paper systematically evaluates the quality of features extracted in JIT-SDP and investigates their impact on predictive performance. Moreover, we propose a novel feature extraction method that accounts for branch dependency to produce high-quality features. Our experiments based on 22 open projects reveal that neglecting branch dependency can significantly affect feature quality; however, this issue has a limited impact on the predictive performance of JIT-SDP models. Our results reveal that, although neglecting branch dependency introduces measurable feature discrepancies, the downstream predictive performance and model rankings of JIT-SDP models remain largely stable. These findings provide strong evidence regarding the empirical robustness of prior JIT-SDP studies, suggesting that the field's foundational conclusions are not compromised by this specific data quality issue. Our enhanced feature extraction method is publicly available at https://github.com/HongJinSecond/Feature-study.</p>\u0000 </div>","PeriodicalId":48898,"journal":{"name":"Journal of Software-Evolution and Process","volume":"38 7","pages":""},"PeriodicalIF":1.6,"publicationDate":"2026-07-20","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148615583","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"A Lightweight Modeling and Verification Framework for Programs Over Weak Memory Models in an Algebraic Semantics Style","authors":"Lili Xiao, Huibiao Zhu, Jonathan P. Bowen","doi":"10.1002/smr.70149","DOIUrl":"https://doi.org/10.1002/smr.70149","url":null,"abstract":"<div>\u0000 \u0000 <p>Modeling and verification of multithreaded programs are difficult since one must consider all the ways that instructions in different threads can be interleaved. Modern hardware architectures and mainstream programming languages employ weak memory models (WMMs) for efficiency reasons, and the additional interleavings from them make the modeling and verification more complex. In addition, different WMMs cause various relaxed-memory effects, and their operational semantics has distinguished expressions and transition rules. On this basis, multiple algorithms are required to be designed to conduct verification on programs over WMMs. In this paper, we propose a lightweight modeling and verification framework for programs over WMMs. Above all, we apply <i>Unifying Theories of Programming</i> (UTP) to investigate the unified algebraic semantics. A set of algebraic laws are explored, which can dynamically generate configuration sequences of programs under WMMs. Two memory models, total store order (TSO) implemented in the <span></span><math>\u0000 <semantics>\u0000 <mrow>\u0000 <mo>×</mo>\u0000 </mrow>\u0000 <annotation>$$ times $$</annotation>\u0000 </semantics></math>86 architecture and SPARC implementations, and ARMv8 supported by the ARM architecture, are used to instantiate the proposed algebraic modeling method. During this process, we record the data state of each configuration and define properties capturing the unique features of the TSO and ARMv8 memory models, and then check whether the properties are satisfied. The algebraic laws are implemented in the rewriting engine Maude, and the verification is also conducted in Maude. The verification results show that the properties all agree with our expectations.</p>\u0000 </div>","PeriodicalId":48898,"journal":{"name":"Journal of Software-Evolution and Process","volume":"38 7","pages":""},"PeriodicalIF":1.6,"publicationDate":"2026-07-12","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148466909","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Enhancing NSGA-II for Refactoring Recommendations: A Constraint-Based Initialization and Adaptive Crossover Approach","authors":"Yang Zhang, Meiyan Zheng","doi":"10.1002/smr.70151","DOIUrl":"https://doi.org/10.1002/smr.70151","url":null,"abstract":"<div>\u0000 \u0000 <p>Search-based refactoring recommendation methods, particularly those utilizing the non-dominated sorting genetic algorithm II (NSGA-II), have shown significant promise in automating software refactoring. However, existing approaches often rely on random initialization of the population and simplistic crossover operators, which can lead to suboptimal solutions and inefficient exploration of the search space. This paper proposes ReReC, a novel refactoring recommendation approach that enhances NSGA-II through constraint-based initialization and an adaptive crossover operator. ReReC introduces heuristic constraint rules for common refactoring types to generate a high-quality initial population, thereby reducing invalid refactoring operations. Furthermore, it employs an adaptive crossover strategy that distinguishes between common and differential refactoring operations, incorporating an elite gene guidance mechanism to accelerate convergence toward optimal solutions. To evaluate the effectiveness of ReReC, we evaluated it on six open-source projects. ReReC achieves an average of 81.95% F1 score, outperforming existing tools JMove and QMove by 7.48% and 5.50%, respectively. The results demonstrate that ReReC effectively improves both the accuracy and efficiency of refactoring recommendations.</p>\u0000 </div>","PeriodicalId":48898,"journal":{"name":"Journal of Software-Evolution and Process","volume":"38 7","pages":""},"PeriodicalIF":1.6,"publicationDate":"2026-07-06","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148462733","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Adaptive Multi-Metric Test Case Selection for Deep Neural Networks Based on Genetic Algorithm","authors":"Weiwei Wang, Qingshuai Chen, Zimo Zhao, Xuejun Liu, Ruilian Zhao","doi":"10.1002/smr.70150","DOIUrl":"https://doi.org/10.1002/smr.70150","url":null,"abstract":"<div>\u0000 \u0000 <p>As deep neural networks (DNNs) are increasingly deployed in safety- and mission-critical domains, their latent defects and security risks have become a growing concern. Test case selection (TCS) aims to identify and label those test cases most likely to be misclassified within a limited annotation budget, thereby maximizing the exposure of real faults at minimal cost and providing high-value data for subsequent localization, repair, and regression testing. However, current TCS methods predominantly utilize single-dimensional metrics, such as neuron coverage or mutant-killing rate, limiting their ability to comprehensively capture diverse fault patterns in complex models. This paper proposes a novel DNN TCS framework based on adaptive multi-metric optimization, which takes both prediction uncertainty and divergence into account and adjust the weights of them to qualify the effectiveness of test cases by genetic algorithm. Concretely, firstly, multiple slightly mutated models are constructed to capture the behavioral discrepancies of each test case across different models. It then establishes two complementary effectiveness metrics-prediction uncertainty and prediction divergence-to quantify a test case's local decision boundary sensitivity and global behavioral diversity, respectively. Finally, a genetic algorithm adaptively optimizes the weights of these metrics, enabling a precise assessment of each test case and the identification of the subset with the highest fault-revealing potential. To verify the effectiveness of our approach, experiments and evaluations are conducted on DNNs across both computer vision and natural language processing domains. And the experimental results show that our approach surpasses the best baseline by improving the fault detection rate by an average of 3.32% on the Top-5% candidate subset and 9.54% on the Top-10% subset.</p>\u0000 </div>","PeriodicalId":48898,"journal":{"name":"Journal of Software-Evolution and Process","volume":"38 7","pages":""},"PeriodicalIF":1.6,"publicationDate":"2026-07-05","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148462793","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"PRTS: Test Sample Selection Based on Category Probability Repair for DNN Testing","authors":"Feifan Gao, Zhiyi Zhang, Tianyu Xu, Shuxian Chen, Zhiqiu Huang","doi":"10.1002/smr.70152","DOIUrl":"https://doi.org/10.1002/smr.70152","url":null,"abstract":"<div>\u0000 \u0000 <p>Selecting appropriate test samples is a critical and frequently utilized step in optimizing deep neural network (DNN) testing. However, DNN models often exhibit high confidence in incorrect predictions, limiting the effectiveness of traditional selection methods that rely on model predictions to detect overconfident faults. To address this issue, this work introduces probability repair–based test sample selection (PRTS), a novel approach that repairs DNN probability distributions to improve test sample selection. We first categorize test samples into trustworthy and untrustworthy subsets based on uncertainty scores. Unlike conventional methods that rely solely on model-predicted class probabilities, we design feature similarity–based class probabilities (FSCP) to quantify uncertainty scores by measuring the discrepancy between these two probability distributions. Additionally, a repair process is incorporated to mitigate overconfidence and underconfidence in FSCP, explicitly targeting untrustworthy data instances to reduce the influence of erroneous neighbors and enhance overall efficiency. Extensive experiments were conducted across diverse datasets and DNN architectures, comparing PRTS with 12 state-of-the-art baseline methods. The results demonstrate that PRTS generally outperforms existing approaches in terms of the number and diversity of revealed faults and improvements in DNN model accuracy, thereby serving as a robust and efficient solution to enhance the reliability of DNN testing.</p>\u0000 </div>","PeriodicalId":48898,"journal":{"name":"Journal of Software-Evolution and Process","volume":"38 7","pages":""},"PeriodicalIF":1.6,"publicationDate":"2026-07-05","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148462794","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"An Effective Software Vulnerability Detection Method Based on Dual-Channel Convolutional Neural Network","authors":"Jinfu Chen, Ziyan Liu, Saihua Cai, Chenrui Zong, Xiaosong Chang, Bingbing Shao","doi":"10.1002/smr.70148","DOIUrl":"https://doi.org/10.1002/smr.70148","url":null,"abstract":"<div>\u0000 \u0000 <p>The detection of software vulnerabilities is an essential security issue in the field of cyberspace security. However, existing detectors often suffer from two main limitations: low computational efficiency owing to high cost and poor scalability when processing large-scale code, and inadequate feature representation capability due to their failure to capture sufficient semantic and structural properties of programs, which ultimately leads to decreased detection accuracy. To address these challenges and improve detection accuracy, we propose MCV-DCNN, a novel software vulnerability detection method based on multi-scale centrality-weighted image generation and dual-channel convolutional neural networks. The proposed MCV-DCNN method leverages multiple node centrality metrics to transform source code into image representations, thereby capturing a more holistic view of the code's structural properties. It also employs a dual-channel convolutional neural network architecture to independently process one image with each channel in parallel. This design mitigates information interference and redundancy between the representations during feature learning. We evaluate MCV-DCNN method on a public dataset with 33,362 C/C++ programs. Experimental results demonstrate that it significantly improves vulnerability detection accuracy while maintaining efficiency. Furthermore, the evaluation on 26,193 real-world functions shows that MCV-DCNN outperforms baselines by detecting 505 additional vulnerabilities, validating its effectiveness in practical settings.</p>\u0000 </div>","PeriodicalId":48898,"journal":{"name":"Journal of Software-Evolution and Process","volume":"38 7","pages":""},"PeriodicalIF":1.6,"publicationDate":"2026-07-02","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148358171","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}