Ishita Kheria, Dhruv Gada, Vijay Harkare, Ruhina Karani
{"title":"Unearthing Code Smells: An In-Depth Exploration of Machine Learning Techniques in Code Smell Detection","authors":"Ishita Kheria, Dhruv Gada, Vijay Harkare, Ruhina Karani","doi":"10.1002/smr.70144","DOIUrl":"https://doi.org/10.1002/smr.70144","url":null,"abstract":"<div>\u0000 \u0000 <p>Code smells are clear indicators of design flaws that compromise maintainability, reliability, and overall performance in software systems. As systems become complex, traditional detection methods become inadequate, prompting the need for automated approaches. In this study, we evaluate the potential of conventional machine learning algorithms for detecting critical code smells. Specifically, we examine Classification and Regression Trees (CART), Gradient Boosting Classifiers, Extreme Gradient Boosting (XGBoost), RuleFit, Adaptive Boosting (AdaBoost), <i>K</i>-Nearest Neighbors (KNN), Stochastic Gradient Descent (SGD), and Random Forest to identify six key code smell categories: Large Class, Long Parameter List, Switch Statements, God Class, Data Class, and Long Method. Advanced resampling techniques such as SMOTE–ENN and SMOTE–Tomek are employed to mitigate dataset imbalances alongside tree-based feature extraction for effective dimensionality reduction. Systematic experimentation demonstrates that SMOTE-based strategies significantly enhance detection performance, yielding robust results across multiple evaluation metrics. By focusing on the core challenges of code smell detection, this study provides a cost-effective alternative to deep learning methods, which often demand extensive computational resources and large labeled datasets.</p>\u0000 </div>","PeriodicalId":48898,"journal":{"name":"Journal of Software-Evolution and Process","volume":"38 6","pages":""},"PeriodicalIF":1.8,"publicationDate":"2026-06-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148173990","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Daniel Rodriguez-Cardenas, David N. Palacio, Anna Schmedding, Yiyang Lu, Aadil Mallick, Bill Hudson, Chris Gourley, Michael Roytman, Chris Shenefiel, Evgenia Smirni, Denys Poshyvanyk
{"title":"On Predicting Vulnerability Severity Using In-Context Learning: An Industrial Case Study","authors":"Daniel Rodriguez-Cardenas, David N. Palacio, Anna Schmedding, Yiyang Lu, Aadil Mallick, Bill Hudson, Chris Gourley, Michael Roytman, Chris Shenefiel, Evgenia Smirni, Denys Poshyvanyk","doi":"10.1002/smr.70126","DOIUrl":"https://doi.org/10.1002/smr.70126","url":null,"abstract":"<p>Modern software systems require earlier and more scalable vulnerability severity assessment to reduce exposure to high-impact security flaws. Security analysts typically assign CVSS scores, but this manual triage does not scale with the growth of disclosed vulnerabilities and often depends on cloud LLM services that raise confidentiality concerns. This paper presents an industrial case study on predicting CVSS v3.1 scores directly from vulnerable C/C++ snippets using <i>in-context learning</i> with locally deployable, <i>open-source</i> LLMs. We compare <i>Propietary</i> data with the <i>Big-Vul</i> dataset, showing sufficiently aligned CVSS distributions to justify <i>Big-Vul</i> as a proxy for industrial data when constructing prompt-based testbeds. We then vary in-context configurations and model parameters, evaluating <i>CodeLlama2-7B</i>, <i>CodeLlama2-13B</i>, <i>Mistral-7B</i>, <i>gpt-oss</i>, and <i>GPT4o-mini</i> using mean squared error (MSE) and feasibility metrics. Our results show that medium-sized open-source code models, particularly <i>CodeLlama2-7B</i>, can approximate the best cloud performance for CVSS regression when guided by lightweight, output-constraining prompts, offering a practical, privacy-preserving building block for severity triage in industrial settings.</p>","PeriodicalId":48898,"journal":{"name":"Journal of Software-Evolution and Process","volume":"38 6","pages":""},"PeriodicalIF":1.8,"publicationDate":"2026-06-02","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://onlinelibrary.wiley.com/doi/epdf/10.1002/smr.70126","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148129424","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Exploring Success Factors and Practices for Mob Programming: Extraction From MLR and Empirical Study","authors":"Rafi Ullah, Muhammad Ilyas","doi":"10.1002/smr.70130","DOIUrl":"https://doi.org/10.1002/smr.70130","url":null,"abstract":"<div>\u0000 \u0000 <p>Mob Programming (MP) is an emerging, relatively new, and unexplored programming technique that is becoming increasingly popular and receiving attention in the software industry. It is a collaborative programming approach where the entire development team works together on a single task. This study explores the success factors (SFs) and their implementation practices for the effective use of MP in the software industry. The research was conducted in two phases: First, a multivocal literature review (MLR) analyzed formal and gray literature (GL), reviewing 76 primary studies to identify SFs and their implementation practices for adopting MP. In the second phase, an empirical survey was conducted, involving 106 software industry experts from 30 countries, to validate the MLR findings. The study identified 12 SFs, among which four factors, “professional communication and team collaboration,” “skills and wills,” “proper feedback sessions,” and “self-organized, self-motivated, and cooperative teams” were ranked as critical success factors (CSFs). Similarly, 126 implementation practices were discovered for the implementation of the identified SFs. The findings suggest that empirical methods are commonly used in MP studies and interest in MP has grown significantly over the past decade. The survey confirmed that the identified SFs are prevalent across projects and organizations of varying sizes. These results offer valuable insights to enhance developers' competence and promote the successful implementation of the MP in the software industry.</p>\u0000 </div>","PeriodicalId":48898,"journal":{"name":"Journal of Software-Evolution and Process","volume":"38 6","pages":""},"PeriodicalIF":1.8,"publicationDate":"2026-06-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148128167","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Yanze Wang, Qi Feng, Jingyue Li, Shanshan Li, He Zhang
{"title":"Model-Driven Full-Lifecycle Support Reference Framework for Blockchain-Based Application Development","authors":"Yanze Wang, Qi Feng, Jingyue Li, Shanshan Li, He Zhang","doi":"10.1002/smr.70111","DOIUrl":"https://doi.org/10.1002/smr.70111","url":null,"abstract":"<div>\u0000 \u0000 <p>Blockchain technology has the potential to improve various traditional industries by enabling decentralized, tamper-resistant data management and transparency. These characteristics, however, make the development of blockchain-based applications demanding and costly. Model-driven development (MDD) offers a paradigm by abstracting low-level details into higher level models, thereby potentially enhancing efficiency and reducing errors. Yet, existing MDD-based solutions are fragmented and lack comprehensive support for the software development lifecycle (SDLC) in a consistent manner. Moreover, the intrinsic complexity of blockchain and its heterogeneous infrastructures introduces further challenges across SDLC stages, highlighting the need for a unified framework that systematically integrates MDD with blockchain-specific development. This research proposes an MDD-based reference framework that supports a consistent modeling process across the entire SDLC of blockchain-based applications, from requirements to deployment. The framework adapts various modeling techniques to address blockchain-specific features such as decentralization, one-time smart-contract deployment, and immutable data record. The model transformation methods are improved to support the seamless full SDLC development process of blockchain-based systems. The proposed reference framework is platform independent at modeling level, enabling compatibility with heterogeneous blockchain platforms. Its adaptability was validated through a prototype implementation supporting both Ethereum and Hyperledger Fabric. Further, a comparative case study was conducted on a cross-border food supply chain. The results demonstrate the feasibility and adaptability of the MDD-based reference framework, as well as improved efficiency for blockchain-based system development, which promotes a broader adoption of blockchain technology.</p>\u0000 </div>","PeriodicalId":48898,"journal":{"name":"Journal of Software-Evolution and Process","volume":"38 6","pages":""},"PeriodicalIF":1.8,"publicationDate":"2026-06-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148128171","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"AIBRD: Automated Identification of Bug Report Description","authors":"Ning Li, Hong Ma, Yuzhou Liu, Peng Zhang","doi":"10.1002/smr.70131","DOIUrl":"https://doi.org/10.1002/smr.70131","url":null,"abstract":"<div>\u0000 \u0000 <p>Bug report descriptions usually contain three types of key information: observed behavior, expected behavior, and steps to reproduce. These elements help developers identify, reproduce, and fix software bugs. Manually identifying them in large bug reports is time-consuming, so automation is needed. However, existing methods mainly rely on features extracted from individual sentences and neglect contextual information, which leads to poor performance in bug report sentence classification. To address this problem, we construct a semantic knowledge base to capture contextual information and then employ a multidimensional fusion strategy to integrate features from multiple sources for bug report sentence classification. The semantic knowledge base consists of three components: a keyword co-occurrence matrix, synonym sets, and causality pair sets, which together capture contextual information. We then fuse the semantic knowledge base features with BERT word embeddings through a cross-attention mechanism, followed by the concatenation of discourse pattern features to improve classification accuracy. We evaluated our method on a context-dependent dataset, achieving an accuracy of 73.4% and a Hamming loss of 0.167. Experimental results indicate that AIBRD outperforms existing methods. Compared with the state-of-the-art baseline DEMIBuD-H, AIBRD improves macro-average accuracy by 3.3% and macro-average F1 score by 7.0%.</p>\u0000 </div>","PeriodicalId":48898,"journal":{"name":"Journal of Software-Evolution and Process","volume":"38 6","pages":""},"PeriodicalIF":1.8,"publicationDate":"2026-06-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148128168","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"From License Understanding to Copyright Violation Detection: A Study of Libraries in GitHub Python Projects","authors":"Wenhua Yang, Yuanting Liu, Minxue Pan","doi":"10.1002/smr.70129","DOIUrl":"https://doi.org/10.1002/smr.70129","url":null,"abstract":"<div>\u0000 \u0000 <p>Open-source licenses grant developers significant flexibility to use, modify, and distribute code, but they also introduce obligations that may pose legal or compliance risks. Among these, copyright-related terms are especially critical. Ensuring compliance with such terms is challenging due to the widespread use of third-party libraries and the diversity of license types. While license-related issues have attracted research interest, most prior work targets a narrow set of licenses and often overlooks copyright terms. To bridge this gap, we propose an automated approach for detecting copyright term violations in open-source projects that incorporate third-party libraries, covering a wide range of Open Source Initiative-approved licenses. Our method extracts licenses from both projects and their dependencies, uses large language models to interpret key copyright terms, and identifies potential violations. We applied our approach to 500 popular Python projects on GitHub and found that about 10% exhibited at least one violation. These results highlight the complexity of third-party license management and the need for improved compliance tools. By analyzing tens of thousands of licenses, we also uncovered common patterns in license usage. To foster further research and promote proactive compliance, we release our tool and dataset to the community.</p>\u0000 </div>","PeriodicalId":48898,"journal":{"name":"Journal of Software-Evolution and Process","volume":"38 6","pages":""},"PeriodicalIF":1.8,"publicationDate":"2026-06-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148128166","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"BPMN Extensions Updated: A Systematic Literature Review","authors":"Milene Cavalcante, Enyo Gonçalves, João Araújo","doi":"10.1002/smr.70127","DOIUrl":"https://doi.org/10.1002/smr.70127","url":null,"abstract":"<p>Business process model and notation (BPMN) is a widely used modeling language for representing business processes. BPMN has been extended to represent particular domains or application areas, or to improve practical aspects. BPMN extensions have been the subject of recent studies. This study aims to identify and analyze the BPMN extensions proposed between 2019 and 2025 (until July). This paper followed the steps of a systematic literature review (SLR) to identify and analyze the BPMN extensions. We first performed automatic and manual search and selection process to obtain the primary studies. We also selected papers using forward and backward snowballing techniques. In this analysis, we checked whether the studies reviewed presented clear meaning for the new extended constructs, including name, formal description, purpose, and application in the proposed context. The research identified 84 studies that resulted in new BPMN extensions, analyzing essential attributes, application areas, type of extension, syntax level, and adherence to Object Management Group (OMG) standards. The results showed that 95.7% of the extensions do not preserve the original BPMN syntax, and 55% do not support the modeling of the new elements, making their adoption difficult. In addition, we found that the presentation of the extensions lacks standardization: 55% of the studies do not provide meaning, 18.8% present partial definitions, and only 26.3% offer complete and structured descriptions. These findings reinforce the need for more consistent guidelines for defining and implementing BPMN extensions, contributing to future research into the extensibility of the notation. In this paper, we presented the results of an updated SLR to identify BPMN extensions. Our goal was to improve the understanding of how BPMN has been extended through evidence in literature. We identified 84 papers that extended BPMN in the last years (2019–2025), and five research questions were answered and discussed.</p>","PeriodicalId":48898,"journal":{"name":"Journal of Software-Evolution and Process","volume":"38 6","pages":""},"PeriodicalIF":1.8,"publicationDate":"2026-06-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://onlinelibrary.wiley.com/doi/epdf/10.1002/smr.70127","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148128209","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"","authors":"","doi":"","DOIUrl":"","url":null,"abstract":"","PeriodicalId":48898,"journal":{"name":"Journal of Software-Evolution and Process","volume":"38 5","pages":""},"PeriodicalIF":1.8,"publicationDate":"2026-05-19","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148082675","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"","authors":"","doi":"","DOIUrl":"","url":null,"abstract":"","PeriodicalId":48898,"journal":{"name":"Journal of Software-Evolution and Process","volume":"38 5","pages":""},"PeriodicalIF":1.8,"publicationDate":"2026-05-17","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148082932","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"","authors":"","doi":"","DOIUrl":"","url":null,"abstract":"","PeriodicalId":48898,"journal":{"name":"Journal of Software-Evolution and Process","volume":"38 5","pages":""},"PeriodicalIF":1.8,"publicationDate":"2026-05-14","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148078980","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}