极端不平衡数据下的信用卡欺诈检测:数据级算法的比较研究

IF 1.7 4区 计算机科学 Q3 COMPUTER SCIENCE, ARTIFICIAL INTELLIGENCE
Amit Singh, R. Ranjan, A. Tiwari
{"title":"极端不平衡数据下的信用卡欺诈检测:数据级算法的比较研究","authors":"Amit Singh, R. Ranjan, A. Tiwari","doi":"10.1080/0952813X.2021.1907795","DOIUrl":null,"url":null,"abstract":"ABSTRACT Credit card fraud is one of the biggest cybercrimes faced by users. Intelligent machine learning based fraudulent transaction detection systems are very effective in real-world scenarios. However, while designing these systems, machine learning approaches suffer from the problem of imbalanced data, i.e. imbalanced class distribution. Therefore, balancing the dataset becomes an imperative sub-task. Investigation of state-of-the-art approaches reveals that there is a need for a systematic study of class imbalance handling strategies to design an intelligent and capable system to detect the fraudulent transaction. This work aims to provide a comparative study of different class imbalance handling methods. To compare the effectiveness and efficiency of different class imbalance approaches in conjunction with state-of-the-art classification approaches, we have performed an extensive experimental study. We compared these methods on many performance indicators such as Precision, Recall, K-fold Cross-validation, AUC-ROC curve and execution time. In this study, we found that the Oversampling followed by Undersampling methods performs well for ensemble classification models such as AdaBoost, XGBoost and Random Forest.","PeriodicalId":15677,"journal":{"name":"Journal of Experimental & Theoretical Artificial Intelligence","volume":"1 1","pages":"571 - 598"},"PeriodicalIF":1.7000,"publicationDate":"2021-04-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"28","resultStr":"{\"title\":\"Credit Card Fraud Detection under Extreme Imbalanced Data: A Comparative Study of Data-level Algorithms\",\"authors\":\"Amit Singh, R. Ranjan, A. Tiwari\",\"doi\":\"10.1080/0952813X.2021.1907795\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"ABSTRACT Credit card fraud is one of the biggest cybercrimes faced by users. Intelligent machine learning based fraudulent transaction detection systems are very effective in real-world scenarios. However, while designing these systems, machine learning approaches suffer from the problem of imbalanced data, i.e. imbalanced class distribution. Therefore, balancing the dataset becomes an imperative sub-task. Investigation of state-of-the-art approaches reveals that there is a need for a systematic study of class imbalance handling strategies to design an intelligent and capable system to detect the fraudulent transaction. This work aims to provide a comparative study of different class imbalance handling methods. To compare the effectiveness and efficiency of different class imbalance approaches in conjunction with state-of-the-art classification approaches, we have performed an extensive experimental study. We compared these methods on many performance indicators such as Precision, Recall, K-fold Cross-validation, AUC-ROC curve and execution time. In this study, we found that the Oversampling followed by Undersampling methods performs well for ensemble classification models such as AdaBoost, XGBoost and Random Forest.\",\"PeriodicalId\":15677,\"journal\":{\"name\":\"Journal of Experimental & Theoretical Artificial Intelligence\",\"volume\":\"1 1\",\"pages\":\"571 - 598\"},\"PeriodicalIF\":1.7000,\"publicationDate\":\"2021-04-03\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"28\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"Journal of Experimental & Theoretical Artificial Intelligence\",\"FirstCategoryId\":\"94\",\"ListUrlMain\":\"https://doi.org/10.1080/0952813X.2021.1907795\",\"RegionNum\":4,\"RegionCategory\":\"计算机科学\",\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"Q3\",\"JCRName\":\"COMPUTER SCIENCE, ARTIFICIAL INTELLIGENCE\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"Journal of Experimental & Theoretical Artificial Intelligence","FirstCategoryId":"94","ListUrlMain":"https://doi.org/10.1080/0952813X.2021.1907795","RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q3","JCRName":"COMPUTER SCIENCE, ARTIFICIAL INTELLIGENCE","Score":null,"Total":0}
引用次数: 28

摘要

信用卡诈骗是用户面临的最大的网络犯罪之一。基于智能机器学习的欺诈交易检测系统在现实世界中非常有效。然而,在设计这些系统时,机器学习方法受到数据不平衡问题的困扰,即类分布不平衡。因此,平衡数据集成为一项势在必行的子任务。对最先进的方法的调查表明,有必要对班级不平衡处理策略进行系统研究,以设计一个智能和有能力的系统来检测欺诈性交易。本文旨在对不同的类不平衡处理方法进行比较研究。为了比较不同的类不平衡方法与最先进的分类方法的有效性和效率,我们进行了广泛的实验研究。我们在精密度、召回率、K-fold交叉验证、AUC-ROC曲线和执行时间等性能指标上对这些方法进行了比较。在本研究中,我们发现过采样后欠采样方法在AdaBoost、XGBoost和Random Forest等集成分类模型中表现良好。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
Credit Card Fraud Detection under Extreme Imbalanced Data: A Comparative Study of Data-level Algorithms
ABSTRACT Credit card fraud is one of the biggest cybercrimes faced by users. Intelligent machine learning based fraudulent transaction detection systems are very effective in real-world scenarios. However, while designing these systems, machine learning approaches suffer from the problem of imbalanced data, i.e. imbalanced class distribution. Therefore, balancing the dataset becomes an imperative sub-task. Investigation of state-of-the-art approaches reveals that there is a need for a systematic study of class imbalance handling strategies to design an intelligent and capable system to detect the fraudulent transaction. This work aims to provide a comparative study of different class imbalance handling methods. To compare the effectiveness and efficiency of different class imbalance approaches in conjunction with state-of-the-art classification approaches, we have performed an extensive experimental study. We compared these methods on many performance indicators such as Precision, Recall, K-fold Cross-validation, AUC-ROC curve and execution time. In this study, we found that the Oversampling followed by Undersampling methods performs well for ensemble classification models such as AdaBoost, XGBoost and Random Forest.
求助全文
通过发布文献求助,成功后即可免费获取论文全文。 去求助
来源期刊
CiteScore
6.10
自引率
4.50%
发文量
89
审稿时长
>12 weeks
期刊介绍: Journal of Experimental & Theoretical Artificial Intelligence (JETAI) is a world leading journal dedicated to publishing high quality, rigorously reviewed, original papers in artificial intelligence (AI) research. The journal features work in all subfields of AI research and accepts both theoretical and applied research. Topics covered include, but are not limited to, the following: • cognitive science • games • learning • knowledge representation • memory and neural system modelling • perception • problem-solving
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
确定
请完成安全验证×
copy
已复制链接
快去分享给好友吧!
我知道了
右上角分享
点击右上角分享
0
联系我们:info@booksci.cn Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。 Copyright © 2023 布克学术 All rights reserved.
京ICP备2023020795号-1
ghs 京公网安备 11010802042870号
Book学术文献互助
Book学术文献互助群
群 号:481959085
Book学术官方微信