VIETNAM NATIONAL UNIVERSITY HO CHI MINH CITY UNIVERSITY OF INFORMATION TECHNOLOGY ADVANCED PROGRAM IN INFORMATION SYSTEMS TRUONG MINH KHIET - 19520628 LUONG TIEN THUAN HAI - 19521462 THESIS CREDIT CARD FRAUD DETECTION USING MACHINE LEARNING AND DEEP LEARNING ALGORITHMS BACHELOR OF ENGINEERING IN INFORMATION SYSTEMS THESIS ADVISOR DR. CAO THI NHAN HO CHI MINH CITY, 2023 ACKNOWLEDGMENTS First, allow my group to express our deep gratitude to Ms. Cao Thi Nhan for her dedication and dedication to guiding the group throughout the process of researching and conducting their graduation thesis. Her patience, extensive knowledge and dedication helped the group overcome difficulties in the research process and complete the project on time as prescribed by the school.
Next, our group would also like to sincerely thank the teachers at the University of Information Technology - National University of Ho Chi Minh City for teaching us valuable knowledge necessary for our group to be able to Complete your thesis, as well as serve your future work well. My group also wants to thank the University of Information Technology - Ho Chi Minh City National University for creating learning and practicing conditions for our group to complete their thesis and course. Finally, our group would like to send our sincere thanks to our families and those who have helped, supported, and encouraged us to feel secure in researching and completing our thesis. TABLE OF CONTENTS cw ACKNOWLEDGMENTTS.- HH HH TH HH Hà HH TT nh ng 1 TABLE OF CONTIENTTS.- 1n HT TH HH HH Tnhh nh nh 2 LIST OF FIGURES.- G1 nHnH HH TT TH HH Hà HH TT ng 5 I0R09)30.- 6 33 v99 HH HH ng ưkp 10 1.
Reason for choosing the {OpIC.-- 5 5 2+ 119119 1 9 1 vn ng tr 10 1.- d1 1E TT TH HH nh 13 I0)) ung i00 0n. Research scope Of the fODIC. - -- 6 SE 1911 11 1 vn HH ng ngư 16 CHAPTER 2. THEORETICAL BACKGROUND AND RELATED WORKS.
Fraud et€CfIOH.- 5 5 x11 TH TH HH nh Hệ 19 2.-- --- 5 << 1x vn HH Hệ 21 2. Extreme learning methO(.-- s6 + +31 91119119 1 911 9v ng ng nề, 21 "2899 soi. LH HH HH Hệ, 22 2.- 5 s2 HH HH gi 23 2. Support vector machine (SVM).-- - HH ng ng kg 24 2.-- 5 G11 E11 vn ng Hết 25 "ĐC: 909.
Convolutional Neural Network (CNN).- c2 110101110 111 ng 0 vn reg 31 I1 0117 --. Experimental environment 000rraẳaẳ.- 3x HH TH HH ng grư 33 c3 0i. Applied machine learning & ensemble learning techn1ques. K-nearest neighbours CKRNN).
Support vector machine (SVM) ou. eee ee eee 5 1kg, 56 E6 bu. Extreme learning method (ELM). Applied deep leaning techniques.
Baseline ion ae. Convolutional neural network (CNN). Ăn vn HH nghiệt 62 3. Performance-evaluation €aSUT€S.
(<< << << 11111EE SĐT ng 055 ket 64 3. si HH ng ớt 64 4. HH TH HH HH HH Hết 64 3. Use machine learning aÏØOTIthH§.
Convolutional neural network (CNNN). Compare results among alGorithms. G1911 110 1h nh HH nh trờ 74 4. Limitations and development diTeCfIONS.- 6 SH TH HH HH HH HH75 4.
Development (IT€CfIOTNS .-- --- 1 11 2kg TH HH giết 77 LIST OF FIGURES Le Figure 1 - Illustration of ÍTaUd.- - << + E1 ng HH nh 19 Figure 2 - ELM's algOrithim.- xxx x19 TH ng HH ng Hưng gà 22 Figure 3 - RFP's aÏEOTItH1.-- 111 1 93 93 2 1 TH HH HH Hưng ngư 24 Figure 4 - SVM's alEOTIthim.- 6 + x2 TH TH TH Hàng HH tiệt 24 Figure 5 - Logistic regression's aÏØOTIfH.- --- 6 + s11 k*91 9 9v vn re, 25 Figure 6 - Logistic regression's aÏØOTIthim.-- ----s «+ + + +*kxeeerseeeereerereere 26 Figure 7 - Pooling ÍAW€T.-- c1 95993111 11h TT HH Hư 28 Figure 8 - CNN output ÏAY€T. G1 TH TH 30 Figure 9 - Application of dropout over neural netWOTK.- -- « +««es<+sx+se+sx+ 3l Figure 10 - Connect data through drive. --- 5 5 + + x** EEEseEEeereereereereerske 39 Figure 11 - Incorrect data handÏIng.- - c6 6+1 1231191119113 1 9v vn tre40 Figure 12 - Check the columns in data .- -- - <1 * 1£ +3 E+*EE+eeereeeeeeeersreere 41 Figure 13 - Chart of Class Distributions. -- --- < + S3 + kseierrrereere 42 Figure Si 0i.
43 Figure 15 - Allocating data for training and f€StInE.- s6 +- + +ss£secssesseesee 44 Figure 16 - Result of Label Distribution. cece ese eseeseeseeseeeeeseseeseeeeeseseeseeaeens 45 Figure 17 S311iiï1 6ï 0. 47 Figure 19 - Imbalance Correlation Matrix .- - s5 c1 + seerseeeeeerereere 48 Figure 20 - Balance Correlation ÏMafT1X.- -- << 1x19 1 ng ng 48 Figure 21 - Negative COrr€ÏfÏOTI.- -- 6 53k TH TH TH TH TH TH HH ng già 49 Figure 22 - Positive COTT€ÏAfIOT. 6 333 vn TH TH TH TH TH HH Hiệp 49 Figure 23 - Result of reduce OU{ÏI€T-.- - --- + + + + 3x E*kESeeEseeereeeeereeeereere 50 Figure 24 - Result of clustering methO(s.-- --- + + + + seeerseeereerereere 51 Figure 25 - Score model Decision Tree .cceesceesseeeseeesneeeseeceseeeeaeeesaeeeseeeeeeesaes 53 Figure 26 - Score model KNN.
LH HH TH HH HH 54 Figure 27 - Score model Ï.-- «xxx 3 3 9191 HH ng HH rưệt 55 Figure 28 - Score model SM. kg TH TH HH HH ng 56 Figure 29 - Score model LOGISTIC REGRESSION. c2 S2sesee 58 Figure 30 - Score model XG BOOS.-- -- c1 1111 1v HH ng re 59 Figure 31 — Score of ELM Model .-- -- << 1139113 1S kg re 61 Figure 32 - Cross-Entropy “s aÏØOTIfhm.- -- - + 1311131 E+EESeeEseeeerrersreere 63 Figure 33 - Formula to calculate aCCUTACY. - c c1 11131113111 1 ve rre 63 Figure 34 - Formula to calculate DF€CISIOTI.- -- <5 +5 *+2+£+*E+veeEeeeeeeereeeers 64 Figure 35 - Formula to calculate Fl- SCOTG.
5 + + + * + E+*vEEseeeeeersreers 64 Figure 36 - Formula to calculate r€CaÏÏÏ.-- -- -- ¿+ +++ + *++‡++*kE+veeEseereeeeeeeerereere 64 Figure 37 - Score the MOel.- -- -- <1 1311131111911 9111 vn HH ng 66 Figure 38 - Score Model .- ---- <5 1 1 HT TH HH Hư 67 LIST OE TABLES Le Table 1.1 - Configure the execution COIDUẨ€T .2 - Configure the Python 3 Google Compute Engine.1 - Accuracy based results of deep learning algorithms.2 - The list of features available in the CCF dafaset.3 - Characteristics of the dataset. 1v ng ng, 36 Table 3.4 - Characteristics of the dataset .- 5 6s vn ngư 37 Table 3.5 - Characteristics of the đafaS€(.- 2G G1 111v SH re 52 Table 3.6 - Table score of Decision “ÏTT€€.7 - Table score Of KNN. LH HH HH HH HH, 54 Table 3.8 - Table score Of RE. c5 11k TH TH HH HH 55 Table 3.9 - Table score Of SVM.
LH HH HH TH TH TH HH HH, 57 Table 3.10 - Table score of Decision “TTT€.- --- 6 sa s11 139113511931 1 31x krse 58 Table 3.11 - Table score of XG BooSt .- 0 n1 ng ng ngư 60 Table 3.12 - Table score Of ELUÌM.-- 5G s11 E91 211191 9111 11t HH ng tr 61 Table 3.13 - Table score of CINN.-- HH HH TH HH HH, 67 Table 3.14 - Table score of BaseÏine. ng ng ng ry 68 Table 3.16 - Results of the algorithms when running on the initial data.17 - Results of algorithms when running on only positive correlations .18 - Results of algorithms when running on only negative correlations .19 - Results of algorithms when running on negative correlations and positive COTT€ÏAfÏOTN§. -- 2 222211101110101033111 11T ng g0 0 HH 70 Table 3.20 - The results of the algorithms when running on the processed data set 71 Table 3.21 - Comparison between 3 models witth different data sefs. 72 ABSTRACT CCF represents an escalating threat to financial institutions, as fraudsters continually develop novel methods to exploit vulnerabilities.
The need for a robust classifier is paramount to adjust to the ever-changing landscape of fraud. The primary goal of a fraud detection system is to precisely forecast instances of fraud while minimizing the occurrence of false positives. The effectiveness of machine learning (ML) approaches varies across diverse business scenarios, where the characteristics of input data play a crucial role in determining the suitable ML techniques. In the realm of CCF detection, pivotal factors influencing model performance encompass the quantity of features, transaction volume, and inter-feature correlations.
Deep learning (DL) techniques, exemplified by Convolutional Neural Networks (CNNs) and their layers, are essential for text processing and serve as a foundational model. Applying these techniques to identify fraudulent credit card activities surpasses the capabilities of traditional algorithms. A comparative evaluation of algorithmic performances highlights the CNN with 20 layers and the XG Boost model as the leading methods, achieving an impressive accuracy of 94. Numerous sampling methods are employed to boost the performance of existing examples, yet they exhibit a noticeable decline in efficacy when applied to unseen data.
Intriguingly, there is an observed enhancement in performance on unseen data with an increased class imbalance. Future research initiatives might explore the integration of more sophisticated deep learning methods to elevate the overall performance of the model proposed in this study. Reason for choosing the topic Credit Card Fraud (CCF) is a form of identity theft where unauthorized individuals use stolen or fake credit card details to conduct transactions. The use of lost, stolen, or forged credit cards can lead to fraudulent activities.
Fraud without physical card possession or the use of credit card information in online transactions is becoming more common due to the growing popularity of online shopping. The increase in fraud, such as CCF, is a result of the expansion of electronic banking and various online payment platforms, resulting in billions of dollars in annual damages. Detecting CCF has become a top priority in this era of digital payments. As a business owner, it is clear that the future is moving towards a cashless society.
As a result, traditional payment methods may not be effective for business expansion, as customers may not always carry cash. Debit and credit card payments are currently subject to fees by companies. As a result, businesses must update their infrastructure to accept all forms of payments. This situation is expected to become even more critical in the coming years [1].
In 2020, approximately 393,207 incidents of Credit Card Fraud (CCF) occurred, accounting for a significant portion of the reported 1.4 million identity theft cases [4]. Presently, CCF ranks as the second most prevalent type of identified identity theft, following closely behind government document and benefits fraud [5]. More specifically, 365,597 instances of fraud in 2020 were associated with the unauthorized opening of new credit card accounts [10]. Comparing the years 2019 and 2020 reveals a substantial 113% increase in identity theft complaints, with reports of credit card identity theft witnessing a notable surge of 44.
On a global scale, payment card theft resulted in an economic loss of $24.26 billion the preceding year. It is noteworthy that the United States, accounting for 38.6% of reported card fraud losses in 2018, emerges as the most susceptible country to credit card theft. Page | 10 Hence, it is essential for financial institutions to prioritize the implementation of an automated system for detecting fraud. The objective of detecting Credit Card Fraud (CCF) entails creating a Machine Learning (ML) model that leverages historical data from credit card payment transactions.
This model aims to distinguish between transactions that are fraudulent and those that are legitimate, using this insight to assess the authenticity of incoming transactions. Overcoming associated challenges requires tackling fundamental issues such as system responsiveness, cost considerations, and feature preprocessing.