VIETNAM NATIONAL UNIVERSITY HOCHIMINH CITY UNIVERSITY OF INFORMATION TECHNOLOGY ADVANCED PROGRAM IN INFORMATION SYSTEMS NGUYEN THI MY LAN- 16520651 LE NGOC UYEN VY- 16521472 AN APPROACH FOR FRAUD DETECTION IN FINANCIAL TRANSACTIONS USING MACHINE LEARNING METHODS BACHELOR OF ENGINEERING IN INFORMATION SYSTEMS THESIS ADVISOR Dr. CAO THI NHAN HO CHI MINH CITY, 2020 ASSESSMENT COMMITTEE The Assessment Committee is established under the Decision. by Rector of the University of Information Technology. ACKNOWLEDGMENTS oElœ The thesis topic will not be completed without any assistance.
So the authors implementing the topic gratefully give acknowledgment to their support and motivation during the graduation thesis project. First of all, we would like to thank the Lecturers of the University of Information Technology as well as the Lecturers of the Information Systems Faculty, who taught and provided the background solid knowledge throughout the time studying at school. This knowledge is an important basis for us to complete our graduation thesis. In particular, we would like to express my endless thanks and gratefulness to my supervisor, Dr.
Cao Thi Nhan. Her dedicated support and constant advice have helped us to complete our thesis well. Her words of encouragement and comments have greatly enriched and improved our work. Without her guidance, the thesis would not have been done effectively.
During the implementation of the thesis, we have tried to apply effectively what we have learned as well as learn new technologies to be able to complete the thesis in the best way. However, in the process of implementation, because of limited knowledge, experience, and study time, it is difficult to avoid shortcomings. Therefore, we hope to receive the comments of teachers to complete the thesis better as well as the necessary knowledge and skills. Thank you so much! Authors Nguyen Thi My Lan — Le Ngoc Uyen Vy TABLE OF CONTENTS woes ACKNOWLEDGMENTS.
TABLE OF CONTENTS .cccscscsssssssssssssssssscsenencnssssssececscsenescnsessesacecsenenesseseeseceees ii LIST OF FIGURES.ccssssssssssssrsssessnsnsessssncsesenensnsssssscsesenencnssssssecacsenensnseseeseseees M LIST OF TA BLES. 5555 S522 Y9911110181010118810101010101001 08ix LIST OF ABBREVIATIONS .cssssssssssssssssssssenenensssesssrecsesenensnsssessececsenensneesenes xi ABSTRACT Chapter 1 INTRODUCTION. cect os saaece MO casas se seseevevesenssssusesnevevesenrnnaees 1 1. Aims and Obj€CfÏV€S.
Languages, Tools and LibrarieS. ¿5525252 S+s+*2+s£+x+xezsexexsx 2 Chapter 2 BACKGROUND. Fraudulent transaction definitiO. Fraud Detection Approach.
Imbalanced dataset problems. Methods to solve the imbalanced problems .-- + + «+ +++<s<+ 7 Chapter 3 MACHINE LEARNING FOR FRAUD DETECTION.- Ăn TT HH HH Hư 10 3. Applying K Nearest Neighbor (KNN). Advantages and disadvantages.
Logistic Regression model. Apply Logistic Regression step-by-step „l5 3. Advantages and disadvantages .---- cà + St sseireeeiey 15 3. Support Vector Machine.
Advantages and Disadvantages. Precision and Recall. Receiver Operating Characteristic CUTV€. Data Analytics Pipeline.
Credit Card Dataset.- 5c St S ĐH HH HH re. cece th HH HH HH HH HH 24 4. Làn 12 HH HH HH HH0 gi 25 4. Synthetic Financial DataS€(S.--- cành HH HH HH it 48 4.
óc 1à TT HT Hà nh Hệ 48 4. PT€DTOC€SSÏNE. án TH HH HH HH gi 56 4. LH TH HT TH TH HT HH TH TH HH Hy 58 4.
HH HH HH HH HH HH Hi 73 Chapter 5 CONCLUSIONS AND FUTURE WORKS. The results achieved. Ăn 01011000010101010000001010101000000000 186 1 iv LIST OF EIGURES voles Figure 1: Identity theft reports in the United States. -------+©-<++ xii Figure 2: Most common types of identity thet.
-- «655 S+c+xsserersreree xiii Figure 3: Credit card fraud reports by Y€AT.- --- +5 + 55252 Sc+x+xzesersrsrseree xiv Figure 3.1: Applying the K Nearest Neighbor (KNN) algorithm step-by-step [8] .2: Common distance metrics [9] .3: Graph of Logistic curve where o=0 and B=1 [1 I].4: Graph of Sigmod function [12] .6: Precision and Recall.7: Receiver Operating Characteristic curve model .1: Data analytics Pipeline .2: The original dataset .3: Dataset Class Distribution .4: Transactions amount distrIDutiOI.5: Transaction Time distributiON. eee cesses ee eeeseseneeeeeeseneees 28 Figure 4.6: Dataset after scaling features “Time” and “Amounf”.7: Confution matrix of testing data when applying KNN algorithm in the original Credit Card dataset .8: Confution matrix of testing data when applying LR algorithm in the original Credit Card đafAS€K. s6 5 111 1n TH TH ng HH nh Hư 32 Figure 4.9: ROC curve of LR in the imbalanced Credit Card dataset.10: Confution matrix of testing data when applying KNN algorithm in the original Credit Card dataset. ¿<5 1xx 91 121111510101 111211101010 1010 ty 34 Figure 4.11: ROC curve of SVM in the imbalanced Credit Card dataset.12: Class distribution in the subsample after using Random Undersampling —.13: Confusion matrix of testing data when applying KNN with undersampling in the Credit Card dataset .14: Confusion matrix of testing data when applying LR with undersampling in the Credit Card dataset .15: ROC curve of LR in the balanced Credit Card dataset using Random Undersammpling.16: Confusion matrix of testing data when applying SVM with undersampling in the Credit Card dataset .17: ROC curve of SVM in the balanced Credit Card dataset using Random ndersampling.18: Class distributions of the Credit Card dataset after applying SMOTE 41 Figure 4.19: Confusion matrix of testing data when applying KNN with SMOTE in the Credit Card dataset .20: Confusion matrix of testing data when applying LR with SMOTE in the Credit Card dataset .21: ROC curve of LR in the balanced Credit Card dataset using SMOTE 44 Figure 4.22: Confusion matrix of testing data when applying SVM with SMOTE in the Credit Card dataset .23: ROC of SVM in balanced Credit Card dataset using SMOTE.24: The comparison chart of the "accuracy" between different algorithms on the Credit Card đia(ASCTL.25: The chart compares the "Fl-score" of the algorithms on the Credit Card datasets .26: shows number of transactions which are the actual fraud per tramsaction tyPe.
St TT HT HH TH HH Hy 52 Figure 4.27: Distribution Of types .--- ¿+ kg ng it 52 Figure 4.28: Distribution of EFauid.-- ceeceeececseseeeneeseseseseeeeeeeeseneeeeeeaeae 53 vi Figure 4.29: Distribution of the feature FlaggedFraud .30: Original Paysim dataset .- cà 1n HH ng it 56 Figure 4.31: Data “isFraud” distributiOI.33: Confusion matrix of testing data when applying K Nearest Neighbors algorithm in the original Paysim dataset .34: Confution matrix of testing data when applying LR algorithm in the original Paysim đa(AS€(L.35: ROC curve of LR in the imbalaced Paysim dataset.36: Confution matrix of testing data when applying SVM algorithm in the Original Paysim điafAS€.37: ROC curve of SVM in the imbalanced Paysim dafaset.38: Class distribution in the subsample after using Random Undersampling Figure 4.39: Confusion matrix of testing data when applying KNN with undersampling in the Paysim dataset .40: Confusion matrix of testing data when applying LR with undersampling in the Paysim dataset .41: ROC curve of LR in the balanced Paysim dataset using Random ndersaimpling.42: Confusion matrix of testing data when applying SVM with undersampling in the Paysim dataset .43: ROC curve of SVM in the balanced Paysim dataset using Ramdom UnderSampÏinng. - - ‹- + 1k 1 1 1 1E TT TT HT TH TH TH Hư 69 Figure 4.44: Confusion matrix of testing data when applying LR with SMOTE in the Paysim dafASCL.45: ROC curve of LR in the balanced Paysim dataset using SMOTE.46: Confusion matrix of testing data when applying SVM with SMOTE in the Paysim dataset. - - 11k TT HH HH HH HH TH HH Hy 72 vii Figure 4.47: ROC curve of SVM in the balanced Paysim dataset using SMOTE .48: The comparison chart of the "accuracy" between different algorithms on the Paysim đafaS€(S. óc S11 121 11191 11H11 TT HT HH HH Hy 74 Figure 4.49: The chart compares the "Fl-score" of the algorithms on the Paysim bố ốc ố ố.
75 viii LIST OF TABLES roles Table 3.1: Commonly used Kernel functions .1: Credit Card Fraud Detection Dataset description .2: Number of columns and records from the dataset.3: Class distribution of the Credit Card Fraud Detection Dataset.4: Dataset check missing ValUC .5: The dataset after scaÏing.6: The Result when using KNN in original Credit Card dataset.7: The Result when using LR in original Credit Card dataset.8: The Result when using SVM in original Credit Card dataset.9: The result when running KNN with balanced Credit Card dataset using Random Undersaimpling.10: The result when running LR with balanced Credit Card dataset using Random Undersampling .11: The result when running SVM with balanced Credit Card dataset using Random Undersampling .12: The dataset classes before and after using SMOTE.13: The result when running KNN with balanced Credit Card dataset using b0 —.14: The result when running LR with balanced Credit Card dataset using b0.15: The result when running SVM with balanced Credit Card dataset using (00 .17: The comparison “F1-Score” between algorithms Credit Card dataset.18: Dataset DesCrip(iOn.19: Paysim Data Types. SH HH HH HH Hư, 50 Table 4.20: check missing value in Paysim đaf(aS€(. --- 6-55 S* sex 51 ix Table 4.21: Quantity statistics by transaction type .22: Class “isFraud” distribution of the Paysim Dataset .23: Statistics of the number of transactions on the isFlaggedFraud D100 ốc ốc cố Cố Cổ CÔ 53 Table 4.24: The Result when using K Nearest Neighbors in original Paysim (a(AS€(L.25: The Result when using LR in original Paysim dataset.26: The Result when using SVM in original Paysim dataset .27: The result when running KNN with balanced Paysim dataset using Random Undersampling .28: The result when running LR with balanced Paysim dataset using Random Ủndersaimpling.29: The result when running SVM with balanced Paysim dataset using Random Ủndersaimpling.30: Number of transactions in Paysim dataset before and after using SMOTE.34: The comparison “F1-score” between algorithms in Paysim dataset.75 LIST OF ABBREVIATIONS No. Word Acronym 1 Machine Learning ML 2 K Nearest Neighbor KNN 3 | Logistic Regression LR 4 | Support Vector Machine SVM 5 Techni Minority Oversampling SMOTE 6 True Positives TP 7 | False Positive FP 8 False Negative FN 9 True Negative TN xi ABSTRACT The financial industry has always dealt with fraud-related problems like missing and damaging in transactions.
In The United States, there are over 270,000 reports which makes credit card fraud become the most common type of identity thief. The number of frauds has increased doubling from 2017 to 2019 [1]. 2015 2016 2017 2018 2019 Figure 1: Identity theft reports in the United States ! + Source : https://www.com/the-ascent/research/identity-theft-credit-card-fraud-statistics xii Credit card fraud Other identity thet P2) Loan or lease fraud Phone or utilities fraud Bank fraud Employment or tax-related fraud Government documents or benefits fraud 4ã 52 50K 100K 150K 200K 250K Figure 2: Most common types of identity theft” Nowadays, there are more and more delicate techniques used by criminals for stealing the money from the user accounts. As a result, detecting fraudulent transactions is becoming more difficult because many illegal transactions look like the normal one.
In addition, the number of fraudulent transactions is higher than in the last few years. ? Source: https://www.com/the-ascent/research/identity-theft-credit-card-fraud-statistics xiii 250K 200K 150K = FSF ° 100K _ 50K « 2014 2015 2016 2017 2018 2019 @ Credit card fraud reports Figure 3: Credit card fraud reports by year? One of the most common challenges when facing fraudulent transactions is the skewed distribution of classes. The proportion of fraud classes is usually smaller many times than the unfraud classes. Though the classifiers should be inclined towards the minority group like fraud, they will focus on the majority group because of their regular appearances.
For dealing with this, in the research, we use Undersampling, Oversampling and One-class Classification techniques. Another problem incurs with an imbalance dataset is how to choose the performance measures used to evaluate models. From this research, we choose the Fl-score, confusion matrix, precision and recall to evaluate the accuracy. 3 Source: https://www.com/the-ascent/research/identity-theft-credit-card-fraud-statistics xiv Chapter 1 INTRODUCTION 1.
Problems For many decades, fraud in financial transactions has caused a lot of serious damages to the economy and the development of many business over the world. Therefore, enterprises always spend large of their resources in detecting fraud in transactions.