Phát Hiện Gian Lận Trong Giao Dịch Tài Chính Sử Dụng Phương Pháp Học Máy

Khóa luận tốt nghiệp nghiên cứu tốt nghiệp an approach for fraud detection in financial transactions using machine learning methods, vận dụng lý thuyết vào thực tế, đề xuất giải

Chuyên ngành

Information Systems

Người đăng

Ẩn danh

Thể loại

Thesis

2020

95
3
0

Phí lưu trữ

35 Point

Mục lục chi tiết

ACKNOWLEDGMENTS

1. CHƯƠNG 1: INTRODUCTION

1.1. Aims and Objectives

1.2. Languages, Tools and Libraries

2. CHƯƠNG 2: BACKGROUND

2.1. Fraudulent transaction definition

2.2. Fraud Detection Approach

2.3. Imbalanced dataset problems

2.4. Methods to solve the imbalanced problems

3. CHƯƠNG 3: MACHINE LEARNING FOR FRAUD DETECTION

3.1. Applying K Nearest Neighbor (KNN)

3.2. Advantages and disadvantages

3.3. Logistic Regression model

3.4. Apply Logistic Regression step-by-step

3.5. Support Vector Machine

3.6. Precision and Recall

3.7. Receiver Operating Characteristic Curve

3.8. Data Analytics Pipeline

3.9. Credit Card Dataset

4. CHƯƠNG 4

4.1. Synthetic Financial Datasets

4.2. Data Processing

5. CHƯƠNG 5: CONCLUSIONS AND FUTURE WORKS

LIST OF FIGURES

LIST OF TABLES

LIST OF ABBREVIATIONS

ABSTRACT

Tóm tắt

I. Tổng Quan Về Phát Hiện Gian Lận Trong Giao Dịch Tài Chính

Phát hiện gian lận trong giao dịch tài chính là một vấn đề nghiêm trọng trong ngành tài chính. Sự gia tăng của các giao dịch trực tuyến đã tạo ra nhiều cơ hội cho tội phạm tài chính. Việc áp dụng học máy trong phát hiện gian lận đã trở thành một xu hướng quan trọng. Các phương pháp học máy giúp phân tích dữ liệu lớn và phát hiện các mẫu gian lận một cách hiệu quả.

1.1. Định Nghĩa Gian Lận Tài Chính

Gian lận tài chính được định nghĩa là các giao dịch không được phép bởi chủ thẻ. Các giao dịch này có thể bao gồm thẻ bị mất, bị đánh cắp hoặc các giao dịch giả mạo. Việc hiểu rõ về gian lận tài chính là bước đầu tiên trong việc phát hiện và ngăn chặn.

1.2. Tầm Quan Trọng Của Phát Hiện Gian Lận

Phát hiện gian lận không chỉ bảo vệ tài sản của khách hàng mà còn bảo vệ danh tiếng của các tổ chức tài chính. Việc phát hiện sớm các giao dịch gian lận giúp giảm thiểu thiệt hại và tăng cường an ninh tài chính.

II. Thách Thức Trong Phát Hiện Gian Lận Tài Chính

Một trong những thách thức lớn nhất trong phát hiện gian lận là sự mất cân bằng trong dữ liệu. Các giao dịch gian lận thường chiếm tỷ lệ rất nhỏ so với các giao dịch hợp pháp. Điều này khiến cho các thuật toán học máy khó khăn trong việc nhận diện các mẫu gian lận.

2.1. Vấn Đề Dữ Liệu Mất Cân Bằng

Dữ liệu mất cân bằng xảy ra khi số lượng giao dịch gian lận rất ít so với giao dịch hợp pháp. Điều này dẫn đến việc các mô hình học máy thường thiên về dự đoán giao dịch hợp pháp, bỏ qua các giao dịch gian lận.

2.2. Khó Khăn Trong Việc Đánh Giá Hiệu Suất

Việc đánh giá hiệu suất của các mô hình phát hiện gian lận là một thách thức. Các chỉ số như độ chính xác có thể không phản ánh đúng khả năng phát hiện gian lận, do đó cần sử dụng các chỉ số khác như F1-scoređộ nhạy.

III. Phương Pháp Học Máy Trong Phát Hiện Gian Lận

Các phương pháp học máy đã được áp dụng để phát hiện gian lận trong giao dịch tài chính. Những phương pháp này bao gồm học sâu, học có giám sáthọc không giám sát. Mỗi phương pháp có những ưu điểm và nhược điểm riêng.

3.1. Học Có Giám Sát

Học có giám sát sử dụng dữ liệu đã được gán nhãn để huấn luyện mô hình. Các thuật toán như Hồi quy logisticSVM thường được sử dụng để phát hiện gian lận.

3.2. Học Không Giám Sát

Học không giám sát không yêu cầu dữ liệu gán nhãn. Các thuật toán như K-meansphân cụm giúp phát hiện các mẫu gian lận mà không cần biết trước về chúng.

IV. Ứng Dụng Thực Tiễn Của Học Máy Trong Phát Hiện Gian Lận

Việc áp dụng học máy trong phát hiện gian lận đã mang lại nhiều kết quả tích cực. Các tổ chức tài chính đã sử dụng các mô hình học máy để giảm thiểu rủi ro và tăng cường an ninh cho các giao dịch của họ.

4.1. Kết Quả Nghiên Cứu

Nghiên cứu cho thấy rằng việc sử dụng các mô hình học máy có thể cải thiện đáng kể khả năng phát hiện gian lận. Các mô hình như KNNSVM đã chứng minh hiệu quả trong việc phát hiện các giao dịch gian lận.

4.2. Ứng Dụng Trong Ngành Tài Chính

Nhiều ngân hàng và tổ chức tài chính đã áp dụng các giải pháp học máy để phát hiện gian lận. Việc này không chỉ giúp bảo vệ tài sản mà còn nâng cao lòng tin của khách hàng.

V. Kết Luận Và Tương Lai Của Phát Hiện Gian Lận

Phát hiện gian lận trong giao dịch tài chính là một lĩnh vực đang phát triển nhanh chóng. Với sự tiến bộ của công nghệ học máy, khả năng phát hiện gian lận sẽ ngày càng chính xác hơn. Tương lai của phát hiện gian lận sẽ phụ thuộc vào việc cải thiện các thuật toán và áp dụng các công nghệ mới.

5.1. Xu Hướng Tương Lai

Các xu hướng như Big DataAI sẽ tiếp tục định hình cách thức phát hiện gian lận. Việc tích hợp các công nghệ này sẽ giúp nâng cao hiệu quả phát hiện gian lận.

5.2. Tầm Quan Trọng Của Đổi Mới

Đổi mới trong công nghệ và phương pháp sẽ là chìa khóa để giải quyết các thách thức trong phát hiện gian lận. Các tổ chức cần đầu tư vào nghiên cứu và phát triển để duy trì lợi thế cạnh tranh.

10/07/2025
Khóa luận tốt nghiệp an approach for fraud detection in financial transactions using machine learning methods

Trích đoạn nội dung tài liệu

VIETNAM NATIONAL UNIVERSITY HOCHIMINH CITY UNIVERSITY OF INFORMATION TECHNOLOGY ADVANCED PROGRAM IN INFORMATION SYSTEMS NGUYEN THI MY LAN- 16520651 LE NGOC UYEN VY- 16521472 AN APPROACH FOR FRAUD DETECTION IN FINANCIAL TRANSACTIONS USING MACHINE LEARNING METHODS BACHELOR OF ENGINEERING IN INFORMATION SYSTEMS THESIS ADVISOR Dr. CAO THI NHAN HO CHI MINH CITY, 2020 ASSESSMENT COMMITTEE The Assessment Committee is established under the Decision. by Rector of the University of Information Technology. ACKNOWLEDGMENTS oElœ The thesis topic will not be completed without any assistance.

So the authors implementing the topic gratefully give acknowledgment to their support and motivation during the graduation thesis project. First of all, we would like to thank the Lecturers of the University of Information Technology as well as the Lecturers of the Information Systems Faculty, who taught and provided the background solid knowledge throughout the time studying at school. This knowledge is an important basis for us to complete our graduation thesis. In particular, we would like to express my endless thanks and gratefulness to my supervisor, Dr.

Cao Thi Nhan. Her dedicated support and constant advice have helped us to complete our thesis well. Her words of encouragement and comments have greatly enriched and improved our work. Without her guidance, the thesis would not have been done effectively.

During the implementation of the thesis, we have tried to apply effectively what we have learned as well as learn new technologies to be able to complete the thesis in the best way. However, in the process of implementation, because of limited knowledge, experience, and study time, it is difficult to avoid shortcomings. Therefore, we hope to receive the comments of teachers to complete the thesis better as well as the necessary knowledge and skills. Thank you so much! Authors Nguyen Thi My Lan — Le Ngoc Uyen Vy TABLE OF CONTENTS woes ACKNOWLEDGMENTS.

TABLE OF CONTENTS .cccscscsssssssssssssssssscsenencnssssssececscsenescnsessesacecsenenesseseeseceees ii LIST OF FIGURES.ccssssssssssssrsssessnsnsessssncsesenensnsssssscsesenencnssssssecacsenensnseseeseseees M LIST OF TA BLES. 5555 S522 Y9911110181010118810101010101001 08ix LIST OF ABBREVIATIONS .cssssssssssssssssssssenenensssesssrecsesenensnsssessececsenensneesenes xi ABSTRACT Chapter 1 INTRODUCTION. cect os saaece MO casas se seseevevesenssssusesnevevesenrnnaees 1 1. Aims and Obj€CfÏV€S.

Languages, Tools and LibrarieS. ¿5525252 S+s+*2+s£+x+xezsexexsx 2 Chapter 2 BACKGROUND. Fraudulent transaction definitiO. Fraud Detection Approach.

Imbalanced dataset problems. Methods to solve the imbalanced problems .-- + + «+ +++<s<+ 7 Chapter 3 MACHINE LEARNING FOR FRAUD DETECTION.- Ăn TT HH HH Hư 10 3. Applying K Nearest Neighbor (KNN). Advantages and disadvantages.

Logistic Regression model. Apply Logistic Regression step-by-step „l5 3. Advantages and disadvantages .---- cà + St sseireeeiey 15 3. Support Vector Machine.

Advantages and Disadvantages. Precision and Recall. Receiver Operating Characteristic CUTV€. Data Analytics Pipeline.

Credit Card Dataset.- 5c St S ĐH HH HH re. cece th HH HH HH HH HH 24 4. Làn 12 HH HH HH HH0 gi 25 4. Synthetic Financial DataS€(S.--- cành HH HH HH it 48 4.

óc 1à TT HT Hà nh Hệ 48 4. PT€DTOC€SSÏNE. án TH HH HH HH gi 56 4. LH TH HT TH TH HT HH TH TH HH Hy 58 4.

HH HH HH HH HH HH Hi 73 Chapter 5 CONCLUSIONS AND FUTURE WORKS. The results achieved. Ăn 01011000010101010000001010101000000000 186 1 iv LIST OF EIGURES voles Figure 1: Identity theft reports in the United States. -------+©-<++ xii Figure 2: Most common types of identity thet.

-- «655 S+c+xsserersreree xiii Figure 3: Credit card fraud reports by Y€AT.- --- +5 + 55252 Sc+x+xzesersrsrseree xiv Figure 3.1: Applying the K Nearest Neighbor (KNN) algorithm step-by-step [8] .2: Common distance metrics [9] .3: Graph of Logistic curve where o=0 and B=1 [1 I].4: Graph of Sigmod function [12] .6: Precision and Recall.7: Receiver Operating Characteristic curve model .1: Data analytics Pipeline .2: The original dataset .3: Dataset Class Distribution .4: Transactions amount distrIDutiOI.5: Transaction Time distributiON. eee cesses ee eeeseseneeeeeeseneees 28 Figure 4.6: Dataset after scaling features “Time” and “Amounf”.7: Confution matrix of testing data when applying KNN algorithm in the original Credit Card dataset .8: Confution matrix of testing data when applying LR algorithm in the original Credit Card đafAS€K. s6 5 111 1n TH TH ng HH nh Hư 32 Figure 4.9: ROC curve of LR in the imbalanced Credit Card dataset.10: Confution matrix of testing data when applying KNN algorithm in the original Credit Card dataset. ¿<5 1xx 91 121111510101 111211101010 1010 ty 34 Figure 4.11: ROC curve of SVM in the imbalanced Credit Card dataset.12: Class distribution in the subsample after using Random Undersampling —.13: Confusion matrix of testing data when applying KNN with undersampling in the Credit Card dataset .14: Confusion matrix of testing data when applying LR with undersampling in the Credit Card dataset .15: ROC curve of LR in the balanced Credit Card dataset using Random Undersammpling.16: Confusion matrix of testing data when applying SVM with undersampling in the Credit Card dataset .17: ROC curve of SVM in the balanced Credit Card dataset using Random ndersampling.18: Class distributions of the Credit Card dataset after applying SMOTE 41 Figure 4.19: Confusion matrix of testing data when applying KNN with SMOTE in the Credit Card dataset .20: Confusion matrix of testing data when applying LR with SMOTE in the Credit Card dataset .21: ROC curve of LR in the balanced Credit Card dataset using SMOTE 44 Figure 4.22: Confusion matrix of testing data when applying SVM with SMOTE in the Credit Card dataset .23: ROC of SVM in balanced Credit Card dataset using SMOTE.24: The comparison chart of the "accuracy" between different algorithms on the Credit Card đia(ASCTL.25: The chart compares the "Fl-score" of the algorithms on the Credit Card datasets .26: shows number of transactions which are the actual fraud per tramsaction tyPe.

St TT HT HH TH HH Hy 52 Figure 4.27: Distribution Of types .--- ¿+ kg ng it 52 Figure 4.28: Distribution of EFauid.-- ceeceeececseseeeneeseseseseeeeeeeeseneeeeeeaeae 53 vi Figure 4.29: Distribution of the feature FlaggedFraud .30: Original Paysim dataset .- cà 1n HH ng it 56 Figure 4.31: Data “isFraud” distributiOI.33: Confusion matrix of testing data when applying K Nearest Neighbors algorithm in the original Paysim dataset .34: Confution matrix of testing data when applying LR algorithm in the original Paysim đa(AS€(L.35: ROC curve of LR in the imbalaced Paysim dataset.36: Confution matrix of testing data when applying SVM algorithm in the Original Paysim điafAS€.37: ROC curve of SVM in the imbalanced Paysim dafaset.38: Class distribution in the subsample after using Random Undersampling Figure 4.39: Confusion matrix of testing data when applying KNN with undersampling in the Paysim dataset .40: Confusion matrix of testing data when applying LR with undersampling in the Paysim dataset .41: ROC curve of LR in the balanced Paysim dataset using Random ndersaimpling.42: Confusion matrix of testing data when applying SVM with undersampling in the Paysim dataset .43: ROC curve of SVM in the balanced Paysim dataset using Ramdom UnderSampÏinng. - - ‹- + 1k 1 1 1 1E TT TT HT TH TH TH Hư 69 Figure 4.44: Confusion matrix of testing data when applying LR with SMOTE in the Paysim dafASCL.45: ROC curve of LR in the balanced Paysim dataset using SMOTE.46: Confusion matrix of testing data when applying SVM with SMOTE in the Paysim dataset. - - 11k TT HH HH HH HH TH HH Hy 72 vii Figure 4.47: ROC curve of SVM in the balanced Paysim dataset using SMOTE .48: The comparison chart of the "accuracy" between different algorithms on the Paysim đafaS€(S. óc S11 121 11191 11H11 TT HT HH HH Hy 74 Figure 4.49: The chart compares the "Fl-score" of the algorithms on the Paysim bố ốc ố ố.

75 viii LIST OF TABLES roles Table 3.1: Commonly used Kernel functions .1: Credit Card Fraud Detection Dataset description .2: Number of columns and records from the dataset.3: Class distribution of the Credit Card Fraud Detection Dataset.4: Dataset check missing ValUC .5: The dataset after scaÏing.6: The Result when using KNN in original Credit Card dataset.7: The Result when using LR in original Credit Card dataset.8: The Result when using SVM in original Credit Card dataset.9: The result when running KNN with balanced Credit Card dataset using Random Undersaimpling.10: The result when running LR with balanced Credit Card dataset using Random Undersampling .11: The result when running SVM with balanced Credit Card dataset using Random Undersampling .12: The dataset classes before and after using SMOTE.13: The result when running KNN with balanced Credit Card dataset using b0 —.14: The result when running LR with balanced Credit Card dataset using b0.15: The result when running SVM with balanced Credit Card dataset using (00 .17: The comparison “F1-Score” between algorithms Credit Card dataset.18: Dataset DesCrip(iOn.19: Paysim Data Types. SH HH HH HH Hư, 50 Table 4.20: check missing value in Paysim đaf(aS€(. --- 6-55 S* sex 51 ix Table 4.21: Quantity statistics by transaction type .22: Class “isFraud” distribution of the Paysim Dataset .23: Statistics of the number of transactions on the isFlaggedFraud D100 ốc ốc cố Cố Cổ CÔ 53 Table 4.24: The Result when using K Nearest Neighbors in original Paysim (a(AS€(L.25: The Result when using LR in original Paysim dataset.26: The Result when using SVM in original Paysim dataset .27: The result when running KNN with balanced Paysim dataset using Random Undersampling .28: The result when running LR with balanced Paysim dataset using Random Ủndersaimpling.29: The result when running SVM with balanced Paysim dataset using Random Ủndersaimpling.30: Number of transactions in Paysim dataset before and after using SMOTE.34: The comparison “F1-score” between algorithms in Paysim dataset.75 LIST OF ABBREVIATIONS No. Word Acronym 1 Machine Learning ML 2 K Nearest Neighbor KNN 3 | Logistic Regression LR 4 | Support Vector Machine SVM 5 Techni Minority Oversampling SMOTE 6 True Positives TP 7 | False Positive FP 8 False Negative FN 9 True Negative TN xi ABSTRACT The financial industry has always dealt with fraud-related problems like missing and damaging in transactions.

In The United States, there are over 270,000 reports which makes credit card fraud become the most common type of identity thief. The number of frauds has increased doubling from 2017 to 2019 [1]. 2015 2016 2017 2018 2019 Figure 1: Identity theft reports in the United States ! + Source : https://www.com/the-ascent/research/identity-theft-credit-card-fraud-statistics xii Credit card fraud Other identity thet P2) Loan or lease fraud Phone or utilities fraud Bank fraud Employment or tax-related fraud Government documents or benefits fraud 4ã 52 50K 100K 150K 200K 250K Figure 2: Most common types of identity theft” Nowadays, there are more and more delicate techniques used by criminals for stealing the money from the user accounts. As a result, detecting fraudulent transactions is becoming more difficult because many illegal transactions look like the normal one.

In addition, the number of fraudulent transactions is higher than in the last few years. ? Source: https://www.com/the-ascent/research/identity-theft-credit-card-fraud-statistics xiii 250K 200K 150K = FSF ° 100K _ 50K « 2014 2015 2016 2017 2018 2019 @ Credit card fraud reports Figure 3: Credit card fraud reports by year? One of the most common challenges when facing fraudulent transactions is the skewed distribution of classes. The proportion of fraud classes is usually smaller many times than the unfraud classes. Though the classifiers should be inclined towards the minority group like fraud, they will focus on the majority group because of their regular appearances.

For dealing with this, in the research, we use Undersampling, Oversampling and One-class Classification techniques. Another problem incurs with an imbalance dataset is how to choose the performance measures used to evaluate models. From this research, we choose the Fl-score, confusion matrix, precision and recall to evaluate the accuracy. 3 Source: https://www.com/the-ascent/research/identity-theft-credit-card-fraud-statistics xiv Chapter 1 INTRODUCTION 1.

Problems For many decades, fraud in financial transactions has caused a lot of serious damages to the economy and the development of many business over the world. Therefore, enterprises always spend large of their resources in detecting fraud in transactions.

Nội dung được bảo vệ bản quyền — Tải xuống đầy đủ

Tài liệu "Phát Hiện Gian Lận Trong Giao Dịch Tài Chính Sử Dụng Phương Pháp Học Máy" cung cấp cái nhìn sâu sắc về cách mà các phương pháp học máy có thể được áp dụng để phát hiện và ngăn chặn gian lận trong lĩnh vực tài chính. Tài liệu này không chỉ giải thích các thuật toán học máy mà còn nêu bật những lợi ích mà chúng mang lại, như khả năng phân tích dữ liệu lớn và phát hiện các mẫu bất thường trong giao dịch. Độc giả sẽ tìm thấy thông tin hữu ích về cách cải thiện độ chính xác trong việc phát hiện gian lận, từ đó giúp các tổ chức tài chính bảo vệ tài sản và uy tín của mình.

Nếu bạn muốn mở rộng kiến thức về ứng dụng của học máy trong các lĩnh vực tài chính khác, hãy tham khảo các tài liệu như Ứng dụng các mô hình học máy machine learning trong dự báo giá cổ phiếu trên sàn chứng khoán hose, nơi bạn có thể tìm hiểu về dự báo giá cổ phiếu. Ngoài ra, tài liệu Ứng dụng mô hình học máy trong dự báo rủi ro phá sản của các doanh nghiệp bất động sản sẽ giúp bạn hiểu rõ hơn về việc sử dụng học máy để đánh giá rủi ro trong lĩnh vực bất động sản. Cuối cùng, tài liệu Nghiên cứu ứng dụng thuật toán học máy tăng cường cho bài toán chấm điểm tín dụng sẽ cung cấp thêm thông tin về cách học máy có thể cải thiện quy trình chấm điểm tín dụng. Những tài liệu này sẽ giúp bạn có cái nhìn toàn diện hơn về ứng dụng của học máy trong tài chính.