Luận Văn: Phát Hiện Gian Lận Thẻ Tín Dụng Sử Dụng Machine Learning và Deep Learning

Khóa luận trình bày phương pháp phát hiện gian lận thẻ tín dụng bằng thuật toán học máy và học sâu, mang lại hiệu quả cao trong bảo mật tài chính.

Người đăng

Ẩn danh

Thể loại

thesis

2023

80
4
0

Phí lưu trữ

30 Point

Mục lục chi tiết

ACKNOWLEDGMENTS

1. CHƯƠNG 1: REASON FOR CHOOSING THE TOPIC

1.1. Research scope of the topic

2. CHƯƠNG 2: THEORETICAL BACKGROUND AND RELATED WORKS

2.1. Fraud detection

2.2. Extreme learning method

2.3. Support vector machine (SVM)

2.4. Convolutional Neural Network (CNN)

3. CHƯƠNG 3: EXPERIMENTAL ENVIRONMENT

3.1. Applied machine learning & ensemble learning techniques

3.1.1. K-nearest neighbours (KNN)

3.1.2. Support vector machine (SVM)

3.1.3. Extreme learning method (ELM)

3.2. Applied deep learning techniques

3.2.1. Baseline

3.2.2. Convolutional neural network (CNN)

3.3. Performance-evaluation results

4. CHƯƠNG 4: USE MACHINE LEARNING ALGORITHMS

4.1. Convolutional neural network (CNN)

4.2. Compare results among algorithms

4.3. Limitations and development directions

4.4. Development directions

LIST OF FIGURES

LIST OF TABLES

ABSTRACT

Tóm tắt

I. Tổng Quan Về Phát Hiện Gian Lận Thẻ Tín Dụng Bằng Machine Learning

Phát hiện gian lận thẻ tín dụng là một thách thức lớn trong ngành tài chính. Với sự gia tăng của các giao dịch trực tuyến, gian lận thẻ tín dụng đã trở thành một vấn đề nghiêm trọng. Các tổ chức tài chính cần các phương pháp hiệu quả để phát hiện và ngăn chặn gian lận. Machine Learning (ML) và Deep Learning (DL) đã nổi lên như những công cụ mạnh mẽ trong việc phát hiện gian lận. Chúng có khả năng phân tích dữ liệu lớn và phát hiện các mẫu bất thường trong giao dịch.

1.1. Định Nghĩa Gian Lận Thẻ Tín Dụng

Gian lận thẻ tín dụng là hành vi sử dụng thông tin thẻ tín dụng bị đánh cắp để thực hiện giao dịch trái phép. Điều này có thể xảy ra qua việc sử dụng thẻ bị mất, bị đánh cắp hoặc thông tin thẻ giả mạo.

1.2. Tầm Quan Trọng Của Phát Hiện Gian Lận

Việc phát hiện gian lận thẻ tín dụng không chỉ bảo vệ người tiêu dùng mà còn giúp các tổ chức tài chính giảm thiểu thiệt hại tài chính. Hệ thống phát hiện gian lận hiệu quả có thể tăng cường lòng tin của khách hàng và bảo vệ danh tiếng của doanh nghiệp.

II. Thách Thức Trong Phát Hiện Gian Lận Thẻ Tín Dụng

Phát hiện gian lận thẻ tín dụng đối mặt với nhiều thách thức. Một trong những thách thức lớn nhất là sự biến đổi liên tục của các phương thức gian lận. Các kẻ gian lận luôn tìm kiếm các lỗ hổng trong hệ thống. Hơn nữa, việc phân tích dữ liệu lớn và không đồng nhất cũng gây khó khăn cho việc phát hiện chính xác.

2.1. Sự Biến Đổi Của Các Phương Thức Gian Lận

Các phương thức gian lận ngày càng tinh vi hơn, từ việc sử dụng thẻ giả đến các giao dịch trực tuyến không hợp lệ. Điều này yêu cầu các hệ thống phát hiện phải liên tục cập nhật và cải tiến.

2.2. Dữ Liệu Không Đồng Nhất

Dữ liệu giao dịch có thể đến từ nhiều nguồn khác nhau và có thể không đồng nhất. Việc xử lý và phân tích dữ liệu này để phát hiện gian lận là một thách thức lớn.

III. Phương Pháp Machine Learning Trong Phát Hiện Gian Lận

Machine Learning cung cấp nhiều phương pháp để phát hiện gian lận thẻ tín dụng. Các thuật toán như Học Tập Cạnh (SVM), Học Tập Gần Nhất (KNN) và Học Tập Tăng Cường (Boosting) đã được áp dụng thành công. Những phương pháp này giúp phân tích dữ liệu và phát hiện các mẫu gian lận một cách hiệu quả.

3.1. Học Tập Cạnh SVM

SVM là một trong những thuật toán phổ biến trong phát hiện gian lận. Nó giúp phân loại các giao dịch thành hợp lệ và gian lận bằng cách tìm kiếm một siêu phẳng tối ưu.

3.2. Học Tập Gần Nhất KNN

KNN là một phương pháp đơn giản nhưng hiệu quả trong việc phát hiện gian lận. Nó dựa trên việc so sánh các giao dịch mới với các giao dịch đã biết để xác định tính hợp lệ.

IV. Ứng Dụng Deep Learning Trong Phát Hiện Gian Lận

Deep Learning, đặc biệt là Mạng Nơ-ron Tích Chập (CNN), đã chứng minh được hiệu quả trong việc phát hiện gian lận thẻ tín dụng. Các mô hình này có khả năng học hỏi từ dữ liệu lớn và phát hiện các mẫu phức tạp mà các phương pháp truyền thống không thể làm được.

4.1. Mạng Nơ ron Tích Chập CNN

CNN là một trong những mô hình Deep Learning mạnh mẽ nhất. Nó có khả năng xử lý và phân tích hình ảnh, nhưng cũng có thể áp dụng cho dữ liệu giao dịch để phát hiện gian lận.

4.2. So Sánh Hiệu Suất Giữa ML và DL

Nghiên cứu cho thấy rằng các mô hình Deep Learning thường đạt được độ chính xác cao hơn so với các mô hình Machine Learning truyền thống trong việc phát hiện gian lận thẻ tín dụng.

V. Kết Quả Nghiên Cứu Về Phát Hiện Gian Lận

Nghiên cứu đã chỉ ra rằng việc áp dụng Machine Learning và Deep Learning có thể cải thiện đáng kể khả năng phát hiện gian lận thẻ tín dụng. Các mô hình hiện đại như CNN và XGBoost đã đạt được độ chính xác lên đến 94%. Điều này cho thấy tiềm năng lớn của các công nghệ này trong việc bảo vệ người tiêu dùng và tổ chức tài chính.

5.1. Đánh Giá Hiệu Suất Các Mô Hình

Các mô hình được đánh giá dựa trên các chỉ số như độ chính xác, độ nhạy và độ đặc hiệu. Kết quả cho thấy rằng các mô hình Deep Learning vượt trội hơn trong việc phát hiện gian lận.

5.2. Ứng Dụng Thực Tiễn

Các tổ chức tài chính đã bắt đầu áp dụng các mô hình Machine Learning và Deep Learning để phát hiện gian lận, giúp giảm thiểu thiệt hại và tăng cường an ninh cho khách hàng.

VI. Tương Lai Của Phát Hiện Gian Lận Thẻ Tín Dụng

Tương lai của phát hiện gian lận thẻ tín dụng sẽ tiếp tục phát triển với sự tiến bộ của công nghệ. Các mô hình Machine Learning và Deep Learning sẽ ngày càng trở nên tinh vi hơn, giúp phát hiện gian lận một cách nhanh chóng và chính xác hơn. Sự kết hợp giữa các công nghệ mới và dữ liệu lớn sẽ mở ra nhiều cơ hội mới trong lĩnh vực này.

6.1. Xu Hướng Công Nghệ Mới

Các công nghệ như trí tuệ nhân tạo và blockchain có thể được tích hợp vào hệ thống phát hiện gian lận, tạo ra những giải pháp an toàn hơn cho người tiêu dùng.

6.2. Nghiên Cứu Tương Lai

Nghiên cứu trong lĩnh vực này sẽ tiếp tục tập trung vào việc cải thiện độ chính xác và hiệu quả của các mô hình phát hiện gian lận, đồng thời tìm kiếm các phương pháp mới để đối phó với các hình thức gian lận ngày càng tinh vi.

10/07/2025
Khóa luận tốt nghiệp credit card fraud detection using machine learning and deep learning algorithms

Trích đoạn nội dung tài liệu

VIETNAM NATIONAL UNIVERSITY HO CHI MINH CITY UNIVERSITY OF INFORMATION TECHNOLOGY ADVANCED PROGRAM IN INFORMATION SYSTEMS TRUONG MINH KHIET - 19520628 LUONG TIEN THUAN HAI - 19521462 THESIS CREDIT CARD FRAUD DETECTION USING MACHINE LEARNING AND DEEP LEARNING ALGORITHMS BACHELOR OF ENGINEERING IN INFORMATION SYSTEMS THESIS ADVISOR DR. CAO THI NHAN HO CHI MINH CITY, 2023 ACKNOWLEDGMENTS First, allow my group to express our deep gratitude to Ms. Cao Thi Nhan for her dedication and dedication to guiding the group throughout the process of researching and conducting their graduation thesis. Her patience, extensive knowledge and dedication helped the group overcome difficulties in the research process and complete the project on time as prescribed by the school.

Next, our group would also like to sincerely thank the teachers at the University of Information Technology - National University of Ho Chi Minh City for teaching us valuable knowledge necessary for our group to be able to Complete your thesis, as well as serve your future work well. My group also wants to thank the University of Information Technology - Ho Chi Minh City National University for creating learning and practicing conditions for our group to complete their thesis and course. Finally, our group would like to send our sincere thanks to our families and those who have helped, supported, and encouraged us to feel secure in researching and completing our thesis. TABLE OF CONTENTS cw ACKNOWLEDGMENTTS.- HH HH TH HH Hà HH TT nh ng 1 TABLE OF CONTIENTTS.- 1n HT TH HH HH Tnhh nh nh 2 LIST OF FIGURES.- G1 nHnH HH TT TH HH Hà HH TT ng 5 I0R09)30.- 6 33 v99 HH HH ng ưkp 10 1.

Reason for choosing the {OpIC.-- 5 5 2+ 119119 1 9 1 vn ng tr 10 1.- d1 1E TT TH HH nh 13 I0)) ung i00 0n. Research scope Of the fODIC. - -- 6 SE 1911 11 1 vn HH ng ngư 16 CHAPTER 2. THEORETICAL BACKGROUND AND RELATED WORKS.

Fraud et€CfIOH.- 5 5 x11 TH TH HH nh Hệ 19 2.-- --- 5 << 1x vn HH Hệ 21 2. Extreme learning methO(.-- s6 + +31 91119119 1 911 9v ng ng nề, 21 "2899 soi. LH HH HH Hệ, 22 2.- 5 s2 HH HH gi 23 2. Support vector machine (SVM).-- - HH ng ng kg 24 2.-- 5 G11 E11 vn ng Hết 25 "ĐC: 909.

Convolutional Neural Network (CNN).- c2 110101110 111 ng 0 vn reg 31 I1 0117 --. Experimental environment 000rraẳaẳ.- 3x HH TH HH ng grư 33 c3 0i. Applied machine learning & ensemble learning techn1ques. K-nearest neighbours CKRNN).

Support vector machine (SVM) ou. eee ee eee 5 1kg, 56 E6 bu. Extreme learning method (ELM). Applied deep leaning techniques.

Baseline ion ae. Convolutional neural network (CNN). Ăn vn HH nghiệt 62 3. Performance-evaluation €aSUT€S.

(<< << << 11111EE SĐT ng 055 ket 64 3. si HH ng ớt 64 4. HH TH HH HH HH Hết 64 3. Use machine learning aÏØOTIthH§.

Convolutional neural network (CNNN). Compare results among alGorithms. G1911 110 1h nh HH nh trờ 74 4. Limitations and development diTeCfIONS.- 6 SH TH HH HH HH HH75 4.

Development (IT€CfIOTNS .-- --- 1 11 2kg TH HH giết 77 LIST OF FIGURES Le Figure 1 - Illustration of ÍTaUd.- - << + E1 ng HH nh 19 Figure 2 - ELM's algOrithim.- xxx x19 TH ng HH ng Hưng gà 22 Figure 3 - RFP's aÏEOTItH1.-- 111 1 93 93 2 1 TH HH HH Hưng ngư 24 Figure 4 - SVM's alEOTIthim.- 6 + x2 TH TH TH Hàng HH tiệt 24 Figure 5 - Logistic regression's aÏØOTIfH.- --- 6 + s11 k*91 9 9v vn re, 25 Figure 6 - Logistic regression's aÏØOTIthim.-- ----s «+ + + +*kxeeerseeeereerereere 26 Figure 7 - Pooling ÍAW€T.-- c1 95993111 11h TT HH Hư 28 Figure 8 - CNN output ÏAY€T. G1 TH TH 30 Figure 9 - Application of dropout over neural netWOTK.- -- « +««es<+sx+se+sx+ 3l Figure 10 - Connect data through drive. --- 5 5 + + x** EEEseEEeereereereereerske 39 Figure 11 - Incorrect data handÏIng.- - c6 6+1 1231191119113 1 9v vn tre40 Figure 12 - Check the columns in data .- -- - <1 * 1£ +3 E+*EE+eeereeeeeeeersreere 41 Figure 13 - Chart of Class Distributions. -- --- < + S3 + kseierrrereere 42 Figure Si 0i.

43 Figure 15 - Allocating data for training and f€StInE.- s6 +- + +ss£secssesseesee 44 Figure 16 - Result of Label Distribution. cece ese eseeseeseeseeeeeseseeseeeeeseseeseeaeens 45 Figure 17 S311iiï1 6ï 0. 47 Figure 19 - Imbalance Correlation Matrix .- - s5 c1 + seerseeeeeerereere 48 Figure 20 - Balance Correlation ÏMafT1X.- -- << 1x19 1 ng ng 48 Figure 21 - Negative COrr€ÏfÏOTI.- -- 6 53k TH TH TH TH TH TH HH ng già 49 Figure 22 - Positive COTT€ÏAfIOT. 6 333 vn TH TH TH TH TH HH Hiệp 49 Figure 23 - Result of reduce OU{ÏI€T-.- - --- + + + + 3x E*kESeeEseeereeeeereeeereere 50 Figure 24 - Result of clustering methO(s.-- --- + + + + seeerseeereerereere 51 Figure 25 - Score model Decision Tree .cceesceesseeeseeesneeeseeceseeeeaeeesaeeeseeeeeeesaes 53 Figure 26 - Score model KNN.

LH HH TH HH HH 54 Figure 27 - Score model Ï.-- «xxx 3 3 9191 HH ng HH rưệt 55 Figure 28 - Score model SM. kg TH TH HH HH ng 56 Figure 29 - Score model LOGISTIC REGRESSION. c2 S2sesee 58 Figure 30 - Score model XG BOOS.-- -- c1 1111 1v HH ng re 59 Figure 31 — Score of ELM Model .-- -- << 1139113 1S kg re 61 Figure 32 - Cross-Entropy “s aÏØOTIfhm.- -- - + 1311131 E+EESeeEseeeerrersreere 63 Figure 33 - Formula to calculate aCCUTACY. - c c1 11131113111 1 ve rre 63 Figure 34 - Formula to calculate DF€CISIOTI.- -- <5 +5 *+2+£+*E+veeEeeeeeeereeeers 64 Figure 35 - Formula to calculate Fl- SCOTG.

5 + + + * + E+*vEEseeeeeersreers 64 Figure 36 - Formula to calculate r€CaÏÏÏ.-- -- -- ¿+ +++ + *++‡++*kE+veeEseereeeeeeeerereere 64 Figure 37 - Score the MOel.- -- -- <1 1311131111911 9111 vn HH ng 66 Figure 38 - Score Model .- ---- <5 1 1 HT TH HH Hư 67 LIST OE TABLES Le Table 1.1 - Configure the execution COIDUẨ€T .2 - Configure the Python 3 Google Compute Engine.1 - Accuracy based results of deep learning algorithms.2 - The list of features available in the CCF dafaset.3 - Characteristics of the dataset. 1v ng ng, 36 Table 3.4 - Characteristics of the dataset .- 5 6s vn ngư 37 Table 3.5 - Characteristics of the đafaS€(.- 2G G1 111v SH re 52 Table 3.6 - Table score of Decision “ÏTT€€.7 - Table score Of KNN. LH HH HH HH HH, 54 Table 3.8 - Table score Of RE. c5 11k TH TH HH HH 55 Table 3.9 - Table score Of SVM.

LH HH HH TH TH TH HH HH, 57 Table 3.10 - Table score of Decision “TTT€.- --- 6 sa s11 139113511931 1 31x krse 58 Table 3.11 - Table score of XG BooSt .- 0 n1 ng ng ngư 60 Table 3.12 - Table score Of ELUÌM.-- 5G s11 E91 211191 9111 11t HH ng tr 61 Table 3.13 - Table score of CINN.-- HH HH TH HH HH, 67 Table 3.14 - Table score of BaseÏine. ng ng ng ry 68 Table 3.16 - Results of the algorithms when running on the initial data.17 - Results of algorithms when running on only positive correlations .18 - Results of algorithms when running on only negative correlations .19 - Results of algorithms when running on negative correlations and positive COTT€ÏAfÏOTN§. -- 2 222211101110101033111 11T ng g0 0 HH 70 Table 3.20 - The results of the algorithms when running on the processed data set 71 Table 3.21 - Comparison between 3 models witth different data sefs. 72 ABSTRACT CCF represents an escalating threat to financial institutions, as fraudsters continually develop novel methods to exploit vulnerabilities.

The need for a robust classifier is paramount to adjust to the ever-changing landscape of fraud. The primary goal of a fraud detection system is to precisely forecast instances of fraud while minimizing the occurrence of false positives. The effectiveness of machine learning (ML) approaches varies across diverse business scenarios, where the characteristics of input data play a crucial role in determining the suitable ML techniques. In the realm of CCF detection, pivotal factors influencing model performance encompass the quantity of features, transaction volume, and inter-feature correlations.

Deep learning (DL) techniques, exemplified by Convolutional Neural Networks (CNNs) and their layers, are essential for text processing and serve as a foundational model. Applying these techniques to identify fraudulent credit card activities surpasses the capabilities of traditional algorithms. A comparative evaluation of algorithmic performances highlights the CNN with 20 layers and the XG Boost model as the leading methods, achieving an impressive accuracy of 94. Numerous sampling methods are employed to boost the performance of existing examples, yet they exhibit a noticeable decline in efficacy when applied to unseen data.

Intriguingly, there is an observed enhancement in performance on unseen data with an increased class imbalance. Future research initiatives might explore the integration of more sophisticated deep learning methods to elevate the overall performance of the model proposed in this study. Reason for choosing the topic Credit Card Fraud (CCF) is a form of identity theft where unauthorized individuals use stolen or fake credit card details to conduct transactions. The use of lost, stolen, or forged credit cards can lead to fraudulent activities.

Fraud without physical card possession or the use of credit card information in online transactions is becoming more common due to the growing popularity of online shopping. The increase in fraud, such as CCF, is a result of the expansion of electronic banking and various online payment platforms, resulting in billions of dollars in annual damages. Detecting CCF has become a top priority in this era of digital payments. As a business owner, it is clear that the future is moving towards a cashless society.

As a result, traditional payment methods may not be effective for business expansion, as customers may not always carry cash. Debit and credit card payments are currently subject to fees by companies. As a result, businesses must update their infrastructure to accept all forms of payments. This situation is expected to become even more critical in the coming years [1].

In 2020, approximately 393,207 incidents of Credit Card Fraud (CCF) occurred, accounting for a significant portion of the reported 1.4 million identity theft cases [4]. Presently, CCF ranks as the second most prevalent type of identified identity theft, following closely behind government document and benefits fraud [5]. More specifically, 365,597 instances of fraud in 2020 were associated with the unauthorized opening of new credit card accounts [10]. Comparing the years 2019 and 2020 reveals a substantial 113% increase in identity theft complaints, with reports of credit card identity theft witnessing a notable surge of 44.

On a global scale, payment card theft resulted in an economic loss of $24.26 billion the preceding year. It is noteworthy that the United States, accounting for 38.6% of reported card fraud losses in 2018, emerges as the most susceptible country to credit card theft. Page | 10 Hence, it is essential for financial institutions to prioritize the implementation of an automated system for detecting fraud. The objective of detecting Credit Card Fraud (CCF) entails creating a Machine Learning (ML) model that leverages historical data from credit card payment transactions.

This model aims to distinguish between transactions that are fraudulent and those that are legitimate, using this insight to assess the authenticity of incoming transactions. Overcoming associated challenges requires tackling fundamental issues such as system responsiveness, cost considerations, and feature preprocessing.

Nội dung được bảo vệ bản quyền — Tải xuống đầy đủ