Khóa luận tốt nghiệp hệ thống thông tin realtime customer reviews analysis using machine learning on big data framework

Khóa luận tốt nghiệp phân tích đánh giá khách hàng thời gian thực bằng học máy trên nền tảng dữ liệu lớn, ứng dụng công nghệ hiện đại.

Chuyên ngành

Information Systems

Người đăng

Ẩn danh

Thể loại

thesis

2021

72
10
0

Phí lưu trữ

30 Point

Mục lục chi tiết

ACKNOWLEDGMENTS

TABLE OF CONTENTS

1. CHƯƠNG 1: PROBLEM STATEMENT

1.1. Rationale

1.2. Aims

1.3. Object and range of study

2. CHƯƠNG 2: RELATED WORK

3. CHƯƠNG 3: METHODOLOGY

3.1. Core technologies of spark and its components

3.2. Machine learning algorithms

3.3. Sentiment Analysis of Twitter Data

4. CHƯƠNG 4: EXPERIMENTAL RESULTS

4.1. Prepare data for the real-time problem

4.2. Dataset for train-test

4.2.1. TF-IDF vectorizer + Logistic Regression Classifier

4.2.2. TF-IDF vectorizer + NaiveBayes Classifier

4.2.3. TF-IDF vectorizer + Decision Tree Classifier

4.2.4. TF-IDF vectorizer + Random Forest Classifier

4.3. Compare Model

4.4. Real-time sentiment analyst

4.5. Compare with UIT-VSMEC

5. CHƯƠNG 5: SUMMARY

6. CHƯƠNG 6: FUTURE WORK

REFERENCES

Tóm tắt

I. Phân Tích Khách Hàng

Phân tích khách hàng là quá trình thu thập và xử lý dữ liệu để hiểu rõ hành vi, nhu cầu và phản hồi của khách hàng. Trong nghiên cứu này, Big DataMachine Learning được sử dụng để phân tích dữ liệu từ mạng xã hội, đặc biệt là Twitter. Các tweet được thu thập và phân loại theo cảm xúc như buồn, vui, giận dữ, sợ hãi, ghê tởm và ngạc nhiên. Phương pháp này giúp doanh nghiệp theo dõi phản hồi của khách hàng một cách hiệu quả và nhanh chóng.

1.1. Thu thập dữ liệu

Dữ liệu được thu thập từ Twitter thông qua công cụ crawler, sau đó được gán nhãn cảm xúc. Bộ dữ liệu này kết hợp với bộ dữ liệu UIT-VSMEC để tăng độ chính xác trong phân tích. Big Data đóng vai trò quan trọng trong việc xử lý khối lượng dữ liệu lớn, đảm bảo tính realtime trong phân tích.

1.2. Phân loại cảm xúc

Các thuật toán Machine Learning như Logistic Regression, Naive Bayes, Decision TreeRandom Forest được sử dụng để phân loại cảm xúc. Kết quả cho thấy Logistic Regression có độ chính xác cao nhất. Phương pháp này giúp doanh nghiệp hiểu rõ phản ứng của khách hàng đối với sản phẩm và dịch vụ.

II. Đánh Giá Khách Hàng

Đánh giá khách hàng là quá trình phân tích và đưa ra nhận định về hành vi và phản hồi của khách hàng. Nghiên cứu này sử dụng Machine Learning để đánh giá cảm xúc của khách hàng thông qua các tweet. Kết quả phân tích được cập nhật liên tục, giúp doanh nghiệp đưa ra quyết định kịp thời.

2.1. Phân tích cảm xúc

Hệ thống phân tích cảm xúc được xây dựng trên nền tảng Apache Spark, cho phép xử lý hàng triệu tweet trong thời gian thực. Kết quả phân tích bao gồm số lượng tweet, tỷ lệ cảm xúc và thời gian phản hồi. Điều này giúp doanh nghiệp nắm bắt xu hướng và phản ứng của khách hàng một cách nhanh chóng.

2.2. Ứng dụng thực tế

Phương pháp này có thể áp dụng trong nhiều lĩnh vực như quản lý thương hiệu, khảo sát sản phẩmphân tích thị trường. Kết quả phân tích giúp doanh nghiệp cải thiện sản phẩm và dịch vụ, tăng cường sự hài lòng của khách hàng.

III. Realtime và Big Data

RealtimeBig Data là hai yếu tố quan trọng trong nghiên cứu này. Hệ thống được xây dựng trên nền tảng Apache Spark, cho phép xử lý dữ liệu lớn và cập nhật kết quả liên tục. Điều này giúp doanh nghiệp theo dõi phản hồi của khách hàng một cách nhanh chóng và chính xác.

3.1. Xử lý dữ liệu lớn

Big Data đóng vai trò quan trọng trong việc xử lý khối lượng dữ liệu lớn từ mạng xã hội. Apache Spark được sử dụng để phân tích dữ liệu một cách hiệu quả, đảm bảo tính realtime trong cập nhật kết quả.

3.2. Cập nhật kết quả liên tục

Hệ thống cập nhật kết quả phân tích liên tục, giúp doanh nghiệp nắm bắt xu hướng và phản ứng của khách hàng một cách nhanh chóng. Điều này đặc biệt hữu ích trong các chiến dịch marketing và khảo sát sản phẩm.

IV. Machine Learning và Phân Tích Dữ Liệu

Machine LearningPhân tích dữ liệu là hai công nghệ chính được sử dụng trong nghiên cứu này. Các thuật toán Machine Learning được áp dụng để phân loại cảm xúc từ các tweet, trong khi Phân tích dữ liệu giúp hiểu rõ hành vi và phản hồi của khách hàng.

4.1. Thuật toán Machine Learning

Các thuật toán Logistic Regression, Naive Bayes, Decision TreeRandom Forest được sử dụng để phân loại cảm xúc. Kết quả cho thấy Logistic Regression có độ chính xác cao nhất, phù hợp với phân tích dữ liệu lớn.

4.2. Phân tích hành vi khách hàng

Phân tích dữ liệu giúp hiểu rõ hành vi và phản hồi của khách hàng. Kết quả phân tích được sử dụng để cải thiện sản phẩm và dịch vụ, tăng cường sự hài lòng của khách hàng.

V. Tối Ưu Hóa Khách Hàng và Dự Đoán Hành Vi

Tối ưu hóa khách hàngDự đoán hành vi là hai mục tiêu chính của nghiên cứu này. Phân tích dữ liệu và Machine Learning giúp dự đoán hành vi khách hàng, từ đó tối ưu hóa chiến lược kinh doanh.

5.1. Dự đoán hành vi

Các thuật toán Machine Learning được sử dụng để dự đoán hành vi khách hàng dựa trên phản hồi từ các tweet. Kết quả dự đoán giúp doanh nghiệp đưa ra chiến lược kinh doanh phù hợp.

5.2. Tối ưu hóa chiến lược

Phân tích dữ liệu giúp tối ưu hóa chiến lược kinh doanh, cải thiện sản phẩm và dịch vụ. Điều này giúp doanh nghiệp tăng cường sự hài lòng của khách hàng và nâng cao hiệu quả kinh doanh.

21/02/2025

Trích đoạn nội dung tài liệu

VIETNAM NATIONAL UNIVERSITY HOCHIMINH CITY UNIVERSITY OF INFORMATION TECHNOLOGY ADVANCED PROGRAM IN INFORMATION SYSTEMS NGUYEN HOANG NHAT REAL-TIME CUSTOMER REVIEWS ANALYSIS USING MACHINE LEARNING ON BIG DATA FRAMEWORK BACHELOR OF ENGINEERING IN INFORMATION SYSTEMS HO CHI MINH CITY, 2021 NATIONAL UNIVERSITY HOCHIMINH CITY UNIVERSITY OF INFORMATION TECHNOLOGY ADVANCED PROGRAM IN INFORMATION SYSTEMS NGUYEN HOANG NHAT - 17520851 REAL-TIME CUSTOMER REVIEWS ANALYSIS USING MACHINE LEARNING ON BIG DATA FRAMEWORK BACHELOR OF ENGINEERING IN INFORMATION SYSTEMS THESIS ADVISOR HOP DO TRONG, Ph.D HO CHI MINH CITY, 2021 ASSESSMENT COMMITTEE The Assessment Committee is established under the Decision. Ho Chi Minh city, date January 18" by Rector of the University of Information Technology. 1 Tho Quan Thanh, ph. 2 Thanh Ngo Duc, Ph.

3 Nhan Cao Thi, Ph.D - Commissary ACKNOWLEDGMENTS First of all, I thank God for giving the support and providing me with the patience and guidance which I used in making this work. I would like to thank my supervisor, Hop Do Trong, Ph. I would like to thank him for his help, continuous encouragement, productive discussion, and valuable suggestions and comments throughout the research and thesis work. It has been a fantastic experience studying at the University of Information Technology.

I was motivated and excited about studying, ultimately fulfilling my expectations. Last but not least important, I owe more than thanks to my family members especially. Thanks to my brother for his support and encouragement throughout my life. TABLE OF CONTENTS cae ACKNOWLEDGMENT .scscsssssssesesstssssesssstsnesesesansnesesesaeanesesesaeaesnesesanatenesesenaeaeaseneranaes i LIST OF FIGURES LIST 9).

Chapter 1 Problem Statemeni. --- «se «HH HH ghi 1 1. Object and range of stud: Chapter 2 Related Work.csssssssesessssssssssnseescsesessssssssneessesssoseesssneeeaneneaeeeseesseneneneaeeeeees 3 Chapter 3 Methodolog). HH HẤ HH nghe7 3.1 Core technologies of spark and its components.

Big Data Analysis Using Spark.1 Machine learning proce: 3.2 Machine learning algorithms 3.2 Sentiment Analysis of Twitter Data 3. Chapter 4 Experimental ReSUItS.1 Prepare data for the real-time problem 4.2 Dataset for train-test.1 TF-IDF vectorizer + Logi 4.2 TF-IDF vectorizer + NaiveBayes Classifi 4. TF-IDF vectorizer + Decision Tree Classifier 4.4 TF-IDF vectorizer + Random Forest Classifier. AA Compare MOdel .5 Real-time sentiment analyst.6 Compare with UIT-VSMEC.c tình nhe 56 Chapter 5 Summary.

Chapter 6 Future WOIFK .- ---- «<< << HH nh ghi 58 REEERENCES.-- 5-5 5< SH HH Họ TH TH HH HH HA HH BE 0.}gg000/(02Ề0sS6ố2n111111111 61 iii LIST OE FIGURES œELlb Figure 3.The three Vs of big đata.2 The ecosystem of Spark.3 Apache Spark data processing .4 A typical process of machine learning [17] .5 Summary of machine learning algorithms [18] .6 Sentiment analyst step.7 Attributes available in snscrape tweet object.8 Code for crawl Vietnamese EW€€ts.9 Dataframe of tW€ẨS.- - nh HH HH HH Hà HH HH HH HH HH HH key 17 Figure 3.10 Pipeline for spark nÍp.- - -- «5+ ềE* St E111 Tàn Tàn Tàn Tàn Tàn TH Tàn Hàn nhe 24 Figure 4.1 Format of data test in r€aÏ-tÏIT.2 Separate the data set Figure 4.3 Cleaning data Figure 4.4 Libs for training model Figure 4.5 Dictionary of emotion Figure 4.6 TF-IDF vectorizer + Logistic Regression Classifier pipeline.7 TF-IDF vectorizer + NaiveBayes Classifier pipeline .8 TF-IDF vectorizer + Decision Tree Classifier pipeline.9 TF-IDF vectorizer + Random Forest Classifier pipeline.10 Step for real-time sentiment.11 Result realtime predict sentiment of tweets .12 Count of Emotion in realtime .13 Visualizing number of emotion in real-time Figure 4.14 Calculate Average of latency in real-time.15 Percentage each EmOtIO.- - ¿c5 St ÉEk‡ÉE kg TH HH ưêc 55 Figure 4.16 Visualizing of Percentage of each Emotion .17 Count sentence is already annaÏySí.--- + c5 sét ey 55 iv LIST OE TABLES œELlb Table 3.1 Step for data preDFOC€SSINE.2 : Statistics of emotion labels of the UIT-VSMEC corpus [Ï].3 Statistics of emotion labels of the Our dataset .4 Statistics of emotion labels of the mix dataset .1 TF-IDF vectorizer + Logistic Regression Classifier Table 4.2 TF-IDF vectorizer + Logistic Regression Classifier with my data train + UIT data test.3 TF-IDF vectorizer + Logistic Regression Classifier with my data train + mix data test.4 Compare accuracy using TF-IDF vectorizer + Logistic Regression Classifier with my data using for training .5 TF-IDF vectorizer + Logis data t€SE.6 TF-IDF vectorizer + Logistic Regression Classifier with UIT data train + UIT data tesf.7 TF-IDF vectorizer + Logistic Regression Classifier with UIT data train + mix data test .8 Compare accuracy using TF-IDF vectorizer + Logistic Regression Classifier with UIT data using for training.9 TF-IDF vectorizer + Logistic Regression Classifier with mix data train + my data test.10 TF-IDF vectorizer + Logistic Regression Classifier with mix data train + UIT data test Table 4.11 TF-IDF vectorizer + Logistic Regression Classifier with mix data train + mix data (€SỂ.12 Compare accuracy using TF-IDF vectorizer + Logistic Regression Classifier with MIX data using for training Table 4.13 TF-IDF vectorizer + NaiveBayes Classifier with my data train + my data test Table 4.14 TF-IDF vectorizer + NaiveBayes Classifier with my data train + UIT data ` —.15 TF-IDF vectorizer + NaiveBayes Classifier with my data train + mix data Ôn .ốốốốỐốỐố ốốốốố ỐC CC Cố CC Cố 0000 36 Table 4.16 TF-IDF vectorizer + NaiveBayes Classifier with UIT data train + my data LOSE eee .17 TF-IDF vectorizer + NaiveBayes Classifier with UIT data train + UIT data ` —.21 TF-IDF vectorizer + NaiveBayes Classifier mix data train + MIX data test .22 Compare accuracy using TF-IDF vectorizer + NaiveBayes Classifier 40 41 Table 4.24 TF-IDF vectorizer + Decision Tree Classifier with my data train + UIT data `.25 TF-IDF vectorizer + Decision Tree Classifier with my data train + mix data {€S{.26 TF-IDF vectorizer + Decision Tree Classifier with UIT data train + my data W€S.27 TF-IDF vectorizer + Decision Tree Classifier with UIT data train + UIT data test.28 TF-IDF vectorizer + Decision Tree Classifier with UIT data train + mix data test.30 TF-IDF vectorizer + Decision Tree Classifier with mix data train + UIT data tet.31 TF-IDF vectorizer + Decision Tree Classifier with mix data train + mix ata t€SỂ.32 Compare accuracy using TF-IDF vectorizer + Decision Tree Classifier.35 TF-IDF vectorizer + Random Forest Classifier with my data train + mix lon T1 .36 TF-IDF vectorizer + Random Forest Classifier with UIT data train + my 6n ae.37 TF-IDF vectorizer + Random Forest Classifier with UIT data train + UIT data test.38 TF-IDF vectorizer + Random Forest Cl data test.39 TF-IDF vectorizer + Random Forest Classifier with mix data train + my data test.40 TF-IDF vectorizer + Random Forest Classifier with mix data train + UIT data test.41 TF-IDF vectorizer + Random Forest Classifier with mix data train + mix data E€SỂ. co nh re Table 4.42 Compare accuracy using TF-IDF vectorizer + Random Forest Classifier.43 Compare each model .44 Compare with UIT-VSMEC.--- tt ng TH Hư 56 vii ABSTRACT For firms to monitor their brand reputations and assess their performance and public feelings about their goods, have many ways, but nowadays, using automation to sentiment analyst of customer reviews is the easiest. We reveal our efforts to develop a machine learning-based system for sentiment analysis of tweets on Twitter and make it a real-time analyst in this study. We use the framework ‘Apache Spark' — Bigdata framework.

Our system outperforms the Emotion Recognition for Vietnamese Social Media Text, 2019 (UIT-VSMEC) [1]. We make a dataset using a tweets crawler and label the dataset that best reflects its sentiment: sadness, enjoyment, anger, disgust, fear, and surprise. We also merge it with project UIT-VSMEC's dataset. That is the Vietnamese language dataset.

System inputs an arbitrary tweet and assigns it to one of the classes that best reflects its sentiment. The significant results were that the Logistic Regression Classifier had the most excellent classification accuracy for this area out of the classification methods studied. After analysis, we create a report such as each emotion count, total tweets, and percentage of each emotion. Our system can analyze the sentiment of millions of Tweets in pseudo - real-time.

Vii Chapter 1 Problem Statement 1.1 Rationale With technology development, social networks have become a "gold mine" to collect customer reviews. Historically, businesses gathered feedback and insight into consumers' feelings about their goods via interviews, questionnaires, and surveys. These traditional methods were often extraordinarily time-consuming, expensive, and must be manual. Tweets are sometimes used to express opinions on a wide range of topics.

These ideas have an important effect in a variety of business decisions as well as in political opinions toward a certain candidate. Develop a sentiment analysis model that can extract customer reviews and detect the emotions that consumers score in order to be successful. From there, you may utilize this information to plan for the consolidation and development of goods and services that are more appealing to customers. Consumers can use sentiment analysis to research products or services before making a purchase., Kindle Marketers can use this to research public opinion of their company and products, or to analyze customer satisfaction., Election Polls Organizations can also use this to gather critical feedback about problems in newly released products., Brand Management (Nike, Adidas) 12 Aims We want to extract attributes from tweets and analyze their emotions, which might be characterized as sadness, enjoyment, anger, disgust, fear, or surprise.

The project UIT-dataset VSMEC's is used in conjunction with our dataset in order to perform emotion classification using the Spark framework. Following that, a real-time sentiment analysis will be performed. When the models are applied to the Vietnamese language, you can see how well they work and how accurate they are.3 Object and range of study We crawl tweets on Twitter and categorize the data so that it can be used as part of our collection. In order to tackle this challenge, we make use of our dataset, which contains more than 4000 tweets, as well as the project UIT-dataset.

VSMEC's The Vietnamese language is represented by the dataset. We used classic machine learning models in Spark (Logistic Regression, Naive Bayes, Decision Tree, and Random Forest) to compare the performance and accuracy of the models. We found that the models performed better and were more accurate. Vietnamese user reviews were used to determine the most appropriate model for sentiment analysis.

The input data is massive, so we need to use a framework for big data in this processing like Spark. Sometimes we need an analyst sentiment in past sentences. Nevertheless, For the most part, we apply sentiment analysis in real-time because sometimes we need results in real-time, which is pretty significant; continuously updating the result is good when we want to get a survey or comment of new services or goods to make a decision thing. Chapter 2 Related work In 2020, Mandloi and Patel [2] show that the evolution of social media platforms drew millions of users, like Twitter, where users may write 280 character tweets.

Tweets' low character count facilitates sentiment analysis. A daily average of 550 million tweets. Sentiment analysis of Twitter data becomes a proxy for societal attitudes. This research using Naive Bayes Classifier, Support Vector Machine (SVM)! Maximum Entropy Method?.

As an outcome, Mandloi and Patel developed a sentiment analysis method and used it in real-time applications. This algorithm is suitable for use in political or other types of review systems. The research demonstrates that machine learning techniques such as Naive Bayes have the best accuracy and may be considered baseline learning methods, although Maximum Entropy approaches are also rather successful in specific circumstances. Bouazizi and Ohtsuki [3] derived the Senta’ approach for classification.

Because new platforms such as Snapchat focused on video- and multimedia-based communication, Twitter kept some properties that make it a fascinating subject of data mining. Twitter creates tremendous amounts of data every day, and the number of users has increased dramatically. They offer a novel technique for sentiment analysis that categorizes tweets into seven types. The findings are promising: the data set utilized for multi-class sentiment analysis had a 60.

However, we think a better training set would be preferable.

Nội dung được bảo vệ bản quyền — Tải xuống đầy đủ

Phân Tích Đánh Giá Khách Hàng Realtime Bằng Machine Learning Trên Nền Tảng Big Data là một tài liệu chuyên sâu về việc ứng dụng machine learning và big data để phân tích và đánh giá khách hàng theo thời gian thực. Tài liệu này cung cấp cái nhìn toàn diện về cách thức xây dựng mô hình, xử lý dữ liệu lớn, và tối ưu hóa quy trình để đưa ra các quyết định kinh doanh nhanh chóng và chính xác. Đặc biệt, nó nhấn mạnh lợi ích của việc tích hợp các công nghệ tiên tiến để nâng cao trải nghiệm khách hàng và tăng hiệu quả hoạt động doanh nghiệp.

Nếu bạn quan tâm đến các phương pháp xử lý dữ liệu và machine learning, bạn có thể khám phá thêm về cải tiến giải thuật KMeans cho bài toán gom cụm dữ liệu chuỗi thời gian để hiểu rõ hơn về các kỹ thuật tối ưu hóa trong xử lý dữ liệu. Ngoài ra, nghiên cứu về các phương pháp học biểu diễn dữ liệu sẽ mang lại góc nhìn sâu sắc về cách biểu diễn thông tin hiệu quả. Cuối cùng, xây dựng mô hình phân lớp với tập dữ liệu nhỏ là một tài liệu hữu ích để tìm hiểu cách áp dụng machine learning trong điều kiện dữ liệu hạn chế.

Những tài liệu này sẽ giúp bạn mở rộng kiến thức và có cái nhìn đa chiều hơn về các ứng dụng của machine learning và big data trong thực tế.