VIETNAM NATIONAL UNIVERSITY HOCHIMINH CITY UNIVERSITY OF INFORMATION TECHNOLOGY ADVANCED PROGRAM IN INFORMATION SYSTEMS NGUYEN HOANG NHAT REAL-TIME CUSTOMER REVIEWS ANALYSIS USING MACHINE LEARNING ON BIG DATA FRAMEWORK BACHELOR OF ENGINEERING IN INFORMATION SYSTEMS HO CHI MINH CITY, 2021 NATIONAL UNIVERSITY HOCHIMINH CITY UNIVERSITY OF INFORMATION TECHNOLOGY ADVANCED PROGRAM IN INFORMATION SYSTEMS NGUYEN HOANG NHAT - 17520851 REAL-TIME CUSTOMER REVIEWS ANALYSIS USING MACHINE LEARNING ON BIG DATA FRAMEWORK BACHELOR OF ENGINEERING IN INFORMATION SYSTEMS THESIS ADVISOR HOP DO TRONG, Ph.D HO CHI MINH CITY, 2021 ASSESSMENT COMMITTEE The Assessment Committee is established under the Decision. Ho Chi Minh city, date January 18" by Rector of the University of Information Technology. 1 Tho Quan Thanh, ph. 2 Thanh Ngo Duc, Ph.
3 Nhan Cao Thi, Ph.D - Commissary ACKNOWLEDGMENTS First of all, I thank God for giving the support and providing me with the patience and guidance which I used in making this work. I would like to thank my supervisor, Hop Do Trong, Ph. I would like to thank him for his help, continuous encouragement, productive discussion, and valuable suggestions and comments throughout the research and thesis work. It has been a fantastic experience studying at the University of Information Technology.
I was motivated and excited about studying, ultimately fulfilling my expectations. Last but not least important, I owe more than thanks to my family members especially. Thanks to my brother for his support and encouragement throughout my life. TABLE OF CONTENTS cae ACKNOWLEDGMENT .scscsssssssesesstssssesssstsnesesesansnesesesaeanesesesaeaesnesesanatenesesenaeaeaseneranaes i LIST OF FIGURES LIST 9).
Chapter 1 Problem Statemeni. --- «se «HH HH ghi 1 1. Object and range of stud: Chapter 2 Related Work.csssssssesessssssssssnseescsesessssssssneessesssoseesssneeeaneneaeeeseesseneneneaeeeeees 3 Chapter 3 Methodolog). HH HẤ HH nghe7 3.1 Core technologies of spark and its components.
Big Data Analysis Using Spark.1 Machine learning proce: 3.2 Machine learning algorithms 3.2 Sentiment Analysis of Twitter Data 3. Chapter 4 Experimental ReSUItS.1 Prepare data for the real-time problem 4.2 Dataset for train-test.1 TF-IDF vectorizer + Logi 4.2 TF-IDF vectorizer + NaiveBayes Classifi 4. TF-IDF vectorizer + Decision Tree Classifier 4.4 TF-IDF vectorizer + Random Forest Classifier. AA Compare MOdel .5 Real-time sentiment analyst.6 Compare with UIT-VSMEC.c tình nhe 56 Chapter 5 Summary.
Chapter 6 Future WOIFK .- ---- «<< << HH nh ghi 58 REEERENCES.-- 5-5 5< SH HH Họ TH TH HH HH HA HH BE 0.}gg000/(02Ề0sS6ố2n111111111 61 iii LIST OE FIGURES œELlb Figure 3.The three Vs of big đata.2 The ecosystem of Spark.3 Apache Spark data processing .4 A typical process of machine learning [17] .5 Summary of machine learning algorithms [18] .6 Sentiment analyst step.7 Attributes available in snscrape tweet object.8 Code for crawl Vietnamese EW€€ts.9 Dataframe of tW€ẨS.- - nh HH HH HH Hà HH HH HH HH HH HH key 17 Figure 3.10 Pipeline for spark nÍp.- - -- «5+ ềE* St E111 Tàn Tàn Tàn Tàn Tàn TH Tàn Hàn nhe 24 Figure 4.1 Format of data test in r€aÏ-tÏIT.2 Separate the data set Figure 4.3 Cleaning data Figure 4.4 Libs for training model Figure 4.5 Dictionary of emotion Figure 4.6 TF-IDF vectorizer + Logistic Regression Classifier pipeline.7 TF-IDF vectorizer + NaiveBayes Classifier pipeline .8 TF-IDF vectorizer + Decision Tree Classifier pipeline.9 TF-IDF vectorizer + Random Forest Classifier pipeline.10 Step for real-time sentiment.11 Result realtime predict sentiment of tweets .12 Count of Emotion in realtime .13 Visualizing number of emotion in real-time Figure 4.14 Calculate Average of latency in real-time.15 Percentage each EmOtIO.- - ¿c5 St ÉEk‡ÉE kg TH HH ưêc 55 Figure 4.16 Visualizing of Percentage of each Emotion .17 Count sentence is already annaÏySí.--- + c5 sét ey 55 iv LIST OE TABLES œELlb Table 3.1 Step for data preDFOC€SSINE.2 : Statistics of emotion labels of the UIT-VSMEC corpus [Ï].3 Statistics of emotion labels of the Our dataset .4 Statistics of emotion labels of the mix dataset .1 TF-IDF vectorizer + Logistic Regression Classifier Table 4.2 TF-IDF vectorizer + Logistic Regression Classifier with my data train + UIT data test.3 TF-IDF vectorizer + Logistic Regression Classifier with my data train + mix data test.4 Compare accuracy using TF-IDF vectorizer + Logistic Regression Classifier with my data using for training .5 TF-IDF vectorizer + Logis data t€SE.6 TF-IDF vectorizer + Logistic Regression Classifier with UIT data train + UIT data tesf.7 TF-IDF vectorizer + Logistic Regression Classifier with UIT data train + mix data test .8 Compare accuracy using TF-IDF vectorizer + Logistic Regression Classifier with UIT data using for training.9 TF-IDF vectorizer + Logistic Regression Classifier with mix data train + my data test.10 TF-IDF vectorizer + Logistic Regression Classifier with mix data train + UIT data test Table 4.11 TF-IDF vectorizer + Logistic Regression Classifier with mix data train + mix data (€SỂ.12 Compare accuracy using TF-IDF vectorizer + Logistic Regression Classifier with MIX data using for training Table 4.13 TF-IDF vectorizer + NaiveBayes Classifier with my data train + my data test Table 4.14 TF-IDF vectorizer + NaiveBayes Classifier with my data train + UIT data ` —.15 TF-IDF vectorizer + NaiveBayes Classifier with my data train + mix data Ôn .ốốốốỐốỐố ốốốốố ỐC CC Cố CC Cố 0000 36 Table 4.16 TF-IDF vectorizer + NaiveBayes Classifier with UIT data train + my data LOSE eee .17 TF-IDF vectorizer + NaiveBayes Classifier with UIT data train + UIT data ` —.21 TF-IDF vectorizer + NaiveBayes Classifier mix data train + MIX data test .22 Compare accuracy using TF-IDF vectorizer + NaiveBayes Classifier 40 41 Table 4.24 TF-IDF vectorizer + Decision Tree Classifier with my data train + UIT data `.25 TF-IDF vectorizer + Decision Tree Classifier with my data train + mix data {€S{.26 TF-IDF vectorizer + Decision Tree Classifier with UIT data train + my data W€S.27 TF-IDF vectorizer + Decision Tree Classifier with UIT data train + UIT data test.28 TF-IDF vectorizer + Decision Tree Classifier with UIT data train + mix data test.30 TF-IDF vectorizer + Decision Tree Classifier with mix data train + UIT data tet.31 TF-IDF vectorizer + Decision Tree Classifier with mix data train + mix ata t€SỂ.32 Compare accuracy using TF-IDF vectorizer + Decision Tree Classifier.35 TF-IDF vectorizer + Random Forest Classifier with my data train + mix lon T1 .36 TF-IDF vectorizer + Random Forest Classifier with UIT data train + my 6n ae.37 TF-IDF vectorizer + Random Forest Classifier with UIT data train + UIT data test.38 TF-IDF vectorizer + Random Forest Cl data test.39 TF-IDF vectorizer + Random Forest Classifier with mix data train + my data test.40 TF-IDF vectorizer + Random Forest Classifier with mix data train + UIT data test.41 TF-IDF vectorizer + Random Forest Classifier with mix data train + mix data E€SỂ. co nh re Table 4.42 Compare accuracy using TF-IDF vectorizer + Random Forest Classifier.43 Compare each model .44 Compare with UIT-VSMEC.--- tt ng TH Hư 56 vii ABSTRACT For firms to monitor their brand reputations and assess their performance and public feelings about their goods, have many ways, but nowadays, using automation to sentiment analyst of customer reviews is the easiest. We reveal our efforts to develop a machine learning-based system for sentiment analysis of tweets on Twitter and make it a real-time analyst in this study. We use the framework ‘Apache Spark' — Bigdata framework.
Our system outperforms the Emotion Recognition for Vietnamese Social Media Text, 2019 (UIT-VSMEC) [1]. We make a dataset using a tweets crawler and label the dataset that best reflects its sentiment: sadness, enjoyment, anger, disgust, fear, and surprise. We also merge it with project UIT-VSMEC's dataset. That is the Vietnamese language dataset.
System inputs an arbitrary tweet and assigns it to one of the classes that best reflects its sentiment. The significant results were that the Logistic Regression Classifier had the most excellent classification accuracy for this area out of the classification methods studied. After analysis, we create a report such as each emotion count, total tweets, and percentage of each emotion. Our system can analyze the sentiment of millions of Tweets in pseudo - real-time.
Vii Chapter 1 Problem Statement 1.1 Rationale With technology development, social networks have become a "gold mine" to collect customer reviews. Historically, businesses gathered feedback and insight into consumers' feelings about their goods via interviews, questionnaires, and surveys. These traditional methods were often extraordinarily time-consuming, expensive, and must be manual. Tweets are sometimes used to express opinions on a wide range of topics.
These ideas have an important effect in a variety of business decisions as well as in political opinions toward a certain candidate. Develop a sentiment analysis model that can extract customer reviews and detect the emotions that consumers score in order to be successful. From there, you may utilize this information to plan for the consolidation and development of goods and services that are more appealing to customers. Consumers can use sentiment analysis to research products or services before making a purchase., Kindle Marketers can use this to research public opinion of their company and products, or to analyze customer satisfaction., Election Polls Organizations can also use this to gather critical feedback about problems in newly released products., Brand Management (Nike, Adidas) 12 Aims We want to extract attributes from tweets and analyze their emotions, which might be characterized as sadness, enjoyment, anger, disgust, fear, or surprise.
The project UIT-dataset VSMEC's is used in conjunction with our dataset in order to perform emotion classification using the Spark framework. Following that, a real-time sentiment analysis will be performed. When the models are applied to the Vietnamese language, you can see how well they work and how accurate they are.3 Object and range of study We crawl tweets on Twitter and categorize the data so that it can be used as part of our collection. In order to tackle this challenge, we make use of our dataset, which contains more than 4000 tweets, as well as the project UIT-dataset.
VSMEC's The Vietnamese language is represented by the dataset. We used classic machine learning models in Spark (Logistic Regression, Naive Bayes, Decision Tree, and Random Forest) to compare the performance and accuracy of the models. We found that the models performed better and were more accurate. Vietnamese user reviews were used to determine the most appropriate model for sentiment analysis.
The input data is massive, so we need to use a framework for big data in this processing like Spark. Sometimes we need an analyst sentiment in past sentences. Nevertheless, For the most part, we apply sentiment analysis in real-time because sometimes we need results in real-time, which is pretty significant; continuously updating the result is good when we want to get a survey or comment of new services or goods to make a decision thing. Chapter 2 Related work In 2020, Mandloi and Patel [2] show that the evolution of social media platforms drew millions of users, like Twitter, where users may write 280 character tweets.
Tweets' low character count facilitates sentiment analysis. A daily average of 550 million tweets. Sentiment analysis of Twitter data becomes a proxy for societal attitudes. This research using Naive Bayes Classifier, Support Vector Machine (SVM)! Maximum Entropy Method?.
As an outcome, Mandloi and Patel developed a sentiment analysis method and used it in real-time applications. This algorithm is suitable for use in political or other types of review systems. The research demonstrates that machine learning techniques such as Naive Bayes have the best accuracy and may be considered baseline learning methods, although Maximum Entropy approaches are also rather successful in specific circumstances. Bouazizi and Ohtsuki [3] derived the Senta’ approach for classification.
Because new platforms such as Snapchat focused on video- and multimedia-based communication, Twitter kept some properties that make it a fascinating subject of data mining. Twitter creates tremendous amounts of data every day, and the number of users has increased dramatically. They offer a novel technique for sentiment analysis that categorizes tweets into seven types. The findings are promising: the data set utilized for multi-class sentiment analysis had a 60.
However, we think a better training set would be preferable.