Bộ GIÁO DỤC VÀ ĐÀO TẠO ĐẠI HỌC KINH TE thành PHố Hồ CHÍ MINH BÁO CÁO TỔNG KẾT ĐỀ TÀI NGHIÊN CỨU KHOA HỌC THAM GIA XÉT GIẢI THƯỞNG “NHÀ NGHIÊN CỨU TRẺ UEH” NĂM 2024 TOWARDS THE EFFECTIVENESS OF ATTENTION-FREE LANGUAGE MODELS FOR E-COMMERCE PLATFORM SENTIMENT ANALYSIS Thuộc nhóm chuyên ngành: Khoa học dữ liệu và trí tuệ nhân tạo TP. Hồ Chí Minh Ngày 17 tháng 2 năm 2024 1 Summary of the research While Transformers has gained its popularity in the modern deep learning stack, it has high complexity due to the Attention operation. Although other language models without the use of the Attention component have been proposed, whether they perform effectively in sentiment analysis tasks for e-commerce platform feedback remains understudied. Furthermore, while there have been many surveys and comparative studies on related methods for sentiment analysis, the majority of these studies focus on machine learning and traditional recurrent-based models such as RNN and LSTM without taking into consideration current attention-free language models.
In this work, we evaluate the performance of such attention-free models, namely BiLSTM, TextCNN, gMLP, and Hyena. We collected the dataset of e-commerce platform feedback and release it under the name UEH-Ecom. We then implement the aforementioned models and give a comprehensive analysis and comparison. Our findings show that the accuracy of Bidirectional LSTM, TextCNN, HyenaDNA, and gMLP achieved comparative results compared to RoBERTa with significantly less number of parameters.
In addition, among the considered attention-free models, even though Bidirectional LSTM obtained the highest accuracy, the difference compared to gMLP is tiny. Otherwise, gMLP also acquired the highest Fl score in the considered attention-free model family.1 Overview of the dataset].1 Cross-entropy loss.2 Experimental design for RoBERTaỊ.3 Experimental design for BÌLSTMỊ.4 Experimental design for TextCNN 50 6.5 Experimental design for HyenaDNA|.6 Experimental design for gMLP[. 53 |7 Results and Discussion] 53 7. 53 7,2 Interpretation of Results in the Context of E-Commerce Feedback] .3 Limitations and Future Work!.
56 18 Conclusion and Future Work! 57 A Proof of Cross-entropy Convexity 58 B Evidence of Paper 60 3 Contents 1 Summary of the research 12 Introduction! 9 2,1 Overview of the development of E-commerce platforms!.2 Research gaps and Motivation 10 2.3 Objective of the Research. 13 13 Literature Review 14 |4 Theoretical Framework! 17 4.1 Fundamentals of Language Processing!.1 The Importance of Language Processing!.2 History of NLP and Important Techniques!.2 Attention mechanism and Transformers modell.2 Sliding Window Attention.5 Preliminaries of Attention-Free Models!. 32 2 List of Tables 1 Samples from our dataset, including the content of the feedback, and its corresponding targets. 42 2 Hyperparameters for finetuning RoBERTubaseI.49 13 Performance metrics of various modelsl.
54 List of Figures 1 Number of visits for popular E-commerce platforms. Figure adapted from. 9 2 Popular E-commerce Platform Market Share in 02/2022. 10 3 A taxonomy of research development for efficient Transformers 1196]].
11 4 Overall workflow of comparing attention-free and transformer-based model. After preprocessing, the data is used to train 4 attention-free models and RoBERTa. The performance of these models is then eval- uated using the test data. 14 5 Attention visualizations [104], The model learns to attend to the words that are most relevant given the input word.
Here, the <E0S> is a special token that marks the end of the sentence and <pad> are padding tokens.! 22 4 6 Illustration of the Attention Mechanism. Here, Query Q and Key K are input matrices of dimensions N X N that represent the data and the aspects to focus on, respectively. Similarity Score s is calculated as the outer product of <2 and K (i. Attention Probabilities A is the softmax function applied to the similarity scores to obtain a probability distribution (i.
This represents the attention each part of the input should receive. Value V is another input matrix that represents the original data. It is used in the final computation of the output. Output Ơ is the output calculated as the dot product of the attention probabilities and the value (i.
This represents the final ‘attended* output. 22 7 An overview of the Transformer architecture [96]]I. 23 18 RoBERTa has the same architecture as BERTI. 25 9 Different attention variants architecture.
Figure adapted from Ị58Ị. While memory-compressed attention utilizes convolution to reduce the computation amount of keys and queries, local attention splits the sequences into separate blocks before feeding into the Masked I Multi-head Attention?]. 27 10 Illustration of sliding window attention according to different window patterns [6]. From left to right: vanilla n2 attention, sliding window attention, dilated sliding window, global+sliding window.
27 5 11 An overall design of Flash Attention. Left: GPU components hierarchy and its bandwidth and memory size. Right: An overall architecture of I FlashAttention. 28 12 LSTM Cell with three components: Input Gate, Forget Gate, and Output Gate 116311.
30 13 TextCNN architecture with 2 channels input [52^1. 36 15 gMLP architecture IỊ57Í 39 16 Distribution of the reviews1 length in the training and test set|. 42 17 An Examination of Sentiment Distribution and Sentence Lengths in the Test Set, (a) Depicts the sentiment distribution in the test set, (b) Illustrates the frequency of sentence lengths for both positive (blue) and negative (red) sentiments 42 18 Top 25 most common words in the training datasetl. 43 119 Tradeoff between Precision and Recall!.
46 20 Cross-entropy loss 48 21 BiLSTM architecture. Best view in digital format. 50 22 TextCNN architecture with our setting 51 6 23 Left: Comparison between Accuracy and Total Parameters. gMLP matches 97% the accuracy of RoBERTa while having 100 times fewer parameters.
Right: Validation loss over epochs of the trained models. We trained RoBERTa and Hyena for only 5 epochs as it starts to converge at this point. For other models, we trained for 30 epochs. 54 24 Models^ Training and Validation Loss after 10 epochs.
The bold lines depict the Training loss while the Validation loss is represented in 7 List of Abbreviations Abbreviation Meaning Al Artificial Intelligence BERT Bidirectional Encoder Representations from Transformers BiGRU Bidirectional Gated Recurrent Unit BiLSTM Bidirectional Long-short Term Memory FFN Feedforward Network FLOPS Floating point operations per second GELL' Gaussian Error Linear Unit gMLP Gated Multilayer Perceptron GPT Generative Pre-trained Transformer HBM High Bandwidth Memory HTML HyperText Markup Language NLP Natural Language Processing POS Part of Speech ReLU Rectified Linear Unit RNNs Recurrent Neural Networks RoBERTa Robustly optimized Bidirectional Encoder Representations from Transformers SGU Spatial Gating Unit SRAM Static random-access memory TextCNN Text Convolutional Neural Network URL Uniform Resource Locator 8 2 Introduction 2.1 Overview of the development of E-commerce platforms The evolution of e-commerce over the past five decades has been marked by signifi cant transformations thanks to the advancements and development of both consumer demands and technological progress. The last decade has witnessed an unprecedented surge in growth and widespread popularity worldwide, making E-commerce platforms an indispensable component in the modern commercial exchange industry. Monthly Visits to Companies Over Time Figure 1: Number of visits for popular E-commerce platforms. Figure adapted from |68| As of 2020, global e-commerce sales reached an impressive $4.13 trillion, reflecting an 18 percent increase from the previous year [64J.
In 2021, more than 2 billion people worldwide regularly engaged in e-commerce transactions. Mobile e-commerce, in particular, witnessed substantial growth, constituting nearly 73 percent of total sales in 2021. 9 Popular E-commerce Platform Market Share in 02/2022 In Vietnam, as illustrated in [Figure 2| Shopee accounted for roughly 66% of the total market share by February 2022. Following that is Lazada, Tiki, and Chotol with nearly 12% and 10% respectively.
E-commerce has evolved into an indispensable component of the global retail land scape. As global internet accessibility and adoption escalate, the number of online shoppers continues to rise. The future of retail is unmistakably shifting towards e- commerce, necessitating businesses to adapt and embrace this transformative change to remain competitive.2 Research gaps and Motivation Since its introduction in 2017 by [104], Transformers has been considered a break through in Al research and applications, leading to an unprecedented development in training language models on a large scale, namely [80, [9, 82,|26j|. Furthermore, this 10 successful development has gone beyond the scope of natural language processing to other fields, including computer vision [29.
Over the last few years, a massive body of research work has proposed more efficient Transformer models in an attempt to improve the performance of the vanilla Transformer, including. 7], a taxonomy of this development can be illustrated in |Figure| Lhartormer OKpn parrec ■"erceiver Rvoeec ai.zuz Jaaq^ at '2021 I ransformer-x Nystromformer Dai Ct al 2019 A ., : Hi HI ZIHV Memory / Memory Recurrence compressec uompressivG Downsampling 1 M w, ^. JU I ft franstormer Set I ransformer Rae et al zuw Lccciai,zoi9 Clusterformer Routine iWanget al 2020) Hinne Poonngtormer 1 ransormer * Hl zV/L, Reformer ranstnrmer (NUtv Cl a. 20231 Performer U4KH4I ZU/U unaKxnanjm 01 ai.ztizu: - IL Jig Uirc Learnable ««141.2020 Zaheer ec 41'2020 Low-Rank rans ormer Wlnauai Ml 7320’ Lonaformer swin lie Jijy e.
N ZWJI ranstormer Clustered Attention 1 IU «1 fll. ZUZ4 s nkhorn iVyas el aU 2020) nfnrmer Low Rank / _onc snort I ransrorrr ’avc<e* .zvzw wane Ct al. 2Ũ2DỒ Kerne s I ranstormer /nj*:*t zw I Fixed/Factorized/ Adaptive Random Patterns Sparse Random feature Attention Synthesizer »ave« ?. aw? DC-Net SSharc fransrormer ^^1 /vz 1 i nr.kw ae ransmrmer vo ice el al.
zuzu; auaai 2019 near Sparse I ransiormer Sparse Transformer Ju etauzuz- MV^rcccws eltf. zvzv image Transformer mttciai.zo 9 Switch p»ma'GUI MIS’ Product Key fransformer Axta I ranstormer ►ecu? CI aLZOZU Memory lamoieet auzuis: Scaling Transformer IMWur at *L20?\ Figure 3: A taxonomy of research development for efficient Transformers |96| Despite the aforementioned advancements, one bottleneck of the Transformer archi tecture lies in its Attention mechanism with &{n~} complexity, leading to inefficient training, thus large computational resources are often required to train Transformer based models. While much effort has been made to improve upon this, including adjusting the self-attention to have fixed patterns 120] to applying low-rank methods 11 115] and connecting the fixed blocks of sequences in a recurrent manner Eni, it has been shown that most of these modifications and proposal of novel Transformer variants does not lead to significantly improved performance |70| and that they have poor performance on long-range sequences modeling tasks |951. Therefore, it is necessary to explore other language models that operate without the use of the Attenion operation.
Recent advancements in attention-free language models have been made with the proposal of the novel capable model, including gMLP [57] which ultilizes Multilayer Perceptron (MLP) in combination with a spatial gating unit. Another featured model is Hyena [78, 72], a novel model that has subquadratic complexity by using long convolutions and data-controlled gating as an alternative for the Attention mechanism. Although there have been many comparative studies regarding the use of deep learning models for sentiment analysis |ịl2; 108[|231, most of them do not take into consideration recently proposed attention-free language models since they are relatively new. The question of the effectiveness of such models therefore remains understudied.
Motivated by this, in this work, we make a comprehensive comparison and analysis of attention- free language models in the context of sentiment analysis tasks for E-commerce platform feedback. In particular, wc first collected feedback from the Google Play Store and App Store for five different popular E-commerce platforms in Vietnam, including Shopee, Tiki, Lazada, Sendo, and Chotot. The final complete dataset includes nearly 90 thousand feedback, with over 72 thousand reviews in the training 12 set and 18 thousand in the test set. We then preprocess this data before training and evaluating the models.
Following that, we implement attention-free language models including traditional ones like BiLSTM |88|, TextCNN IĨ5ÕỊ1. and more advanced models, namely gMLP 1571 and Hyena [78]. We train these models from scratch and show that gMLP is the most capable attention-free model for E-commerce platform feedback sentiment analysis, with 87.