VIETNAM NATIONAL UNIVERSITY HO CHI MINH CITY HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY NGUYEN CONG THANH ADVERTISEMENT ASSIGNMENT SYSTEM Major: Computer Science Major code: 8480101 MASTER’S THESIS HO CHI MINH CITY, June 2024 THIS THESIS IS COMPLETED AT HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY – VNU-HCM Supervisor: 1. Thoai Nam, Ph. Nguyen Quang Hung, Ph.D Examiner 1: Ha Viet Uyen Sinh, Ph.D Examiner 2: Nguyen Le Duy Lai, Ph.D This master’s thesis is defended at HCM City University of Technology, VNU- HCM City on June 18, 2024 Master’s Thesis Committee: 1. Chairman: Le Thanh Sach, Ph.
Secretary: Le Thanh Van, Ph. Examiner 1: Ha Viet Uyen Sinh, Ph. Examiner 2: Nguyen Le Duy Lai, Ph. Huynh Tuong Nguyen, Ph.D Approval of the Chair of Master’s Thesis Committee and Dean of Faculty of Computer Science and Engineering after the thesis being corrected (If any).
CHAIR OF THESIS COMMITTEE DEAN OF FACULTY OF COMPUTER SCIENCE AND ENGINEERING VIETNAM NATIONAL UNIVERSITY - HO CHI MINH CITY SOCIALIST REPUBLIC OF VIETNAM HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY Independence – Freedom - Happiness THE TASK SHEET OF MASTER’S THESIS Full name: Nguyen Cong Thanh Student ID: 2170573 Date of birth: 23/06/1999 Place of birth: Vinh Long Major: Computer Science Major ID: 8480101 I. THESIS TITLE (In Vietnamese): Hệ thống gán quảng cáo cho người dùng II. THESIS TITLE (In English): Advertisement assignment system III. TASKS AND CONTENTS: Examine the effects of metadata enrichment and its applications on improving the performance of personalized recommendation systems that support advertisement assigning purposes.
THESIS START DAY: 15/01/2024 V. THESIS COMPLETION DAY: 20/05/2024 VI. Thoai Nam, Ph.D and Nguyen Quang Hung, Ph.D Ho Chi Minh City, 15/08/2024 SUPERVISOR 1 SUPERVISOR 2 CHAIR OF PROGRAM COMMITTEE DEAN OF FACULTY OF COMPUTER SCIENCE AND ENGINEERING Acknowledgements First and foremost, we would like to express our gratitude to our supervisor, Ph. Thoai Nam, for his guidance throughout the course of this research.
Without his help, we would not know where to start nor how to move on every time the research encounters an obstacle. Thanh would also like to thank his family for their unwavering support and encouragement. His parents have a constant incredible belief in him, no matter the ups and downs of the journey. His older sister and his brother-in-law have always taken care of everything, he has not yet had to worry about anything else but the research and his job.
His two nephews have always loved his company even though he spent most of his time around them not playing but just studying on his laptop. Finally, Thanh would like to thank his girlfriend, Quy, for her patience and understanding. It is tricky and sometimes annoying to be with someone who is always busy with work, especially when her herself too has a lot on her plate. Despite all that, she has never stopped taking care of him and bringing him warmth in his difficult moments.
He has not yet heard the slightest complaint from her, and he is grateful for that. The students implementing the topic i Abstract English Advertisement assignment systems make heavy use of recommendation sys- tems to make sure the ads that are shown to users are the most relevant ones, hence maximizing the revenue of the system and the enhancing the experience of the users. Personalized recommendation systems are a type of recommendation system whose decisions take into account the users that one is recommending to. Since the proposition of Factorization Machines by Rendle et al.
in 2010, many of its variants such as the Neural Collaborative Filtering by He et al. and the DeepFM by Guo et al. have shown to be effective in capturing the interactions between users and items in the recommendation system, even in the extremely low density of interaction. One of the greatest premises of Factorization Machines and its variants is the ability to incorporate more metadata about the users and items involved in the interactions, rather than just the mere presence of the interactions themselves.
Yet, there is a lack of research on the true impact of incorporating more metadata on the performance of these models. Hence, this research aims to investigate such impact by extensively enriching the metadata of a recommendation dataset then analyzing the necessary configuration and performance of the models with several levels of metadata enrichment. The results display promising improvements as a result of careful use of metadata, with some insights on how well combining multiple features can perform given how well each feature performs on its own. ii Tiếng Việt Hầu hết các hệ thống gán quảng cáo cho người dùng đều cần một hệ thống gợi ý hiệu quả để đảm bảo rằng quảng cáo được hiện thị cho người dùng phù hợp.
Điều này sẽ giúp tối ưu lợi nhuận cho doanh nghiệp và tạo ra trải nghiệm người dùng tốt nhất. Các hệ thống gợi ý theo dạng cá nhân hoá là một dạng của hệ thống gợi ý, tại đó hệ thống có cân nhắc tới bản thân người dùng khi nó đưa ra gợi ý cho người dùng đó. Từ khi Rendle et al. đề xuất mô hình Factorization Machines vào năm 2010, nhiều biến thể của mô hình này như Neural Collaborative Filtering của He et al.
và DeepFM của Guo et al. đã cho thấy sự hiệu quả trong việc mô hình hóa tương tác giữa người dùng và sản phẩm trong hệ thống gợi ý, ngay cả khi mật độ tương tác rất thấp. Một trong những điều hứa hẹn nhất của Factorization Machines và các biến thể của nó là khả năng tích hợp thêm metadata về người dùng và sản phẩm tham gia vào tương tác, thay vì chỉ dựa vào việc có hay không các tương tác giữa các thực thể này. Vậy mà trên thực tế, chưa có nhiều nghiên cứu về tác động thực sự của việc tích hợp metadata vào hiệu suất của các mô hình này.
Vì vậy, công trình nghiên cứu của chúng tôi nhắm tới việc tìm hiểu về tác động đó. Để làm điều đó, chúng tôi đã làm giàu metadata của một tập dữ liệu gợi ý rồi phân tích những cải tiến về khả năng gợi ý khi metadata được làm giàu, cũng như những thiết lập cần phải làm ở các kích thước khác nhau của tập dữ liệu. Chúng tôi nhận thấy sự cải thiện đáng kể về khả năng gợi ý chỉ từ việc làm giàu metadata một cách có tính toán, cũng như thấy được việc kết hợp các đặc trưng giúp ích như thế nào so với việc sử dụng từng đặc trưng đơn lẻ. iii Commitment We hereby declare that all the work presented in this report, as well as the source code, is entirely our own - except for referenced knowledge and sample code provided by the manufacturer, and is not copied from any other source.
If this commitment is contrary to the truth, we are willing to take full responsibility before the Department Head and the School Administration. The students implementing the topic iv Contents Acknowledgements i Abstract ii Commitment iv Images x Tables xi Abbreviations & Acronyms xii 1 Introduction 1 1.2 The Personalized Recommendation System Problem Statement .1 Artificial Neural Networks .2 Hyperbolic Tangent (tanh) Function .3 Rectified Linear Unit (ReLU) Function .4 Leaky ReLU Function .3 Evaluation Metrics for Recommendation Systems .2 Normalized Discounted Cumulative Gain (nDCG) .1 Neural Collaborative Filtering .4 Feature Quantity Impact on Movielens-100K .1 The MovieLens 20M Dataset. 21 5 Metadata Enrichment For Personalized Recommendation Systems 24 5.3 Test And Validation Data .5 Multivalued Feature Outliers Elimination .1 Hyperparameters Fine-tuning .1 The Effect of Multivalued Feature Outliers Control .2 Feature Importance Analysis .3 Feature Quantity Analysis .4 Comparison With Other Models. 40 References 41 vii List of Figures 3.1 An example illustrates MF’s limitation taken from the work of He et al.
From data matrix (a), u4 is most similar to u1 , following by u3 , and lastly u2. However, in the latent space (b), placing p4 closet to p1 makes p4 closer to p2 than p3 resulting in a large ranking loss.2 The architecture of Neural Collaborative Filtering taken from the work of He et al.3 The architecture of DeepFM taken from the work of Guo et al. [2] which they described as the wide & deep architecture of DeepFM. The wide and deep components share the same input feature vec- tors, which enables DeepFM to learn low- and high-order feature interactions simultaneously.4 The spectrum of various Factorization Machines variants taken from the work of Guo et al.5 The list of algorithms used in the study taken from the work of Wegmeth et al.
[7] with their category. Although not discussed in depth in the paper, the parameter tuning process for each algorithm can be found in the last column.6 The list of feature sets used in the study taken from the work of Wegmeth et al. The feature sets are generated from the base Movielens-100K [ 13] data set either by cutting or enriching features. The columns denote the contained features in the named feature sets and the total number of features.
18 viii List of Figures 3.7 The performance of the algorithms taken from the work of Weg- meth et al. The performance of the algorithms is measured by the RMSE metric. The performance of the ML and AutoML algorithms is visibly improved as the number of features increases, while the performance of the RecSys algorithms does not have any recognizable differences.8 The feature importance of the random forest regressor taken from the work of Wegmeth et al. The feature importance is measured by the Gini importance metric using their trained random forest regressor, ordered from the most important to the least important feature.
Statistical features are refixed with ‘i’ and ‘u’ to denote if the feature is an item or a user feature.1 Distributions of ratings per user and per movie. Both distributions are long-tail distributions.2 Distribution of the Number of Genres per Movie .1 The process of collecting additional metadata for the movies in the MovieLens-20M dataset .2 Distribution of the number of cast members in each movie. The horizontal axis is the number of cast members, and the vertical axis is the number of movies with that number of cast members.3 Distribution of the number of members in camera department, visual effects, art department and producers of each movie. The horizontal axis is the number of members, and the vertical axis is the number of movies with that number of members.4 The adaptation of DeepFM architecture as the number of features increases.
As the number of features increases, the number of inputs in the Factorization Machine and the feed-forward neural network also increases. 30 ix List of Figures 5.5 The improvement in RMSE when using various maximum number of cast members in each movie subtracted by that of the baseline. The horizontal axis is the maximum number of cast members, and the vertical axis is the improvement in RMSE. Higher is better.
The best learning rate (0.0511 RMSE reduction) is annotated with an light pink dot. The zero-improvement baseline is shown as a horizontal dash line.6 The improvement in HR@10 when using various features subtracted by that of the baseline. The horizontal axis is the feature, and the vertical axis is the improvement in HR@10. Higher is better.7 The improvement in HR@10 when using various number of features subtracted by that of the baseline.
The horizontal axis is the number of features, and the vertical axis is the improvement in HR@10. Higher is better. The best number of features (6, with 84.123% HR@10) is annotated with an light pink dot. 36 x List of Tables 4.1 Statistics of the MovieLens-20M Dataset .1 The mumber of features collected from IMDb and TMDb, respectively 25 5.2 Number of features before and after data collection .