VIETNAM NATIONAL UNIVERSITY, HANOI UNIVERSITY OF ENGINEERING AND TECHNOLOGY ---------- Nguyen Minh Chau LIFELONG MACHINE LEARNING METHODS AND ITS APPLICATION IN MULTI-LABEL CLASSIFICATION Major: Computer Science HANOI – 2019 VIETNAM NATIONAL UNIVERSITY, HANOI UNIVERSITY OF ENGINEERING AND TECHNOLOGY ---------- Nguyen Minh Chau LIFELONG MACHINE LEARNING METHODS AND ITS APPLICATION IN MULTI-LABEL CLASSIFICATION Major: Computer Science Supervisor: Assoc. Ha Quang Thuy HANOI – 2019 AUTHORSHIP “I hereby declare that the work contained in this thesis is of my own and has not been previously submitted for a degree or diploma at this or any other higher education institution. To the best of my knowledge and belief, the thesis contains no materials previously published or written by another person except where due reference or acknowledgement is made.” Signature:……………………………………………… i SUPERVISOR’S APPROVAL “I hereby approve that the thesis in its current form is ready for committee examination as a requirement for the Bachelor of Computer Science degree at the University of Engineering and Technology.” Signature: ……………………………………………… ii ACKNOWLEDGEMENT First of all, I would like to express my sincere and deepest gratitude to the teacher, Assoc. Ha Quang Thuy, who dedicated and instructed, encouraged and guided me during the research process.
Secondly, I would like to thank the teachers and students in Knowledge Technology Laboratory, especially Dr. Pham Thi Ngan and Mr. Nguyen Van Quang, for their enthusiasm to work with, comment and guide me while doing research as members of the research team. Thirdly, I sincerely thank the teachers and staff of the University of Engineering and Technology, Vietnam National University, Hanoi for creating favorable conditions for me to do research.
Finally, I want to thank my family and friends, especially my parents, those who always give me love, faith and encouragement. iii ABSTRACT Multi-label classification is a classification problem that classifies data which can have more than one label. Multi-label learning is very useful in text classification applications, however, there are many challenges in building training examples. One challenge is that we may not have a large amount of data for training when we face a new task.
In addition, even if we spend time to collect a large amount of data, labelling multi- label data is a very time-consuming task. Hence, we should have a model that could work well when we only have a small amount of training data. Lifelong Machine Learning (LML) is one possible approach to solve this problem. Lifelong Machine Learning (or Lifelong Learning) is an advanced machine learning paradigm that learns continuously, gathers the information learned in past tasks, and utilizations it to support future learning.
All the while, the learner turns out to be increasingly educated and compelling at learning. This learning capacity is one of the signs of human intelligence. Notwithstanding, the current dominant machine learning paradigm learns in isolation: given a training dataset, it runs machine learning algorithm on the dataset to create a model. It makes no endeavor to retain the information and use it in future learning.
In spite of the fact that this isolated learning paradigm has been exceptionally effective, it requires an extensive number of training examples, and is appropriate for well-characterized and restricted tasks. In correlation, we people can adapt adequately with a few examples since we have aggregated a great amount of information in the past which empowers us to learn with little information or exertion. Lifelong learning means to accomplish this ability. As statistical machine learning develops, the time has come to endeavor to break the isolated learning tradition and to examine lifelong learning figuring out how to bring machine learning higher than ever.
Applications, for example, intelligent assistants, chatbots, and physical robots that cooperate with humans and systems in real-life environments are also calling for such lifelong learning capabilities. Without the capacity to amass the information and use it to adapt more learning gradually, a system will presumably never be truly intelligent. iv TABLE OF CONTENTS ABSTRACT. iv TABLE OF CONTENTS.
v List of Figures. vii List of tables. Contributions and thesis format. Lifelong machine learning.
Definition of lifelong learning. System architecture of lifelong learning. Lifelong topic modeling. LTM: a lifelong topic model.
AMC: a lifelong topic model for small data. The closeness of previous datasets to the current dataset. The closeness of two datasets. Proposed model of lifelong topic modeling using close domain knowledge for multi- label classification.
RESULTS AND DISCUSSIONS. Experimental results and discussions. 40 vi List of Figures Figure 1. The system architecture of lifelong machine learning .8 Figure 2: The Lifelong Topic Model (LTM) system architecture.13 Figure 3: The AMC model system architecture.15 Figure 4: Entropy function with n = 2.
The lifelong topic model using close domain knowledge for multi-label classification .34 vii List of tables Table 1. Data division details. The experimental results with 50 reviews in D4 and using kNN, Decision Tree as classifying methods. The experimental results with 50 reviews in D4 and using Random Forest, MLP, AdaBoost, Gaussian Naïve Bayes as classifying methods.
The experimental results with 100 reviews in D4 and using kNN, Decision Tree as classifying methods. The experimental results with 100 reviews in D4 and using Random Forest, MLP, AdaBoost, Gaussian Naïve Bayes as classifying methods .39 viii TÓM TẮT Phân loại đa nhãn là lớp bài toán phân lớp mà đối tượng dữ liệu có thể có nhiều hơn một nhãn. Bộ học phân lớp đa nhãn rất hữu ích trong các ứng dụng phân loại văn bản, tuy nhiên, có nhiều thách thức trong việc xây dựng bộ ví dụ đào tạo. Một thách thức là chúng ta có thể không có một lượng lớn dữ liệu để đào tạo khi chúng ta đối mặt với một tác vụ mới.
Ngoài ra, ngay cả khi chúng ta dành thời gian để thu thập một lượng lớn dữ liệu, việc gán nhãn dữ liệu đa nhãn là một công việc rất tốn thời gian. Do đó, chúng ta nên có một mô hình có thể hoạt động tốt khi chúng ta chỉ có một lượng nhỏ dữ liệu đào tạo. Học máy suốt đời (Lifelong Machine Learning - LML) là một cách tiếp cận khả thi để giải quyết vấn đề này. Học máy suốt đời (hay học suốt đời) là một mô hình học máy tiên tiến, học liên tục, thu thập thông tin học được trong các tác vụ trước đây và sử dụng nó để hỗ trợ cho việc học trong tương lai.
Trong khi đó, bộ học ngày càng được có nhiều kiến thức và hiệu quả hơn trong việc học. Năng lực học tập này là một trong những dấu hiệu của trí tuệ con người. Mặc dù vậy, mô hình học máy phổ biến hiện tại lại học một cách cô lập: nó được cung cấp một tập dữ liệu đào tạo và chạy thuật toán học máy trên tập dữ liệu để tạo ra một mô hình. Nó không có nỗ lực để giữ lại thông tin và sử dụng nó trong việc học tập trong tương lai.
Mặc dù thực tế là mô hình học tập cô lập này có hiệu quả, nhưng nó đòi hỏi một số lượng lớn các ví dụ đào tạo, và chỉ phù hợp cho các bài toán được định nghĩa rõ ràng. Trong khi đó, con người chúng ta có thể học chỉ với một vài ví dụ vì chúng ta đã tổng hợp một lượng lớn thông tin trong quá khứ cho phép chúng ta học hỏi với ít thông tin hoặc nỗ lực. Học máy suốt đời có thể thực hiện khả năng này. Đã đến lúc nỗ lực phá vỡ truyền thống học máy cô lập và tiến tới học máy suốt đời để tìm ra cách đưa học máy lên một nấc cao hơn.
Các ứng dụng như trợ lý thông minh, chatbot và robot vật lý tương tác vớicon người và các hệ thống trong môi trường thực tế cũng cần đến khả năng học tập suốt đời. Nếu không có khả năng tích lũy thông tin và sử dụng nó để phục vụ việc học hỏi dần dần, một hệ thống có lẽ sẽ không bao giờ được coi là thực sự thông minh. ix ABBREVATIONS LML Lifelong Machine Learning ML Machine Learning AI Artificial Intelligence kNN k-nearest neighbors NBC Naive Bayes Classifier CART Classification and Regression Trees MLP Multilayer Perceptrons x Chapter 1 INTRODUCTION 1.1 Motivation Multi-label text classification is a problem with many practical applications. For example, you own a hotel and you care about what your customers refer to about your hotel service on the hotel website or on a social network (e.
Those aspects could be the attitude of the staff, the view from the hotel, the price of the room, the quality of the hotel food,. There are thousands or even tens of thousands of reviews on your website or on social network. However, it is very time consuming to read and classify each review. A good approach is that you have a classification of these reviews, however, assigning labels to thousands of reviews is also a waste of time and effort.
If there is a model that could classify well with just a small amount of labeled data, you will save your time and effort. The commonly used approach is machine learning. Chen and Liu [x] state that “Machine learning (ML) has been instrumental for the advances of both data analysis and artificial intelligence (AI)”. The ongoing accomplishment of profound learning conveys it 1 to another tallness.
ML algorithms have been utilized in practically all zones of software engineering and numerous zones of regular science, building, and sociologies. Commonsense applications are much progressively far reaching. Without effective ML algorithms, numerous ventures would not have prospered, e., Internet commerce and Web search. The current dominant paradigm for machine learning is to run a ML algorithm on a dataset to create a model.
The model is then applied in tasks in real-life. For both supervised learning and unsupervised learning, this is true. It is called isolated learning because it does not think about some other related data or the learned information. The crucial issue with this isolated learning paradigm is that it does not retain and accumulate knowledge learned before and use it in future learning.
This is not the way human learns. We people never learn isolation. We generally retain the information learned before and use it to support future learning and problem solving. That is the reason at whatever point we experience another circumstance or issue, we may see that numerous parts of it are not really new because of the fact that we have seen them in the past in some different contexts.
Without the capacity to collect information, a ML algorithm ordinarily needs a big number of training examples to learn effectively. For supervised learning, marking of data labels is regularly done physically, which is exceptionally time-consuming and tedious. Since the world is too complex with too many possible tasks, it is almost impossible to label a large number of examples for every possible task or application for an ML algorithm to learn. To make matters worse, everything around us additionally changes always, and the labeling in this way should be done continuously, which is an overwhelming task for people.
Even for unsupervised learning, gathering a substantial volume of data may not be possible in many cases. In contrast, we human beings seem to learn quite differently. We accumulate and maintain the knowledge learned from previous tasks and use it seamlessly in learning new tasks and solving new problems. Over time we learn more and more and become more and more knowledgeable, and more and more effective at learning.
Lifelong Machine Learning (LML) (or simply lifelong learning) aims to mimic this human learning process and capability.