Luận Văn Thạc Sĩ: Các Phương Pháp Học Sâu Tiên Tiến và Ứng Dụng Trong Hệ Hỏi Đáp Miền Mở

Luận văn thạc sĩ VNU UET trình bày các phương pháp học sâu tiên tiến và ứng dụng trong hệ hỏi đáp miền mở, mang lại cái nhìn sâu sắc về công nghệ.

Chuyên ngành

Computer Science

Người đăng

Ẩn danh

Thể loại

master thesis

2019

67
1
0

Phí lưu trữ

30 Point

Mục lục chi tiết

Abstract

Acknowledgements

Declaration

1. CHƯƠNG 1: INTRODUCTION

1.1. Open-domain Question Answering

1.2. Difficulties and Challenges

1.3. Deep learning

List of Figures

List of Tables

List of Publications

Acronyms

Tóm tắt

I. Tổng Quan Về Các Phương Pháp Học Sâu Trong Hệ Hỏi Đáp Miền Mở

Trong thời đại thông tin hiện nay, việc truy xuất thông tin từ các hệ thống là rất quan trọng. Các phương pháp học sâu đã trở thành một công cụ mạnh mẽ trong việc phát triển các hệ thống hỏi đáp miền mở. Hệ thống này không chỉ giúp người dùng tìm kiếm thông tin mà còn cung cấp câu trả lời chính xác và ngắn gọn cho các câu hỏi. Sự phát triển của machine learningtrí tuệ nhân tạo đã thúc đẩy nghiên cứu trong lĩnh vực này, mở ra nhiều cơ hội mới cho việc cải thiện khả năng truy xuất và hiểu biết tài liệu.

1.1. Khái Niệm Về Hệ Hỏi Đáp Miền Mở

Hệ hỏi đáp miền mở cho phép người dùng đặt câu hỏi mà không bị giới hạn bởi một lĩnh vực cụ thể. Điều này có nghĩa là hệ thống có thể truy cập vào một lượng lớn dữ liệu không cấu trúc từ nhiều nguồn khác nhau, như Wikipedia hay các cơ sở dữ liệu trực tuyến khác.

1.2. Vai Trò Của Học Sâu Trong Hệ Hỏi Đáp

Các phương pháp học sâu giúp cải thiện khả năng hiểu ngữ nghĩa của câu hỏi và tài liệu. Chúng cho phép hệ thống học từ dữ liệu lớn, từ đó nâng cao độ chính xác trong việc trả lời câu hỏi.

II. Những Thách Thức Trong Hệ Hỏi Đáp Miền Mở

Mặc dù có nhiều tiến bộ, nhưng việc phát triển hệ thống hỏi đáp miền mở vẫn gặp phải nhiều thách thức. Một trong những vấn đề lớn nhất là khả năng truy xuất tài liệu chính xác từ một kho dữ liệu khổng lồ. Hệ thống cần phải nhanh chóng và hiệu quả trong việc tìm kiếm thông tin, đồng thời đảm bảo độ chính xác cao.

2.1. Vấn Đề Về Quy Mô Dữ Liệu

Với lượng dữ liệu khổng lồ trên Internet, việc xử lý và truy xuất thông tin trở nên khó khăn. Hệ thống cần có khả năng xử lý dữ liệu lớn và thay đổi liên tục để cung cấp câu trả lời chính xác.

2.2. Độ Chính Xác Của Hệ Thống

Một thách thức khác là đảm bảo rằng các tài liệu được truy xuất không chỉ liên quan mà còn phải cung cấp thông tin chính xác. Hệ thống cần có khả năng hiểu ngữ nghĩa của câu hỏi để tránh trả lời sai.

III. Phương Pháp Học Sâu Tiên Tiến Trong Hệ Hỏi Đáp

Các phương pháp học sâu tiên tiến như mô hình học tự chú ýhọc để xếp hạng đã được áp dụng để cải thiện hiệu suất của hệ thống hỏi đáp miền mở. Những phương pháp này cho phép hệ thống học từ các tài liệu và câu hỏi một cách hiệu quả hơn.

3.1. Mô Hình Học Tự Chú Ý

Mô hình học tự chú ý giúp hệ thống tập trung vào các phần quan trọng của tài liệu khi trả lời câu hỏi. Điều này cải thiện khả năng hiểu và phân tích ngữ nghĩa của câu hỏi.

3.2. Học Để Xếp Hạng

Phương pháp học để xếp hạng cho phép hệ thống đánh giá và xếp hạng các tài liệu dựa trên độ liên quan của chúng với câu hỏi. Điều này giúp cải thiện độ chính xác trong việc chọn tài liệu phù hợp.

IV. Ứng Dụng Thực Tiễn Của Hệ Hỏi Đáp Miền Mở

Hệ thống hỏi đáp miền mở có nhiều ứng dụng thực tiễn trong các lĩnh vực như giáo dục, chăm sóc sức khỏe và dịch vụ khách hàng. Chúng giúp người dùng nhanh chóng tìm kiếm thông tin và giải quyết vấn đề một cách hiệu quả.

4.1. Ứng Dụng Trong Giáo Dục

Trong giáo dục, hệ thống này có thể hỗ trợ học sinh và sinh viên tìm kiếm thông tin nhanh chóng, giúp họ giải quyết các câu hỏi trong quá trình học tập.

4.2. Ứng Dụng Trong Chăm Sóc Sức Khỏe

Trong lĩnh vực chăm sóc sức khỏe, hệ thống hỏi đáp miền mở có thể cung cấp thông tin y tế chính xác và kịp thời cho bệnh nhân và bác sĩ.

V. Kết Luận Và Tương Lai Của Hệ Hỏi Đáp Miền Mở

Hệ thống hỏi đáp miền mở đang trên đà phát triển mạnh mẽ nhờ vào các phương pháp học sâu tiên tiến. Tương lai của lĩnh vực này hứa hẹn sẽ mang lại nhiều cải tiến trong khả năng truy xuất và hiểu biết thông tin.

5.1. Xu Hướng Phát Triển

Các nghiên cứu trong tương lai sẽ tập trung vào việc cải thiện độ chính xác và tốc độ của hệ thống, đồng thời mở rộng khả năng truy xuất thông tin từ nhiều nguồn khác nhau.

5.2. Tác Động Đến Người Dùng

Hệ thống hỏi đáp miền mở sẽ tiếp tục cải thiện trải nghiệm người dùng, giúp họ dễ dàng tìm kiếm và truy xuất thông tin một cách hiệu quả hơn.

22/07/2025
Luận văn thạc sĩ vnu uet advanced deep learning methods and applications in opendomain question answering các phương pháp học sâu tiên tiến và ứng dụng vào bài toán hệ hỏi đáp miền mở

Trích đoạn nội dung tài liệu

VIETNAM NATIONAL UNIVERSITY, HANOI UNIVERSITY OF ENGINEERING AND TECHNOLOGY Nguyen Minh Trang ADVANCED DEEP LEARNING METHODS AND APPLICATIONS IN OPEN-DOMAIN QUESTION ANSWERING MASTER THESIS Major: Computer Science HA NOI - 2019 LUAN VAN CHAT LUONG download : add luanvanchat@agmail.com VIETNAM NATIONAL UNIVERSITY, HANOI UNIVERSITY OF ENGINEERING AND TECHNOLOGY Nguyen Minh Trang ADVANCED DEEP LEARNING METHODS AND APPLICATIONS IN OPEN-DOMAIN QUESTION ANSWERING MASTER THESIS Major: Computer Science Supervisor: Assoc. Ha Quang Thuy Ph. Nguyen Ba Dat HA NOI - 2019 LUAN VAN CHAT LUONG download : add luanvanchat@agmail.com Abstract Ever since the Internet has become ubiquitous, the amount of data accessible by information retrieval systems has increased exponentially. As for information con- sumers, being able to obtain a short and accurate answer for any query is one of the most desirable features.

This motivation, along with the rise of deep learning, has led to a boom in open-domain Question Answering (QA) research. An open- domain QA system usually consists of two modules: retriever and reader. Each is developed to solve a particular task. While the problem of document compre- hension has received multiple success with the help of large training corpora and the emergence of attention mechanism, the development of document retrieval in open-domain QA has not gain much progress.

In this thesis, we propose a novel encoding method for learning question-aware self-attentive document represen- tations. Then, these representations are utilized by applying pair-wise ranking approach to them. The resulting model is a Document Retriever, called QASA, which is then integrated with a machine reader to form a complete open-domain QA system. Our system is thoroughly evaluated using QUASAR-T dataset and shows surpassing results compared to other state-of-the-art methods.

Keywords: Open-domain Question Answering, Document Retrieval, Learning to Rank, Self-attention mechanism. iii LUAN VAN CHAT LUONG download : add luanvanchat@agmail.com Acknowledgements Foremost, I would like to express my sincere gratitude to my supervisor Assoc. Ha Quang Thuy for the continuous support of my Master study and research, for his patience, motivation, enthusiasm, and immense knowledge. His guidance helped me in all the time of research and writing of this thesis.

I would also like to thank my co-supervisor Ph. Nguyen Ba Dat who has not only provided me with valuable guidance but also generously funded my re- search. My sincere thanks also goes to Assoc. Chng Eng-Siong and M.

Vu Thi Ly for offering me the summer internship opportunities in NTU, Singapore and leading me working on diverse exciting projects. I thank my fellow labmates in KTLab: M. Le Hoang Quynh, B. Can Duy Cat, B.

Tran Van Lien for the stimulating discussions, and for all the fun we have had in the last two years. Last but not the least, I would like to thank my parents for giving birth to me at the first place and supporting me spiritually throughout my life. iv LUAN VAN CHAT LUONG download : add luanvanchat@agmail.com Declaration I declare that the thesis has been composed by myself and that the work has not be submitted for any other degree or professional qualification. I confirm that the work submitted is my own, except where work which has formed part of jointly- authored publications has been included.

My contribution and those of the other authors to this work have been ex- plicitly indicated below. I confirm that appropriate credit has been given within this thesis where reference has been made to the work of others. The work pre- sented in Chapter 3 was previously published in Proceedings of the 3rd ICMLSC as “QASA: Advanced Document Retriever for Open Domain Question Answering by Learning to Rank Question-Aware Self-Attentive Document Representations” by Trang M. Vu, Eng-Siong Chng.

This study was conceived by all of the authors. My contributions include: proposing the method, carrying out the experiments, and writing the paper. Master student Nguyen Minh Trang v LUAN VAN CHAT LUONG download : add luanvanchat@agmail.com Table of Contents Abstract. v Table of Contents.

viii List of Figures. x List of Tables .1 Open-domain Question Answering .2 Difficulties and Challenges .3 Objectives and Thesis Outline. 8 2 Background knowledge and Related work .1 Deep learning in Natural Language Processing .2 Long Short-Term Memory network .2 Employed Deep learning techniques .1 Rectified Linear Unit activation function .2 Mini-batch gradient descent .3 Adaptive Moment Estimation optimizer. 20 vi LUAN VAN CHAT LUONG download : add luanvanchat@agmail.3 Pairwise Learning to Rank approach.

24 3 Material and Methods .2 Question Encoding Layer .3 Document Encoding Layer .2 Training Process and Integrated System. 39 4 Experiments and Results .1 Tools and Environment. 50 List of Publications. 52 vii LUAN VAN CHAT LUONG download : add luanvanchat@agmail.com Acronyms Adam Adaptive Moment Estimation AoA Attention-over-Attention BiDAF Bi-directional Attention Flow BiLSTM Bi-directional Long Short-Term Memory CBOW Continuous Bag-Of-Words EL Embedding Layer EM Exact Match GA Gated-Attention IR Information Retrieval LSTM Long Short-Term Memory NLP Natural Language Processing QA Question Answering QASA Question-Aware Self-Attentive QEL Question Encoding Layer R3 Reinforced Ranker-Reader ReLU Rectified Linear Unit RNN Recurrent Neural Network viii LUAN VAN CHAT LUONG download : add luanvanchat@agmail.com SGD Stochastic Gradient Descent TF-IDF Term Frequency – Inverse Document Frequency TREC Text Retrieval Conference ix LUAN VAN CHAT LUONG download : add luanvanchat@agmail.com List of Figures 1.1 An overview of Open-domain Question Answering system.2 The pipeline architecture of an Open-domain QA system.3 The relationship among three related disciplines.4 The architecture of a simple feed-forward neural network.1 Embedding look-up mechanism.2 Recurrent Neural Network.3 Long short-term memory cell.4 Attention mechanism in the encoder-decoder architecture.5 The Rectified Linear Unit function.1 The architecture of the Document Retriever.2 The architecture of the Embedding Layer.1 Example of a question with its corresponding answer and contexts from QUASAR-T.2 Distribution of question genres (left) and answer entity-types (right).3 Top-1 accuracy on the validation dataset after each epoch.4 Loss diagram of the training dataset calculated after each epoch.

48 x LUAN VAN CHAT LUONG download : add luanvanchat@agmail.com List of Tables 1.1 An example of problems encountered by the Document Retriever.4 Evaluation of retriever models on the QUASAR-T test set.5 The overall performance of various open-domain QA systems. 49 xi LUAN VAN CHAT LUONG download : add luanvanchat@agmail.com Chapter 1 Introduction 1.1 Open-domain Question Answering We are living in the Information Age where many aspects of our lives are driven by information and technology. With the boom of the Internet few decades ago, there is now a colossal amount of data available and this number continues to grow exponentially. Obtaining all of these data is one thing, how to efficiently use and extract information from them is one of the most demanding requirements.

Generally, the activity of acquiring useful information from a data collection is called Information Retrieval (IR). A search engine, such as Google or Bing, is a type of IR. Search engines are extensively used that it is hard to imagine our lives today without them. Despite their applicability, current search engines and similar IR systems can only produce a list of relevant documents with respect to the user’s query.

To find the exact answer needed, users still have to manually examine these documents. Because of this, although IR systems have been handy, retrieving desirable information is still a time consuming process. The users can express their information needs in natural language instead of a series of keywords as in search engines. Furthermore, instead of a list of documents, QA systems try to return the most concise and coherent answers possible.

With the vast amount of data nowadays, QA systems can reduce count- less effort in retrieving information. Depending on usage, there are two types of QA: closed-domain and open-domain. Unlike closed-domain QA, which is re- 1 LUAN VAN CHAT LUONG download : add luanvanchat@agmail.com stricted to a certain domain and requires manually constructed knowledge bases, open-domain QA aims to answer questions about basically anything. Hence, it mostly relies on world knowledge in the form of large unstructured corpora, e.

Wikipedia, but databases are also used if needed.1 shows an overview of an open-domain QA system.1: An overview of Open-domain Question Answering system. The research about QA systems has a long history tracing back to the 1960s when Green et al. [20] first proposed BASEBALL. About a decade after that, Woods et al.

Both of these systems are closed-domain and they use manually defined language patterns to transform the questions into structured database queries. Since then, knowledge bases and closed-domain QA systems had become dominant [27]. They allow users to ask questions about cer- tain things but not all. Not until the beginning of this century that open-domain QA research has become popular with the launch of the annual Text Retrieval Conference (TREC) [44] started in 1999.

Ever since, TREC competitions, espe- cially the open-domain QA tracks, have progressed in size and complexity of the dataset provided, and evaluation strategies are improved. The attention is now shifting to open-domain QA and in recent years, the number of studies on the subject has increased exceedingly. 2 LUAN VAN CHAT LUONG download : add luanvanchat@agmail.1 Problem Statement In QA systems, the questions are natural language sentences and there are a many types of them based on their semantic categories such as factoid, list, causal, confirmation, hypothetical questions, etc. The most common ones that attract most studies in the literature are factoid questions which usually begin with Wh- interrogated words, i.

What, When, Where, Who [27]. With open-domain QA, the questions are not restricted to any particular domain but the users can ask whatever they want. Answers to these questions are facts and they can simply be expressed in text format. From an overview perspective, as presented in Figure 1.1, the input and out- put of an open-domain QA system are straightforward.

The input is the question, which is unrestricted, and the output is the answer, both are coherent natural lan- guage sentences and presented by text sequences. The system can use resources from the web or available databases. Any system like this can be considered as an open-domain QA system. However, open-domain QA is usually broken down into smaller sub-tasks since being able to give concise answers to any questions is not trivial.

Corresponding to each sub-task, there is a component dedicated to it. Typically, there are two sub-tasks: document retrieval and document com- prehension (or machine comprehension). Accordingly, open-domain QA systems customarily comprise of two modules: a Document Retriever and a Document Reader. Seemingly, the Document Retriever handles the document retrieval task and the Document Reader deals with the machine comprehension task.

The two modules can be integrated in a pipeline manner, e. [7, 46], to form a complete open-domain QA system. This architecture is depicted in Figure 1.2: The pipeline architecture of an Open-domain QA system. 3 LUAN VAN CHAT LUONG download : add luanvanchat@agmail.com The input of the system is still a question, namely q, and the output is an answer a.

Given q, the Document Retriever acquires top-k documents from a search space by ranking them based on their relevance to q. Since the require- ment for open-domain systems is that they should be able to answer any question, the hypothetical search space is massive as it must contains the world knowledge. However, an unlimited search space is not practical, so, knowledge sources like the Internet, or specifically Wikipidia, are commonly used. In the document re- trieval phase, a document is considered relevant to question q if it helps answer q correctly, meaning that it must at least contains the answer within its content.

Nevertheless, containing the answer alone is not enough because the document returned should also be comprehensible by the Reader and consistent with the se- mantic of the question. The relevance score is quantifiable by the Retriever so that all the documents can be ranked using it. Let D represent all documents in the search space, the set of top-k highest-scored documents is: ! D? = argmax Õ f (d, q) (1.1) X∈[D]k d∈X where f (·) is the scoring function.

Nội dung được bảo vệ bản quyền — Tải xuống đầy đủ