Xây Dựng Hệ Thống Hỏi Đáp Sử Dụng Mô Hình Sinh Sinh Sâu và Tăng Cường Truy Vấn

Tài liệu nghiên cứu Using retrieval augmentation and deep generative models to build question answering systems, tổng hợp lý thuyết và thực hành, cung cấp kiến thức chuyên sâu về .

Chuyên ngành

Computer Science

Người đăng

Ẩn danh

Thể loại

Thesis

2023

71
2
0

Phí lưu trữ

30 Point

Mục lục chi tiết

Declaration Of Authenticity

Acknowledgement

Abstract

1. Introduction

1.1. High computational complexity of deep language models

1.2. Cost of training deep language models

1.3. Limited amount of Vietnamese training data

1.4. Objective

1.4.1. Academic significance

1.4.2. Practical significance

1.5. Scope

1.6. Thesis Structure

2. Theoretical Background

2.1. The architecture of a Transformer model

2.1.1. The encoder

2.1.2. The decoder

2.2. Use

2.2.1. BARTpho

2.2.1.1. Architecture
2.2.1.2. Pre-training data
2.2.1.3. Pre-training objective
2.2.1.4. Use

2.2.2. mBERT

2.2.2.1. Architecture

3. Design, experiment and implementation process

4. Applying the results of this work into real-life applications

5. Evaluations and comparisons of the results of this work to other similar work

6. Conclusions, results of this thesis and future plans

List of Tables

List of Figures

Tóm tắt

I. Tổng Quan Về Hệ Thống Hỏi Đáp Bằng Mô Hình Sinh Sinh Sâu

Hệ thống hỏi đáp là một trong những ứng dụng quan trọng trong lĩnh vực công nghệ thông tin. Với sự phát triển của công nghệ AI, việc xây dựng hệ thống hỏi đáp bằng mô hình sinh sinh sâu đã trở thành một xu hướng nổi bật. Hệ thống này cho phép người dùng đặt câu hỏi và nhận được câu trả lời chính xác từ dữ liệu có sẵn. Việc áp dụng học sâu trong hệ thống này không chỉ giúp cải thiện độ chính xác mà còn tăng cường khả năng hiểu ngữ nghĩa của ngôn ngữ tự nhiên.

1.1. Khái Niệm Về Hệ Thống Hỏi Đáp

Hệ thống hỏi đáp là một ứng dụng cho phép người dùng tương tác và tìm kiếm thông tin. Chúng sử dụng các thuật toán xử lý ngôn ngữ tự nhiên để phân tích và hiểu câu hỏi của người dùng.

1.2. Vai Trò Của Mô Hình Sinh Sinh Sâu

Mô hình sinh sinh sâu giúp cải thiện khả năng sinh ra câu trả lời tự nhiên và chính xác hơn. Nó cho phép hệ thống học từ dữ liệu lớn và cải thiện khả năng hiểu ngữ nghĩa.

II. Thách Thức Trong Việc Xây Dựng Hệ Thống Hỏi Đáp

Mặc dù có nhiều lợi ích, việc xây dựng hệ thống hỏi đáp bằng mô hình sinh sinh sâu cũng gặp phải nhiều thách thức. Một trong những vấn đề lớn nhất là độ phức tạp tính toán của các mô hình này. Điều này có thể dẫn đến chi phí cao trong việc triển khai và bảo trì hệ thống.

2.1. Độ Phức Tạp Tính Toán Cao

Các mô hình sinh sinh sâu yêu cầu tài nguyên tính toán lớn, điều này có thể gây khó khăn cho các tổ chức nhỏ trong việc triển khai.

2.2. Thiếu Dữ Liệu Huấn Luyện Tiếng Việt

Sự thiếu hụt dữ liệu huấn luyện chất lượng cao cho tiếng Việt là một thách thức lớn. Điều này ảnh hưởng đến khả năng của hệ thống trong việc hiểu và trả lời câu hỏi chính xác.

III. Phương Pháp Xây Dựng Hệ Thống Hỏi Đáp Hiệu Quả

Để xây dựng một hệ thống hỏi đáp hiệu quả, cần áp dụng các phương pháp tiên tiến trong học sâuxử lý ngôn ngữ tự nhiên. Việc sử dụng tăng cường truy vấn có thể giúp cải thiện độ chính xác của câu trả lời.

3.1. Sử Dụng Mô Hình Tăng Cường Truy Vấn

Mô hình tăng cường truy vấn cho phép hệ thống tìm kiếm thông tin từ nhiều nguồn khác nhau, giúp cải thiện độ chính xác của câu trả lời.

3.2. Tạo Dữ Liệu Huấn Luyện Mới

Việc tạo ra các bộ dữ liệu hỏi đáp mới cho tiếng Việt sẽ giúp cải thiện khả năng của hệ thống trong việc xử lý và trả lời câu hỏi.

IV. Ứng Dụng Thực Tiễn Của Hệ Thống Hỏi Đáp

Hệ thống hỏi đáp có thể được áp dụng trong nhiều lĩnh vực khác nhau như giáo dục, y tế và dịch vụ khách hàng. Việc áp dụng công nghệ AI trong các hệ thống này giúp nâng cao trải nghiệm người dùng và tiết kiệm thời gian.

4.1. Hệ Thống Tư Vấn Học Tập

Hệ thống hỏi đáp có thể hỗ trợ sinh viên trong việc tìm kiếm thông tin học tập, giúp họ dễ dàng tiếp cận kiến thức.

4.2. Hệ Thống Hỗ Trợ Khách Hàng

Trong lĩnh vực dịch vụ khách hàng, hệ thống hỏi đáp giúp giảm thiểu thời gian chờ đợi và nâng cao sự hài lòng của khách hàng.

V. Kết Luận Và Tương Lai Của Hệ Thống Hỏi Đáp

Hệ thống hỏi đáp bằng mô hình sinh sinh sâu đang ngày càng trở nên phổ biến và có tiềm năng lớn trong tương lai. Việc cải thiện công nghệ và phát triển dữ liệu huấn luyện sẽ giúp nâng cao hiệu quả của các hệ thống này.

5.1. Tiềm Năng Phát Triển

Với sự phát triển không ngừng của công nghệ, hệ thống hỏi đáp có thể trở thành một phần quan trọng trong cuộc sống hàng ngày.

5.2. Hướng Nghiên Cứu Tương Lai

Nghiên cứu về cách cải thiện độ chính xác và khả năng hiểu ngữ nghĩa của hệ thống sẽ là một lĩnh vực quan trọng trong tương lai.

08/07/2025

Trích đoạn nội dung tài liệu

VIETNAM NATIONAL UNIVERSITY HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY FACULTY OF COMPUTER SCIENCE AND ENGINEERING —————————————– GRADUATION THESIS Using Retrieval Augmentation and Deep Generative Models to Build Question Answering Systems THESIS COMMITTEE: COMPUTER SCIENCE 3 SUPERVISOR: ASSOC. PHAM TRAN VU REVIEWER: DR. LE THANH VAN ——— STUDENT: NGUYEN KHAC HAO - 1852346 HO CHI MINH CITY, JANUARY 2023 Declaration Of Authenticity I declare that the research for this thesis is my own work, conducted under the su- pervision and guidance of Assoc. Pham Tran Vu.

The result of this research is legitimate and has not been published in any form prior to this. All materials used within this research are collected by myself by various sources and are appropriately listed in the references section. In any case of plagiarism, I stand by my actions and will be responsible for it. Ho Chi Minh City University of Technology therefore is not responsible for any copyright infringements conducted within my research.

i Acknowledgement I would like to thank my instructor, Assoc. Pham Tran Vu, not only for his academic guidance and assistance, but also for his patience and personal support which made me truly grateful. I also express my gratitude to Ho Chi Minh City University of Technology for giving me the opportunity to work on this research. Finally, I would like to express my deep sense of gratitude to the research teams and individuals working in computer science who made their incredible work publicly accessi- ble.

Open research papers, datasets and tools are the fundamental elements that helped made this thesis. ii Abstract The recent developments in information technology has given rise to a new generation of conversational applications. A large number of these applications are question answer- ing systems, where the user can ask an information-seeking question, and the application will reply with the corresponding information. Realizing the growth in this kind of application, I set out to build a general-purpose Vietnamese dialogue system that can be quickly adapted into any domain, reducing the amount of development work needed to build a question answering application.

In this work, I introduce a Vietnamese retrieval-augmented question answering system, and ex- plore ways to improve question answering accuracy in Vietnamese. I also introduce 5 new Vietnamese question answering datasets, created using machine translation and two question answering applications. iii Contents 1 Introduction 1 1.1 High computational complexity of deep language models .2 Cost of training deep language models .3 Limited amount of Vietnamese training data .3 Pre-training data .4 Pre-training objective .3 Pre-training data .4 Pre-training objective .3 Pre-training data .4 Pre-training objective .1 Architecture and pre-training .1 Reading comprehension question answering .2 Closed-book question answering .3 Closed-domain question answering .4 Open-domain question answering .1 Creating Vietnamese Question Answering Datasets .2 Reading Comprehension Model .1 Experimenting with fine-tuning data mixtures .2 Further analysis on multilingual fine-tuning .3 Combining best strategies .4 Model experiment 1: BARTpho VQA .5 Model experiment 2: mBERT VQA .6 Model experiment 3: XLM-R VQA .3 Retrieval-Augmentation System .1 Academic Counselling System .1 Vietnamese Question Answering Accuracy .2 English Question Answering Accuracy .3 Answer Retrieval Accuracy .1 Mean reciprocal rank .4 Real-life Testing. 53 6 Conclusion and Future Development 54 6.

55 List of Tables 1.1 Structure of this thesis.1 Detokenized and case-sensitive ROUGE evaluation score for BARTpho and mBART, in Vietnamese text summarization task.2 BLEU scores on PhoMT’s validation set.3 Primary differences between mT5 and other pre-trained multilingual lan- guage models.1 Statistics of the training set of the new Vietnamese question answering dataset. Extractive samples are samples where the answer is a span of text extracted from the document, while abstractive samples are samples where the answer is not an extracted text span from the document.2 Evaluation scores for each model in full-data multilingual training experiment.3 F1 scores (%) of Vi and En+Vi models, evaluated on different data splits.4 English question answering scores for BARTpho En and BARTpho En+Vi.1 Evaluation scores on Vietnamese question answering for baseline models, BARTpho VQA, mBERT VQA and mT5. Scores for mT5 models taken from [16].2 English question answering scores for XLM-R VQA and mT5. Scores for mT5 models taken from [16].3 Mean reciprocal rank scores for answer retrieval of the academic counselling system.4 Exact match scores for answer retrieval of the academic counselling system.

53 vi List of Figures 1.1 Example of question answering system: e-commerce chatbots.2 The linear growth in number of parameters of language models in recent years.1 The architecture of a Transformer model.2 The encoder of the Transformer model.3 The decoder of the Transformer model.4 F1 scores of best models for question answering to date, such as XL-Net [5] and LUKE [6], which are Transformer model.5 An illustration of BERT and mBERT during pre-training.6 Input representation of BERT and mBERT.7 Amount of data in GiB (log-scale) for the 88 languages that appear in both the Wiki-100 corpus used to pre-train mBERT, and the CommonCrawl-100 corpus used to pre-train XLM-R.8 Illustration of the MLM objective used to pre-train XLM-R and XLM.9 Recall@100 scores for unsupervised retrieval.10 An illustration comparing (a) black-box language models and (b) retrieval- augmented systems [17].11 RAG [20], an example implementation or retrieval-augmented model.12 Example of reading comprehension question answering input and output.13 Example of closed-book question answering input and output.1 A diagram of the steps performed during translation of English datasets into Vietnamese.2 Example of HTML tags and whitespaces removal during preprocessing.3 Example of answer span extraction on the SQuAD dataset during prepro- cessing.4 Example of neural machine translation of a preprocessed English data sample.5 Examples of the edits made during translation post-processing.6 Example of a document - question - answer data sample from the final dataset.7 Example of input and output in the multilingual training strategy.8 An example of input and output in the same-dataset multitask training strategy.9 An example of input and output in the multi-dataset multitask training strategy.10 F1 scores for each models trained with different data mixtures, after each training epoch.11 Examples where BARTpho En+Vi answered correctly while BARTpho Vi answered incorrectly.12 F1 scores of Vi and En+Vi models, evaluated on different data splits.13 F1 score of models trained with different data mixtures. NQA is short for NarrativeQA.14 A diagram of the answer span prediction head.15 A diagram of the answer extraction process in encoder-only question an- swering models.16 Architecture of the basic retrieval-augmented question answering system.17 A diagram of the relevance score calculation process in the Retriever.18 An example of the Generator’s input and output.1 Use-case diagram for the user and administrator of the Academic Coun- selling system.2 A diagram of the architecture of the academic counselling application.3 Format of an academic document in the Academic Counselling System.4 Example of Academic Counselling System answering questions by reading academic documents and generating an answer.5 Example of Academic Counselling System answering questions related to transfers.6 Example of Academic Counselling System answering questions related to courses and scheduling.7 Example of Academic Counselling System answering questions related to grading.8 Example of alternative answers in Academic Counselling System.1 Examples of the retrieval testing data for the Academic Counselling Sys- tem. The data is collected from real-life frequently asked questions.1 Problem Statement In the recent years, businesses and organizations have been increasingly using dialogue systems to communicate with their users. Some examples of these systems are e-commerce chatbots, counselling systems and health assistant.

Fundamentally, these are question- answering systems that work to find the right information to respond to the user’s queries. With the increasing use of these systems, there is an increasing need for a general-purpose question answering system that can be quickly adapted into any domain and application, saving development time.1: Example of question answering system: e-commerce chatbots. With deep language models, we can build more flexible, general question-answering systems. However, there is currently a number of problems and limitations that engineers need to overcome before they can bring deep question-answering systems into mainstream use.1 High computational complexity of deep language models Figure 1.2: The linear growth in number of parameters of language models in recent years.

Deep language model’s memory use and processing time increases very quickly with the input size. Language models such as T5 [1] has its memory use quadrupled when doubling the input sequence length. This limits developers from giving deep language models a large data input. For systems such as e-commerce chatbots, where accessing a large company database is required, deep language models become an unsuitable solution.2 Cost of training deep language models Depending on the size of the language model, training can be expensive or unfeasible.

Some implementations today use generative language models that answer user’s queries using knowledge stored within their parameters, learned during training time. When the information becomes outdated, the only way to update these models is to train them again. The cost of developing a language model from scratch and maintenance makes them unsuitable for real-life use.3 Limited amount of Vietnamese training data Deep language models requires a large amount of data for training and fine-tuning. Currently, there is not as many Vietnamese datasets for training language models as English.

For question answering, there is no large-scale, publicly-accessible Vietnamese dataset available. The limited amount of Vietnamese dataset available is one of the rea- sons traditional, specific-domain question answering systems is still currently the prefer- 2 able choice.2 Objective The objective of this work is to create a Vietnamese retrieval-augmented question answering system. This system can be quickly adapted to specific domains, reducing engineering work. When adapting this system to a new domain, no training is needed.

And when developers need to update the application’s data, they can update the data files without re-training the language model. This work is also a study on how to to apply deep language models in Vietnamese ap- plications. In this work, I explore with creating new dataset, new data mixture strategies and new Vietnamese question answering models.1 Academic significance • The new Vietnamese question answering datasets and question answering models introduced in this work can become useful resources for future studies in Vietnamese question answering. • The optimal data mixture experiment in this work can be useful for future studies.2 Practical significance • The retrieval-augmentation question answering system can be quickly modified to create applications for specific use, such as counselling systems and customer support systems.

This can help reduce development time. I demonstrated this significance by implementing two applications using this system. • Vietnamese question answering models and datasets introduced in this work can help other developments in bringing deep language models to real-life applications.4 Scope Currently, there are many types of dialogue system with different capabilities. There are question answering systems that can work with images, question answering systems that can do reasoning and calculations, and personal assistants that can accomplish tasks such as sending messages and e-mails.

The scope of this topic is limited to single-turn question answering systems that finds the information for the user. This system can work with information sources of any domain, but cannot do reasoning, calculations or accomplishing tasks other than question answering.5 Thesis Structure There are 6 chapters in this thesis: Chapter Content 1 Introduction to the objectives of this thesis Theoretical background as foundation knowledge of the algorithms, 2 models and techniques used in this work. The design, experiment and implementation process for the objectives 3 of this work. 4 Applying the results of this work into real-life applications.

Evaluations and comparisons of the results of this work to other similar 5 work. 6 Conclusions, results of this thesis and future plans.1: Structure of this thesis. 4 Chapter 2 Theoretical Background 2.1: The architecture of a Transformer model. The Transformer model [2] is a deep language model proposed in June 2017.

Trans- former adopts the mechanism of self-attention, learning from not only individual words but also the context of each word. The additional training parallelization allows training on larger datasets than was once possible. Since the introduction of the Transformer 5 model, many variations based on this model have been introduced.

Nội dung được bảo vệ bản quyền — Tải xuống đầy đủ

Tài liệu "Xây Dựng Hệ Thống Hỏi Đáp Bằng Mô Hình Sinh Sinh Sâu và Tăng Cường Truy Vấn" trình bày một phương pháp tiên tiến trong việc phát triển hệ thống hỏi đáp, sử dụng các mô hình sinh sinh sâu để cải thiện khả năng truy vấn và trả lời thông tin. Bài viết nhấn mạnh tầm quan trọng của việc áp dụng công nghệ trí tuệ nhân tạo trong việc tối ưu hóa trải nghiệm người dùng, giúp họ dễ dàng tìm kiếm và nhận được thông tin chính xác hơn.

Để mở rộng kiến thức của bạn về lĩnh vực này, bạn có thể tham khảo thêm tài liệu Khóa luận tốt nghiệp công nghệ thông tin xây dựng hệ thống hỏi đáp dựa trên đọc hiểu tự động cho tiếng việt, nơi bạn sẽ tìm thấy những ứng dụng cụ thể của công nghệ đọc hiểu tự động trong hệ thống hỏi đáp. Ngoài ra, tài liệu Khóa luận tốt nghiệp công nghệ thông tin xây dựng hệ thống hỏi đáp tiếng việt dựa trên các mô hình sinh ngôn ngữ sẽ cung cấp cho bạn cái nhìn sâu sắc về các mô hình sinh ngôn ngữ và cách chúng có thể được áp dụng trong việc phát triển hệ thống hỏi đáp. Cuối cùng, bạn cũng có thể tìm hiểu thêm về Khóa luận tốt nghiệp khoa học máy tính xây dựng hệ thống hỏi đáp về quy định đào tạo đại học, tài liệu này sẽ giúp bạn hiểu rõ hơn về cách xây dựng hệ thống hỏi đáp trong bối cảnh giáo dục đại học. Những tài liệu này sẽ giúp bạn mở rộng kiến thức và khám phá thêm nhiều khía cạnh thú vị trong lĩnh vực này.