Hệ Thống Hỏi Đáp Dựa Trên Đồ Thị Tri Thức COVID-19

Khóa luận tốt nghiệp nghiên cứu hệ thống thông tin hỏi đáp dựa trên đồ thị tri thức về COVID-19, cung cấp giải pháp thông tin hiệu quả.

Người đăng

Ẩn danh

Thể loại

thesis

2021

113
3
0

Phí lưu trữ

35 Point

Mục lục chi tiết

ACKNOWLEDGMENTS

ABSTRACTS

TABLE OF CONTENTS

1. CHAPTER 1: INTRODUCTION

1.1. Research questions and targets

1.2. Targets

1.3. Problem Statement

2. CHAPTER 2: BACKGROUND AND RELATED WORK

2.1. Semantic web technologies

2.2. Web Ontology Language (OWL)

2.3. Discrimination of knowledge graph

2.4. Question Answering over Knowledge Graph

3. CHAPTER 3: TOWARDS QUESTION ANSWERING OVER KNOWLEDGE GRAPH BENCHMARK CORPORA

3.1. Question Type Classification

3.2. GROUP BY-HAVING

4. CHAPTER 4: VIETNAMESE DBPEDIA KNOWLEDGE GRAPH

4.1. Vietnamese DBpedia generation

4.2. Raw Infobox Extraction

4.3. Mapping-Based Infobox Extraction

4.4. Relations and Predicates

5. CHAPTER 5: QUESTION ANSWERING OVER KNOWLEDGE GRAPH METHODOLOGY

5.1. Subgraph matching techniques

5.2. Template-based technique

5.3. Information retrieval-based technique

6. CHAPTER 6: EXPERIMENTS AND RESULTS

6.1. End-to-end Comparison

6.2. Influence of Question Taxonomy

6.3. Effects of Quantity of Question

6.4. The trade-off between English and Vietnamese

6.5. Summary

7. CHAPTER 7: QUESTION ANSWERING APPLICATION FOR QUESTION ANSWERING SYSTEMS OVER KNOWLEDGE GRAPH

7.1. Description of COVID-KGQA application

8. CHAPTER 8: CONCLUSION

8.1. Summary of the work

8.2. Future Directions

REFERENCES

LIST OF ABBREVIATIONS

LIST OF TABLES

LIST OF FIGURES

Tóm tắt

I. Tổng quan về Hệ Thống Hỏi Đáp Dựa Trên Đồ Thị Tri Thức COVID 19

Hệ thống hỏi đáp dựa trên đồ thị tri thức COVID-19 là một công nghệ tiên tiến giúp người dùng truy cập thông tin chính xác và nhanh chóng về đại dịch. Công nghệ này sử dụng các đồ thị tri thức để tổ chức và phân tích dữ liệu liên quan đến COVID-19, từ đó cung cấp câu trả lời cho các câu hỏi tự nhiên. Việc áp dụng hệ thống này không chỉ giúp nâng cao hiệu quả trong việc tìm kiếm thông tin mà còn hỗ trợ trong việc ra quyết định trong bối cảnh khủng hoảng sức khỏe toàn cầu.

1.1. Định nghĩa và vai trò của hệ thống hỏi đáp

Hệ thống hỏi đáp (QA) là công cụ cho phép người dùng đặt câu hỏi và nhận câu trả lời từ cơ sở dữ liệu. Đối với COVID-19, hệ thống này giúp cung cấp thông tin chính xác về virus, vaccine và các biện pháp phòng ngừa.

1.2. Lợi ích của việc sử dụng đồ thị tri thức

Đồ thị tri thức giúp tổ chức thông tin một cách có cấu trúc, cho phép truy xuất dữ liệu nhanh chóng và hiệu quả. Điều này đặc biệt quan trọng trong bối cảnh thông tin về COVID-19 liên tục thay đổi.

II. Thách thức trong việc phát triển hệ thống hỏi đáp COVID 19

Mặc dù hệ thống hỏi đáp dựa trên đồ thị tri thức mang lại nhiều lợi ích, nhưng vẫn tồn tại nhiều thách thức trong quá trình phát triển. Các vấn đề như độ chính xác của dữ liệu, khả năng hiểu ngữ nghĩa của câu hỏi và tốc độ phản hồi là những yếu tố cần được cải thiện. Đặc biệt, việc xử lý các câu hỏi phức tạp liên quan đến COVID-19 đòi hỏi hệ thống phải có khả năng phân tích và suy luận tốt.

2.1. Độ chính xác của dữ liệu

Độ chính xác của dữ liệu là yếu tố quan trọng nhất trong hệ thống hỏi đáp. Dữ liệu không chính xác có thể dẫn đến thông tin sai lệch, ảnh hưởng đến quyết định của người dùng.

2.2. Khả năng hiểu ngữ nghĩa

Hệ thống cần có khả năng hiểu ngữ nghĩa của câu hỏi để cung cấp câu trả lời phù hợp. Điều này đòi hỏi các thuật toán xử lý ngôn ngữ tự nhiên phải được cải tiến liên tục.

III. Phương pháp phát triển hệ thống hỏi đáp COVID 19 hiệu quả

Để phát triển một hệ thống hỏi đáp hiệu quả, cần áp dụng các phương pháp tiên tiến trong lĩnh vực trí tuệ nhân tạo và học máy. Việc xây dựng một cơ sở dữ liệu đồ thị tri thức phong phú và đa dạng là rất quan trọng. Ngoài ra, việc sử dụng các mô hình học sâu để cải thiện khả năng hiểu ngữ nghĩa cũng là một yếu tố then chốt.

3.1. Xây dựng cơ sở dữ liệu đồ thị tri thức

Cơ sở dữ liệu đồ thị tri thức cần được xây dựng từ nhiều nguồn thông tin khác nhau, bao gồm dữ liệu từ các tổ chức y tế, nghiên cứu khoa học và thông tin từ cộng đồng.

3.2. Ứng dụng mô hình học sâu

Mô hình học sâu có thể giúp cải thiện khả năng phân tích ngữ nghĩa và cung cấp câu trả lời chính xác hơn cho các câu hỏi phức tạp.

IV. Ứng dụng thực tiễn của hệ thống hỏi đáp COVID 19

Hệ thống hỏi đáp dựa trên đồ thị tri thức COVID-19 đã được áp dụng trong nhiều lĩnh vực khác nhau, từ giáo dục đến y tế. Các ứng dụng này không chỉ giúp người dùng tìm kiếm thông tin mà còn hỗ trợ các chuyên gia trong việc ra quyết định. Việc sử dụng hệ thống này đã chứng minh được hiệu quả trong việc nâng cao nhận thức cộng đồng về COVID-19.

4.1. Hỗ trợ giáo dục

Hệ thống hỏi đáp có thể được sử dụng trong các khóa học trực tuyến để cung cấp thông tin chính xác về COVID-19 cho sinh viên và giảng viên.

4.2. Hỗ trợ ra quyết định trong y tế

Các chuyên gia y tế có thể sử dụng hệ thống để truy xuất thông tin nhanh chóng, từ đó đưa ra các quyết định kịp thời trong công tác phòng chống dịch.

V. Kết luận và tương lai của hệ thống hỏi đáp COVID 19

Hệ thống hỏi đáp dựa trên đồ thị tri thức COVID-19 đã cho thấy tiềm năng lớn trong việc cung cấp thông tin chính xác và kịp thời. Tương lai của hệ thống này sẽ phụ thuộc vào việc cải tiến công nghệ và mở rộng cơ sở dữ liệu. Việc tích hợp các công nghệ mới như trí tuệ nhân tạo và học máy sẽ giúp nâng cao hiệu quả của hệ thống trong việc phục vụ cộng đồng.

5.1. Tiềm năng phát triển

Với sự phát triển không ngừng của công nghệ, hệ thống hỏi đáp có thể mở rộng ra nhiều lĩnh vực khác nhau, không chỉ giới hạn trong COVID-19.

5.2. Tích hợp công nghệ mới

Việc tích hợp các công nghệ mới sẽ giúp cải thiện khả năng phản hồi và độ chính xác của hệ thống, từ đó phục vụ tốt hơn cho người dùng.

10/07/2025
Khóa luận tốt nghiệp hệ thống thông tin question answering over knowledge graphs for covid 19

Trích đoạn nội dung tài liệu

VIETNAM NATIONAL UNIVERSITY HO CHI MINH CITY UNIVERSITY OF INFORMATION TECHNOLOGY FACULTY OF INFORMATION SYSTEMS KEL QUESTION ANSWERING OVER KNOWLEDGE GRAPHS FOR COVID-19 BACHELOR OF ENGINEERING IN INFORMATION SYSTEMS HO CHI MINH CITY- 2021 VIETNAM NATIONAL UNIVERSITY HO CHI MINH CITY UNIVERSITY OF INFORMATION TECHNOLOGY FACULTY OF INFORMATION SYSTEMS QUESTION ANSWERING OVER KNOWLEDGE GRAPHS FOR COVID-19 BACHELOR OF ENGINEERING IN INFORMATION SYSTEMS THESIS SUPERVISOR Dr. NGUYEN LUU THUY NGAN HO CHI MINH CITY- 2021 ACKNOWLEDGMENTS I would like to show my thanks to a large number of persons who graciously assisted me in completing the work contained in this thesis as well as who supported me during my university studies. My heartfelt thanks to my thesis supervisor, Dr. Nguyén Luu Thuy Ngan at UIT, VNU-HCM.

Through her valuable instruction and offer of independence, I gained vital skills for future study and steered my research in the appropriate path. Her assistance is unmatched and has aided me in through the most difficult period of my life. Without her, I would not have been able to complete my thesis and choose the best course of action for my future career. I appreciate your patience and tolerance with me.

My heartfelt thanks to my thesis advisor, Dr. Nguyén Thanh Tam from Griffith University for sharing a portion of his work at the Leibinz AI Lab at Leibniz Uni- versity Hannover and for their insightful comments and conversations that helped me better my dissertation. I appreciate your patience and effort spent organizing my university thesis. My wholehearted appreciation to thesis committee, Dr.

Ngô Đức Thanh at UIT, VNU-HCM for his feedback and opinions that helped me improve my dissertation. Following that, I would like to express my gratitude to all members of the Faculty of Information Systems and the University of Information Technology for assisting me in obtaining the optimal circumstances for studying and completing this thesis. It was a joyous time in my life to work with my friends, brothers, and lecturers at The UIT Natural Language Processing Group, and I am eternally grateful for their wonderful support and companionship. I apologize for being unable to complete all activities due to work and time restrictions.

I’d also want to express my gratitude to the outstanding educator, colleague, and co-author with whom I had the honor of working throughout my undergraduate studies. Many thanks Msc. Nguyén Van Kiét, lecturer I would be unable to continue my research work without your great thoughts. You are one of the most wonderful guidance I have ever received, and I want to express my gratitude to my lab-mates and friends at The UIT Natural Language Processing Group during my time at UIT, who make my working days so enjoyable and enjoyable through various organizations.

Ứd want to convey my heartfelt appreciation to my family and close friends for checking in on me on a regular basis to see how I was doing in a place so far away. They regularly expressed concern for my career and personal life, despite the fact that I was unable to spend much time with them during this hectic period. When we connected, I always felt both warm and strong, as they expressed their immense trust in me and the necessity of my job. Without their unwavering love and support, this dissertation and this remarkable journey would not have been possible.

I am really appreciative to my parents. They will always be there for me, to support and to love me, regardless of my accomplishments or failures. ii ABSTRACTS In order to gather and organize factual information, several large-scale knowl- edge graphs have been established. The efficient retrieval of meaningful informa- tion from knowledge bases has become an increasingly essential issue as the scope of KGs has grown.

It has become commonplace to employ Knowledge Graph Ques- tion Answering (KGQA), a technology that automatically responds to natural lan- guage questions (NLQ) made by users, to better comprehend users’ intents and ex- tract relevant knowledge from knowledge graph (KGs). For the purpose of making the KGQA more practical in real-world situations, researchers have shifted their fo- cus away from simple questions and toward complicated questions that better fulfill users’ increasingly sophisticated requirements. Compound questions, as opposed to simple questions—which are inquiries that relate to a single fact contained in a knowledge graph—frequently need the use of logical, quantitative, and compara- tive reasoning across several knowledge graph triples. KGQA has been a research hot-spot in recent years, as a result of this increased interest in the subject.

The primary purpose of this thesis is to design and develop multilingual ques- tion answering over the knowledge graph, for use in the COVID-19 construction pipeline design and construction. There have been a number of previous papers on this subject, but none of them had COVID-19 information and multilingual property. As a result, the goal of this research is to provide a multilingual KGQA dataset for evaluating the efficacy of KGQA systems in translating articulated enquiries into a specific data query language. This dataset was also used to evaluate four state-of- the-art KGQA systems, one of which was used our proposed approach and get better results in some benchmark datasets.

Furthermore, in order to aid in Vietnamese KGQA, we are developing a Vietnamese DBpedia-based KG a from unstructured text, as well as to assist with Vietnamese KGQA assessment. 11 TABLE OF CONTENTS ACKNOWLEDGMENTS i ABSTRACTS iii TABLE OF CONTENTS iv ABBREVIATIONS ix LIST OF TABLES Xx LIST OF FIGURES xi Chapter 1 INTRODUCTION 1 1. Research questions and targets.6 Thesis OrØganIsatiOn. 00002 ee 7 Chapter 2 BACKGROUND AND RELATED WORK 8 2.02 2 eee ee ee 8 1V TABLE OF CONTENTS 2.2 Semantic web technologies .3 Web Ontology Language (OWL) .2 Discrimination of knowledge graph .2 eee ee eee 16 2.2 Question Answering over Knowledge Graph.

ee 20 Chapter 3 TOWARDS QUESTION ANSWERING OVER KNOWLEDGE GRAPH BENCHMARK CORPORA 21 3.2 Question Type Classification .00 00085 28 TABLE OF CONTENTS 3.8 GROUP BY-HAVING. Q Q Q Q Q Q HH ee eee 29 3. ee ee 32 Chapter 4 VIETNAMESE DBPEDIA KNOWLEDGE GRAPH 35 4.2 Vietnamese DBpedia generation .1 Raw Infobox Extraction .2 Mapping-Based Infobox Extraction. ee ee es 49 4.3 Relations and Predicates.

Q Q Q Q Q HQ HH ee ee 51 Chapter 5 QUESTION ANSWERING OVER KNOWLEDGE GRAPH METHODOLOGY 52 5.2 Subgraph matching techniques. 52 VI TABLE OF CONTENTS “Am.3 Template-based technque.4 Information retrieval-based technque. Ặ Q0 SH HH va 68 Chapter 6 EXPERIMENTS AND RESULTS 69 6. eee eee eee 69 6.1 End-to-end Comparison.3 Influence of Question Taxonomy .4 Effects of Quantity ofQuestion.5 The trade-off between English and Vietnamese.

HQ ee va 79 65 SUmMATV. ee ee 80 Chapter 7 QUESTION ANSWERING APPLICATION FOR QUESTION ANSWERING SYSTEMS OVER KNOWLEDGE GRAPH 81 7.2 Description of COVID-KGQA application.ee 85 Chapter 8 CONCLUSION 86 8.1 Summary of the work .0 00000002 ee eee 86 Vil TABLE OF CONTENTS 8.2 Future Directions REFERENCES Vili LIST OF ABBREVIATIONS QA = Question Answering KB Knowledge Base KG Knowledge Graph KBQA Question Answering over Knowledge Base KGQA Question Answering over Knowledge Graph SOTA State-of-the-art COVID-19 Coronavirus Disease of 2019 NLQ Natural Language Question OWL Web Ontology Language RDF Resource Description Framework 1X LIST OF TABLES 1.1 Example of Question Answering over Knowledge Graphs .1 List of datasets 2.1 Comparison between DBpedia knowledge graph .2 Overview information about ViDBpedia and DBpedia .1 General comparisons between KBQA techniques.1 Overall Performance of QA Systems.2 Evaluation of QA Systems on Precision .3 Evaluation of QA Systems on Recall.4 Evaluation of QA SystemsonFl.5 Evaluation of QA Systems on Time(s).6 Comparison between English and Vietnamese in COVID-KGQA. 78 LIST OF FIGURES 1.1 Question answering over knowledge graphs.1 Resource Description Framework (RDF) triple example.2 Web Ontology Language (OWL) example!. Q Q eee ee ee 13 2.4 A diagram of composition of three members of KG .5 Data sources that are interlinked with DBpedia [32].1 Data construction pfoCess.2 Example of our dataset .3 Distribution of the most popular question prefixes in QALD-9.4 Distribution of the most popular question prefixes in LC-QUAD.5 Distribution of the most popular question prefixes in COVID-KGQA 34 4.1 Statistic of Wikipedia pages from November 2018 to November 202127 38 4.2 Overview of Vietnamese DBpedia generation architecture .3 The AST representation of the image link [72].4 Wikipedia information about COVID-19 pandemic in Vietnam? .5 Comparison between ViDBpedia and existing default Vietnamese in DBpedia.6 Comparison between ViDBpedia, existing default Vietnamese and Total Vietnamese in DBpedia .1 Overview of gAnswermodel.2 Overview of semantic relation generation framework of gAnswer in question understading step[79].3 Overview of TeBaQA model [82] .4 Overview of QAspardglmodel.

62 XI LIST OF FIGURES 6.1 Comparison 4 systems on Test set of 5corpora .2 End to end time comparison on Testset. Comparison FI on COVID-KGQA with influence of question type from test set, the underscore on the charts represent valueQO .4 Comparison Time on COVID-KGQA with influence of question type from testset.5 Comparison F1 on LCQUAD with influence of question type from test set, the underscore on the charts represent valueO .6 Comparison Time on LCQUAD with influence of question number fromtestset 2.7 Comparison on COVID-KGQA with influence of question number from test set.8 Fl and Time comparison on LCQUAD with influence of question quantity fromtestset 2. cee eee ee eee 78 7.1 System architecture of COVID-KGQA QA application. Screenshot of QA systems .3 Screenshot of QA systems (Com).1 Context Question Answering (QA) is a long-standing discipline within the field of natural language processing (NLP), which is concerned with providing answers to questions posed in natural language on data sources, and also draws on techniques from lin- guistics, database processing, and information retrieval [1].

One important type of data sources is knowledge bases, also known as knowledge graphs, which have been automatically constructed from web data and have become a key asset for search engines and many applications [2]. The number of knowledge graphs (KGs) has increased at an unprecedented rate over the past 15 years [3, 4, 5]. These KGs include a wealth of information that may possibly be utilized for QA. Finding answers for a question in a KG, on the other hand, is not always straightforward.

The user needs to have a thorough understand- ing of the KG as well as a structured query language in order to express their queries in a structured manner that can be utilized to locate matches in the KG. In order to address this issue, a significant number of QA systems that allow users to express their information requirements using natural language have been created. According to [6], factoid question answering has two main approaches, information retrieval (IR) based QA and knowledge-based QA. In the first approach, many datasets and systems have been established recent years with the advancement of machine reading comprehension (MRC) task.

One of the most famous datasets is SQuAD [7] conducted in English. MRC is also become famous in multilingualism such as Vietnamese [8]. The second approach is question answering over knowledge base (KBQA) with precision of question answering over knowledge graph (KG). Many studies have been conducted with KB such as [9, 10] and recently is WabiQA [11], which is KBQA system in Thai language.

Conducting question answering over a knowledge base (KB) implies that one CHAPTER 1. INTRODUCTION wishes to locate an answer to a query within the KB. Consider the following hy- pothetical situation: "Where is the first case of COVID-19 in Vietnam?" In this case, we want to extract the precise information contained within the triple depicted above. This will include the answer you seek.

The primary challenge in question answering over Knowledge Bases is develop- ing an algorithm capable of automatically searching through a collection of triples and determining the answer to a question. This is the type of QA that is discussed in this thesis. Despite the fact that the primary issue exists, there are several re- lated issues.

Nội dung được bảo vệ bản quyền — Tải xuống đầy đủ