Xây Dựng Đồ Thị Tri Thức Trong Hệ Thống Thông Tin

Luận văn tốt nghiệp nghiên cứu tốt nghiệp constructing knowledge graphs with triple extraction techniques, điều tra thực trạng, phân tích số liệu, đề xuất biện pháp cải tiến thực

Chuyên ngành

Information Systems

Người đăng

Ẩn danh

Thể loại

Graduation Thesis

2020

112
6
0

Phí lưu trữ

35 Point

Mục lục chi tiết

ACKNOWLEDGEMENTS

ABSTRACT

1. CHƯƠNG 1: CONTEXT AND PROBLEM STATEMENT

1.1. Context

1.2. The problem and its significance

1.3. Overview

2. CHƯƠNG 2: BACKGROUND AND THEORY

2.1. Triple extraction

2.2. Semantic role labeling

2.3. Knowledge graph construction

3. CHƯƠNG 3: SYSTEM DESIGN AND IMPLEMENTATION

4. CHƯƠNG 4: EVALUATION AND RESULTS

5. CHƯƠNG 5: CONCLUSION AND FUTURE WORK

APPENDIX A: EXAMPLES OF SENTENCE DECOMPOSER MODULE

APPENDIX B: AN EXAMPLE OF PROPERTY TRIPLES

APPENDIX C: EXAMPLES OF QUESTION GENERATION

LIST OF ABBREVIATIONS

Tóm tắt

I. Tổng Quan Về Xây Dựng Đồ Thị Tri Thức Trong Hệ Thống Thông Tin

Đồ thị tri thức là một công cụ mạnh mẽ trong việc tổ chức và quản lý thông tin. Nó cho phép lưu trữ và truy xuất thông tin một cách hiệu quả, giúp cải thiện khả năng tìm kiếm và phân tích dữ liệu. Trong bối cảnh hiện đại, việc xây dựng đồ thị tri thức trở nên cần thiết hơn bao giờ hết, đặc biệt trong các hệ thống thông tin lớn. Đồ thị tri thức không chỉ giúp tổ chức thông tin mà còn hỗ trợ trong việc phát triển các ứng dụng trí tuệ nhân tạo.

1.1. Định Nghĩa Đồ Thị Tri Thức Là Gì

Đồ thị tri thức là một cấu trúc dữ liệu mô tả các thực thể và mối quan hệ giữa chúng. Nó thường được biểu diễn dưới dạng các bộ ba (head, relation, tail), cho phép người dùng dễ dàng truy vấn và phân tích thông tin.

1.2. Vai Trò Của Đồ Thị Tri Thức Trong Hệ Thống Thông Tin

Đồ thị tri thức đóng vai trò quan trọng trong việc cải thiện khả năng tìm kiếm và phân tích dữ liệu. Nó giúp tổ chức thông tin một cách có hệ thống, từ đó nâng cao hiệu quả trong việc ra quyết định và phát triển ứng dụng.

II. Vấn Đề Trong Việc Xây Dựng Đồ Thị Tri Thức

Mặc dù đồ thị tri thức mang lại nhiều lợi ích, nhưng việc xây dựng chúng cũng gặp phải nhiều thách thức. Một trong những vấn đề lớn nhất là việc trích xuất thông tin từ dữ liệu không cấu trúc. Thông tin có thể bị mất mát trong quá trình này, dẫn đến việc xây dựng đồ thị không đầy đủ hoặc không chính xác.

2.1. Thách Thức Trong Việc Trích Xuất Thông Tin

Trích xuất thông tin từ dữ liệu không cấu trúc là một nhiệm vụ phức tạp. Các hệ thống hiện tại thường gặp khó khăn trong việc nhận diện và phân loại các thực thể, dẫn đến việc mất mát thông tin quan trọng.

2.2. Vấn Đề Về Độ Chính Xác Của Đồ Thị

Độ chính xác của đồ thị tri thức phụ thuộc vào chất lượng của dữ liệu đầu vào. Nếu dữ liệu không chính xác hoặc không đầy đủ, đồ thị tri thức sẽ không phản ánh đúng thực tế, gây khó khăn trong việc sử dụng.

III. Phương Pháp Xây Dựng Đồ Thị Tri Thức Hiệu Quả

Để xây dựng đồ thị tri thức hiệu quả, cần áp dụng các phương pháp hiện đại trong lĩnh vực xử lý ngôn ngữ tự nhiên và học máy. Việc kết hợp các kỹ thuật như phân tích vai trò ngữ nghĩa và nhận diện thực thể sẽ giúp cải thiện chất lượng của đồ thị.

3.1. Sử Dụng Phân Tích Vai Trò Ngữ Nghĩa

Phân tích vai trò ngữ nghĩa (Semantic Role Labeling) giúp xác định các vai trò của các thực thể trong câu, từ đó tạo ra các bộ ba chính xác hơn cho đồ thị tri thức.

3.2. Kết Hợp Với Nhận Diện Thực Thể

Nhận diện thực thể (Named Entity Recognition) là một kỹ thuật quan trọng giúp xác định và phân loại các thực thể trong văn bản, từ đó cải thiện độ chính xác của thông tin được trích xuất.

IV. Ứng Dụng Thực Tiễn Của Đồ Thị Tri Thức

Đồ thị tri thức có nhiều ứng dụng trong các lĩnh vực khác nhau như tìm kiếm thông tin, hệ thống gợi ý và phân tích dữ liệu. Chúng giúp cải thiện khả năng truy vấn và phân tích thông tin, từ đó nâng cao hiệu quả trong các ứng dụng thực tiễn.

4.1. Ứng Dụng Trong Tìm Kiếm Thông Tin

Đồ thị tri thức giúp cải thiện khả năng tìm kiếm thông tin bằng cách cung cấp các kết quả chính xác hơn và liên quan hơn đến yêu cầu của người dùng.

4.2. Hỗ Trợ Trong Phân Tích Dữ Liệu

Trong phân tích dữ liệu, đồ thị tri thức giúp tổ chức và trực quan hóa thông tin, từ đó hỗ trợ người dùng trong việc ra quyết định.

V. Kết Luận Về Tương Lai Của Đồ Thị Tri Thức

Tương lai của đồ thị tri thức hứa hẹn sẽ phát triển mạnh mẽ với sự tiến bộ của công nghệ. Việc áp dụng trí tuệ nhân tạo và học máy sẽ giúp cải thiện khả năng xây dựng và sử dụng đồ thị tri thức, mở ra nhiều cơ hội mới trong việc quản lý và phân tích thông tin.

5.1. Xu Hướng Phát Triển Công Nghệ

Công nghệ xây dựng đồ thị tri thức sẽ tiếp tục phát triển với sự hỗ trợ của các kỹ thuật học sâu và trí tuệ nhân tạo, giúp nâng cao độ chính xác và hiệu quả.

5.2. Tác Động Đến Các Ngành Công Nghiệp

Đồ thị tri thức sẽ có tác động lớn đến nhiều ngành công nghiệp, từ tài chính đến y tế, giúp cải thiện khả năng phân tích và ra quyết định.

10/07/2025
Khóa luận tốt nghiệp constructing knowledge graphs with triple extraction techniques

Trích đoạn nội dung tài liệu

VIETNAM NATIONAL UNIVERSITY - HO CHI MINH CITY UNIVERSITY OF INFORMATION TECHNOLOGY FACULTY OF INFORMATION SYSTEMS NGUYEN MINH HIEU GRADUATION THESIS BACHELOR OF ENGINEERING IN INFORMATION SYSTEMS HO CHI MINH CITY, 2020 VIETNAM NATIONAL UNIVERSITY - HO CHI MINH CITY UNIVERSITY OF INFORMATION TECHNOLOGY FACULTY OF INFORMATION SYSTEMS NGUYEN MINH HIEU — 16520399 GRADUATION THESIS CONSTRUCTING KNOWLEDGE GRAPHS WITH BACHELOR OF ENGINEERING IN INFORMATION SYSTEMS THESIS ADVISOR ASSOC. DO PHUC DR. NGO DUC THANH HO CHI MINH CITY, 2020 ASSESSMENT COMMITTEE The Assessment Committee is established under the Decision. by Rector of the University of Information Technology.

Quan Thanh Tho - Chairman 2 Dr. Duong Minh Duc - Secretary 3 Dr. Do Trong Hop - Member ACKNOWLEDGEMENTS I would like to express the deepest appreciation and sincere gratitude to my advisors Associate Professor Do Phuc and Doctor Ngo Duc Thanh for the valuable guidance and continuous support of my bachelor thesis. I have never seen such a more conscientious professor than Assoc.

Do Phuc who is always active to keep track with my progress and use his immense knowledge to carefully explain my questions. With precious advices and ideas from Dr. Ngo Duc Thanh, my ability to analyze and solve problems had a great enhancement. This dissertation could not have been successfully completed without their persistent help and concern.

Next, I wish to thank all members in Faculty of Information Systems as well as members of the University of Information Technology for their assistance to provide me best conditions to study and finish this thesis. I would also like to thank Mr. Hung Le for spending time on meetings with me to share his experience and suggestions. Besides, I am really grateful to Alex Sands and Ajay Patel, founders of Plasticity, who granted me free credits to use their system’s services.

They were extremely enthusiastic to support and help me deal with urgent cases during my system implementation process. Finally, I must express my very profound respect to my parents and my friends for encouragement throughout my years of study and during the time working on this thesis. December 2020 Nguyen Minh Hieu TABLE OF CONTENTS œ4ÍE]k› ,. i TABLE OF CONTENTS.

HH HH HH TH TH HH Thu nh HH HH nh TH rệt il IV. v TABLE OF TABLES100. ccceececcccccsesceeceeseeseeseeseesecaecaeeeseeseesecsecaececeseeseeaessesseseesneeeaeeseeaeens ix LIST OF ABBREVIATIONS. S11 ng HH nh nh nhiệt 1 1.2 The problem and its significance .- -G- 2G 3 191 vn TT HH ng nh nh nhện 2 1.2 FNgiU [si40v: 0s LucỪŨỪŨŨIDẬDAẠAẦỌDOAỘAẠAỌẬỆỄẲÊẦIẬNIẬIẤầẳầdỪỒỪ.3 Semantic role labeÏing.4 Knowledge graph COnSfTUCfIOH.

- G- G11 TH ng HH nh nh nành 6 1.7 Chapter SUMMary 1 8 CHAPTER 2. BACKGROUND AND THEORYY. - - - - - - - << << <315511111111 111K KĐT ng 22355555 10 PP Ji nh 5.3 Graph data models 117.1 Resource Description Framework.2 Labeled Property Graph .2 Neo4j graph dataasSe. - TT TH TH TH HH nh TH Hà HH Tnhh 17 P ¡hình .3 Named Entity RÑ€COBTIIOH.

32122111211 11 111111 119 11 11H11 TH TH HH Hiệp 22 2.- c2 St S111 v11 11 211211 11111 11 H1 HT TH TH TH ng HT ng 24 2.6 BERT — Bidirectional Encoder Representations from Transformers .7 BERT for semantic role labeÏing. SYSTEM DESIGN AND IMPLEMTATION.2 Extractor SVS{€TH.G- Q1 TH ke 40 SN» 200v.-- Q H H TH TH HH 45 3.3 Question-ansWer Pair Ø€T€TAẨOT.4 Question-answer pair S€aTC€T.-- TT TH nh TH HH Thu nu HH Hàn nh nh TH 57 3.3 Knowledge graph cOnStrUCfIOT. c1 11121113 1111 11121111911 9111 TH HH kkt 81 h9. CONCLUSION AND FUTURE WORK.- nh TT TH HH Thu TH HH HT TH HH ch TH 83 APPENDIX A: EXAMPLES OF SENTENCE DECOMPSOSER MODULE.

85 APPENDIX B: AN EXAMPLE OF PROPERTY TRIPLES .- ¿5555 5<<s<<ssss+ 87 APPENDIX C: EXAMPLES OF QUESTION GENERATION.ececccecceceesceseeseeseeseeeceeceececeesecsecaeceececeeaesaecaesaeseeceeeeseeaecaecaeseseeeeaeeaeeates 93 iv TABLE OF FIGURES ca Le Figure 2.1: Towards a Definition of Knowledge Graphs: An abstract KG architecture.2 An example about a knowledge graph.eccecceesceeseeseeseeeneeeeeeeeeseeeseenseeens II Figure 2.3 An RDF graph describing Joe Smith’s homepage 1nformafion.4 An RDF graph describing Joe Simi(Ì.- 6 S13 sinh re 13 Figure 2.5 Joe Smith information RDF graph syntax .6 A CD-list RDF SYTIfAX. - G1 TH nh nh TH TH Hà Hành nh nh nh 15 Figure 2.7 CD list RDF graphn.8 An example about LP.10 The result from read qU€FV.11 The result from update QU€TV.-- --- 5 2 33213 *2EE++EEE++eEEeeeeEeeeerereereeers 21 Figure 2.12 The result from delete QU€FY.13 An example using SpaCy to extract named entities.14 List of SpaCy named entity types 00. cece - - SH HH Hư 23 Figure 2.15 An example using AllenNLP coreference resolution for a paragraph.16 Some entities in discourse model form the paragraph in Figure 2.17 Frame file for the verb ““faÏÏ””.-- -- + c + 1112111111311 5115111111111 ke rke 27 Figure 2.18 BERT pre-training and Íine-tunn1ng.- - ---- - -s + -s + + ++se+seeeseereeeereeees 31 Figure 2.19 BERT input r€DT€S€TfAfIOTI.-- G- 2G c9 911911911 1 11 vn nh ngư 32 Figure 2.20 BERT-based relation extraction mođel.-- --: 5s + + + ++sx+s+eexssexssss 34 V Figure 2.21 Predicate argument identification and classification model .1 System OV€TVICW. LH HT TH TH T1 H11 H1 E1 H1 TH TH TH HH ng Hiệp 39 Figure 3.-- c0 3121131211112 1 1911118111101 1 101111011181 HH rệt 40 Figure 3.3 A sentence containing ClaUSeS.

--- 5 + tt 9 9v ng ng nh nnưy 41 Figure 3.4 A response from Sapien Language Eng1ne.5 Sapien graph object.- -ó- su TH HH Hà nh nh nh nh 43 Figure 3.7 A example about extracting named entities with SpaCy.8 A prediction of AllenNLP SRL prediCfOr.9 A JSON-formatted property fIDÏG.10 A sentence with mentions of the same coreference cluster .11 Graph insertion €SUÏÍ.12 The graph after mapping entities with coreference table.13 A pruned knowledge gøraph.14 An enriched €TIẨV.15 A subject-based question-answer pal .16 An object-based question-ansWer Pat .17 A modifier-based question-anSwer Pal .18 A question-answer pair searcher result .19 Visualization system compOn€nfS. - ¿+ sc + 3213 EEEeEsrresrsrrsree 58 Figure 3.20 Graph visualization: neo4jd3 Vs €OVIS.21 AllenNLP coreference resolution vIsual1ZatiOT.- -- sec xsssserss 60 Mi Figure 3.22 Visualization sySf€T4 OV€TVICW. Án 1H S1 SH 1T 1n HH rệt 61 Figure 3.23 Raw text input S€CfIOH.24 Knowledge graph VI€WT.25 Sentence decomposer VI€WCT.- ác TH nh TH Hà Hành nh nh nh 63 Figure 3.26 Coreference resolution VI€W€Y.27 KG stats: Relation distrIDufIOTI. --- «xxx x11 ng tt rưy 65 Figure 3.28 KG stats: Named entity distrIbutIOn.29 KG stats: Predicate modifier distrIbufIon.30 Question-answer pair table .31 Question-answer pair Searcher VICWCT.1 SRL-KG’s knowledge graph on example text 1 oo.

eee eeeeeeeeeeeceeeereeeeees 73 Figure 4.2 Team UWA’s knowledge graph on example text Ï.3 Modifiers of predicate “debuted”? 0. eee eecceccccescceseceeseeeeeeeeseeseaeeeeseeeseeenees 74 Figure 4.4 Team UWA’s knowledge graph on example text 2.5 SRL-KG’s knowledge graph on example text 2.6 Team UWA’s knowledge graph on example text 3.7 SRL-KG’s knowledge graph on example text 3.8 Modifiers of predicate ““T€fUTTS””.9 SRL-KG’s knowledge graph on example text 4.1 Knowledge graph about Fernando TOFT€S. - c6 6c 319v TT HH TH nh nh TH HH Hành ch nh ngờ 88 Figure 6800) xuan.4 Predicate vill TABLE OF TABLES œ4ÍE]k› 0020.2 PropBank annotations na.1 An example about coreference †abÏ€.1 Performance of SRL-KG on OIE-2016 (3200 senfences).2 Performance of SRL-KG on CaRB (641 sentences).- -- ¿5s ss<sssss2 71 Table 4.3 SRL-KG time perÍOrImaTnC.-- c2 +19 9E ng HH HH nh nh nành 79 Table 4.4 BLEU score on SQuAD of different question generation models. 81 1X LIST OF ABBREVIATIONS œ4ÍE]k› AI Artificial Intelligence CNN Convolutional Neural Network KG Knowledge Graph LSTM Long short-term memory NER Named Entity Recognition NLP Natural Language Processing RNN Recurrent Neural Network SRL Semantic Role Labeling ABSTRACT It is obvious that text data accounts for a great proportion of the total data available.

With the growth about the amount of data, deriving valuable information from text becomes more and more challenging. Knowledge graph is considered as a powerful mechanism to store the text data in an organized way in order to serve further applications in Artificial Intelligence. Most of popular knowledge graphs, e. Google Knowledge graph, DBpedia, Wikidata and YAGO, requires a large quantity of data as input and depends on human intervention processes on structural text data.

The main goal of this thesis is to design and implement an automatic knowledge graph construction pipeline from unstructured text. There have been several existed works on this topic, however information lost is one big problem that they have been still dealing with. Therefore, the center of attention in this thesis is to enhance the ability to make use of raw text in order to preserve the provided information as much as possible by using semantic role labeling approach cooperated other Natural Language Processing techniques like named entity recognition and coreference resolution. Moreover, to prove the usefulness of constructed knowledge graph from the proposed pipeline, there is additional tasks relating to question generation and question semantic searching.1 Context In recent years, a buzzword of industry 4.0 has been mentioned and appeared dominantly on social media with its promise to take the advantages of data to revolutionize manufacturing.

It can be said that data is the foundation of megatrends in information technology fields relating to Artificial Intelligence (AI) which can strongly support the industry. On the grounds that the number of online activities keeps arising, the amount of data has been increasing exponentially and becomes more various in nature. We obviously know that data may contain underlying information and knowledge. However, data seems useless if it is not processed properly.

Therefore, a big challenge is to extract value from collected raw data for further uses. In order to make use of information from text data, knowledge graphs appeared as a good way of knowledge representation. The term knowledge graph became obviously notable due to the introduction of Google’s Knowledge Graph in 2012 in which the basic motto is to focusing on searching things not strings [1] by utilizing knowledge graphs. Special concerns about this technology led to the fact that many other companies began to develop their own knowledge graphs such as: DBpedia [2], YAGO [3], Wikidata [4] or Freebase [5].

Due to knowledge graph’s powerful capability of knowledge representation and reasoning, it was used as the backbone of artificial intelligence to serve both academic and industrial purposes. To construct a knowledge graph, a set of triples (head, relation, tail) must be produced from text data. Therefore, triple extraction is one of the key basic steps. Until now, methods of triple extraction have been being developed with many different techniques which relate to statistical based and neural network based models.

However, it is supposed to be unclean and insufficient to instantly apply extracted triples through these methods for knowledge graph construction. As a result, the main purpose of this thesis is to make triples extracted from unstructured text more useful and valuable in order that high-quality knowledge graphs can be created from them to support advanced uses.2 The problem and its significance Triple extraction, a subset of information extraction, is an aspect in Natural Language Processing that takes an important role in knowledge graph construction. One of the most well-known information extraction system is Open Information System (OpenIE) [6] which was first introduce in 2007. After that, a wide range of OpenIE system has been developed with different approaches in which modern ones are associated with neural networks so as to enhance performance.

However, by doing experiments with some OpenlE systems such as: OpenIE-5 (including CALMIE [7], BONIE [8], RelNoun [9] and SRLIE [10]), MinIE [11], Supervised-OIE or RnnOIE [12] and IMoJIE [13], it can be found that those systems tend to produce many triples from a sentence to cover all kinds of results. Subjects or objects extracted from those techniques are regularly long sequences of words which includes semantic roles.

Nội dung được bảo vệ bản quyền — Tải xuống đầy đủ

Tài liệu có tiêu đề Xây Dựng Đồ Thị Tri Thức Trong Hệ Thống Thông Tin cung cấp cái nhìn sâu sắc về cách thức xây dựng và ứng dụng đồ thị tri thức trong các hệ thống thông tin hiện đại. Nội dung chính của tài liệu tập trung vào các phương pháp và kỹ thuật để tổ chức và quản lý thông tin, giúp người đọc hiểu rõ hơn về cách thức mà đồ thị tri thức có thể cải thiện khả năng truy xuất và phân tích dữ liệu.

Một trong những lợi ích lớn nhất mà tài liệu mang lại là khả năng giúp người đọc nắm bắt được các xu hướng mới trong lĩnh vực công nghệ thông tin, từ đó áp dụng vào thực tiễn công việc của mình. Để mở rộng thêm kiến thức, bạn có thể tham khảo tài liệu Luận án tiến sĩ khoa học máy tính dự đoán liên kết trên đồ thị tri thức sử dụng nhúng dịch chuyển và mạng tích chập, nơi trình bày các phương pháp tiên tiến trong việc dự đoán liên kết trên đồ thị tri thức. Ngoài ra, tài liệu Phương pháp xây dựng đồ thị tri thức theo miền dựa trên nguồn dữ liệu từ wikipedia sẽ giúp bạn hiểu rõ hơn về cách khai thác dữ liệu từ các nguồn mở để xây dựng đồ thị tri thức hiệu quả. Những tài liệu này không chỉ cung cấp thông tin bổ ích mà còn mở ra nhiều cơ hội để bạn khám phá sâu hơn về lĩnh vực này.