VIETNAM NATIONAL UNIVERSITY - HO CHI MINH CITY UNIVERSITY OF INFORMATION TECHNOLOGY FACULTY OF INFORMATION SYSTEMS NGUYEN MINH HIEU GRADUATION THESIS BACHELOR OF ENGINEERING IN INFORMATION SYSTEMS HO CHI MINH CITY, 2020 VIETNAM NATIONAL UNIVERSITY - HO CHI MINH CITY UNIVERSITY OF INFORMATION TECHNOLOGY FACULTY OF INFORMATION SYSTEMS NGUYEN MINH HIEU — 16520399 GRADUATION THESIS CONSTRUCTING KNOWLEDGE GRAPHS WITH BACHELOR OF ENGINEERING IN INFORMATION SYSTEMS THESIS ADVISOR ASSOC. DO PHUC DR. NGO DUC THANH HO CHI MINH CITY, 2020 ASSESSMENT COMMITTEE The Assessment Committee is established under the Decision. by Rector of the University of Information Technology.
Quan Thanh Tho - Chairman 2 Dr. Duong Minh Duc - Secretary 3 Dr. Do Trong Hop - Member ACKNOWLEDGEMENTS I would like to express the deepest appreciation and sincere gratitude to my advisors Associate Professor Do Phuc and Doctor Ngo Duc Thanh for the valuable guidance and continuous support of my bachelor thesis. I have never seen such a more conscientious professor than Assoc.
Do Phuc who is always active to keep track with my progress and use his immense knowledge to carefully explain my questions. With precious advices and ideas from Dr. Ngo Duc Thanh, my ability to analyze and solve problems had a great enhancement. This dissertation could not have been successfully completed without their persistent help and concern.
Next, I wish to thank all members in Faculty of Information Systems as well as members of the University of Information Technology for their assistance to provide me best conditions to study and finish this thesis. I would also like to thank Mr. Hung Le for spending time on meetings with me to share his experience and suggestions. Besides, I am really grateful to Alex Sands and Ajay Patel, founders of Plasticity, who granted me free credits to use their system’s services.
They were extremely enthusiastic to support and help me deal with urgent cases during my system implementation process. Finally, I must express my very profound respect to my parents and my friends for encouragement throughout my years of study and during the time working on this thesis. December 2020 Nguyen Minh Hieu TABLE OF CONTENTS œ4ÍE]k› ,. i TABLE OF CONTENTS.
HH HH HH TH TH HH Thu nh HH HH nh TH rệt il IV. v TABLE OF TABLES100. ccceececcccccsesceeceeseeseeseeseesecaecaeeeseeseesecsecaececeseeseeaessesseseesneeeaeeseeaeens ix LIST OF ABBREVIATIONS. S11 ng HH nh nh nhiệt 1 1.2 The problem and its significance .- -G- 2G 3 191 vn TT HH ng nh nh nhện 2 1.2 FNgiU [si40v: 0s LucỪŨỪŨŨIDẬDAẠAẦỌDOAỘAẠAỌẬỆỄẲÊẦIẬNIẬIẤầẳầdỪỒỪ.3 Semantic role labeÏing.4 Knowledge graph COnSfTUCfIOH.
- G- G11 TH ng HH nh nh nành 6 1.7 Chapter SUMMary 1 8 CHAPTER 2. BACKGROUND AND THEORYY. - - - - - - - << << <315511111111 111K KĐT ng 22355555 10 PP Ji nh 5.3 Graph data models 117.1 Resource Description Framework.2 Labeled Property Graph .2 Neo4j graph dataasSe. - TT TH TH TH HH nh TH Hà HH Tnhh 17 P ¡hình .3 Named Entity RÑ€COBTIIOH.
32122111211 11 111111 119 11 11H11 TH TH HH Hiệp 22 2.- c2 St S111 v11 11 211211 11111 11 H1 HT TH TH TH ng HT ng 24 2.6 BERT — Bidirectional Encoder Representations from Transformers .7 BERT for semantic role labeÏing. SYSTEM DESIGN AND IMPLEMTATION.2 Extractor SVS{€TH.G- Q1 TH ke 40 SN» 200v.-- Q H H TH TH HH 45 3.3 Question-ansWer Pair Ø€T€TAẨOT.4 Question-answer pair S€aTC€T.-- TT TH nh TH HH Thu nu HH Hàn nh nh TH 57 3.3 Knowledge graph cOnStrUCfIOT. c1 11121113 1111 11121111911 9111 TH HH kkt 81 h9. CONCLUSION AND FUTURE WORK.- nh TT TH HH Thu TH HH HT TH HH ch TH 83 APPENDIX A: EXAMPLES OF SENTENCE DECOMPSOSER MODULE.
85 APPENDIX B: AN EXAMPLE OF PROPERTY TRIPLES .- ¿5555 5<<s<<ssss+ 87 APPENDIX C: EXAMPLES OF QUESTION GENERATION.ececccecceceesceseeseeseeseeeceeceececeesecsecaeceececeeaesaecaesaeseeceeeeseeaecaecaeseseeeeaeeaeeates 93 iv TABLE OF FIGURES ca Le Figure 2.1: Towards a Definition of Knowledge Graphs: An abstract KG architecture.2 An example about a knowledge graph.eccecceesceeseeseeseeeneeeeeeeeeseeeseenseeens II Figure 2.3 An RDF graph describing Joe Smith’s homepage 1nformafion.4 An RDF graph describing Joe Simi(Ì.- 6 S13 sinh re 13 Figure 2.5 Joe Smith information RDF graph syntax .6 A CD-list RDF SYTIfAX. - G1 TH nh nh TH TH Hà Hành nh nh nh 15 Figure 2.7 CD list RDF graphn.8 An example about LP.10 The result from read qU€FV.11 The result from update QU€TV.-- --- 5 2 33213 *2EE++EEE++eEEeeeeEeeeerereereeers 21 Figure 2.12 The result from delete QU€FY.13 An example using SpaCy to extract named entities.14 List of SpaCy named entity types 00. cece - - SH HH Hư 23 Figure 2.15 An example using AllenNLP coreference resolution for a paragraph.16 Some entities in discourse model form the paragraph in Figure 2.17 Frame file for the verb ““faÏÏ””.-- -- + c + 1112111111311 5115111111111 ke rke 27 Figure 2.18 BERT pre-training and Íine-tunn1ng.- - ---- - -s + -s + + ++se+seeeseereeeereeees 31 Figure 2.19 BERT input r€DT€S€TfAfIOTI.-- G- 2G c9 911911911 1 11 vn nh ngư 32 Figure 2.20 BERT-based relation extraction mođel.-- --: 5s + + + ++sx+s+eexssexssss 34 V Figure 2.21 Predicate argument identification and classification model .1 System OV€TVICW. LH HT TH TH T1 H11 H1 E1 H1 TH TH TH HH ng Hiệp 39 Figure 3.-- c0 3121131211112 1 1911118111101 1 101111011181 HH rệt 40 Figure 3.3 A sentence containing ClaUSeS.
--- 5 + tt 9 9v ng ng nh nnưy 41 Figure 3.4 A response from Sapien Language Eng1ne.5 Sapien graph object.- -ó- su TH HH Hà nh nh nh nh 43 Figure 3.7 A example about extracting named entities with SpaCy.8 A prediction of AllenNLP SRL prediCfOr.9 A JSON-formatted property fIDÏG.10 A sentence with mentions of the same coreference cluster .11 Graph insertion €SUÏÍ.12 The graph after mapping entities with coreference table.13 A pruned knowledge gøraph.14 An enriched €TIẨV.15 A subject-based question-answer pal .16 An object-based question-ansWer Pat .17 A modifier-based question-anSwer Pal .18 A question-answer pair searcher result .19 Visualization system compOn€nfS. - ¿+ sc + 3213 EEEeEsrresrsrrsree 58 Figure 3.20 Graph visualization: neo4jd3 Vs €OVIS.21 AllenNLP coreference resolution vIsual1ZatiOT.- -- sec xsssserss 60 Mi Figure 3.22 Visualization sySf€T4 OV€TVICW. Án 1H S1 SH 1T 1n HH rệt 61 Figure 3.23 Raw text input S€CfIOH.24 Knowledge graph VI€WT.25 Sentence decomposer VI€WCT.- ác TH nh TH Hà Hành nh nh nh 63 Figure 3.26 Coreference resolution VI€W€Y.27 KG stats: Relation distrIDufIOTI. --- «xxx x11 ng tt rưy 65 Figure 3.28 KG stats: Named entity distrIbutIOn.29 KG stats: Predicate modifier distrIbufIon.30 Question-answer pair table .31 Question-answer pair Searcher VICWCT.1 SRL-KG’s knowledge graph on example text 1 oo.
eee eeeeeeeeeeeceeeereeeeees 73 Figure 4.2 Team UWA’s knowledge graph on example text Ï.3 Modifiers of predicate “debuted”? 0. eee eecceccccescceseceeseeeeeeeeseeseaeeeeseeeseeenees 74 Figure 4.4 Team UWA’s knowledge graph on example text 2.5 SRL-KG’s knowledge graph on example text 2.6 Team UWA’s knowledge graph on example text 3.7 SRL-KG’s knowledge graph on example text 3.8 Modifiers of predicate ““T€fUTTS””.9 SRL-KG’s knowledge graph on example text 4.1 Knowledge graph about Fernando TOFT€S. - c6 6c 319v TT HH TH nh nh TH HH Hành ch nh ngờ 88 Figure 6800) xuan.4 Predicate vill TABLE OF TABLES œ4ÍE]k› 0020.2 PropBank annotations na.1 An example about coreference †abÏ€.1 Performance of SRL-KG on OIE-2016 (3200 senfences).2 Performance of SRL-KG on CaRB (641 sentences).- -- ¿5s ss<sssss2 71 Table 4.3 SRL-KG time perÍOrImaTnC.-- c2 +19 9E ng HH HH nh nh nành 79 Table 4.4 BLEU score on SQuAD of different question generation models. 81 1X LIST OF ABBREVIATIONS œ4ÍE]k› AI Artificial Intelligence CNN Convolutional Neural Network KG Knowledge Graph LSTM Long short-term memory NER Named Entity Recognition NLP Natural Language Processing RNN Recurrent Neural Network SRL Semantic Role Labeling ABSTRACT It is obvious that text data accounts for a great proportion of the total data available.
With the growth about the amount of data, deriving valuable information from text becomes more and more challenging. Knowledge graph is considered as a powerful mechanism to store the text data in an organized way in order to serve further applications in Artificial Intelligence. Most of popular knowledge graphs, e. Google Knowledge graph, DBpedia, Wikidata and YAGO, requires a large quantity of data as input and depends on human intervention processes on structural text data.
The main goal of this thesis is to design and implement an automatic knowledge graph construction pipeline from unstructured text. There have been several existed works on this topic, however information lost is one big problem that they have been still dealing with. Therefore, the center of attention in this thesis is to enhance the ability to make use of raw text in order to preserve the provided information as much as possible by using semantic role labeling approach cooperated other Natural Language Processing techniques like named entity recognition and coreference resolution. Moreover, to prove the usefulness of constructed knowledge graph from the proposed pipeline, there is additional tasks relating to question generation and question semantic searching.1 Context In recent years, a buzzword of industry 4.0 has been mentioned and appeared dominantly on social media with its promise to take the advantages of data to revolutionize manufacturing.
It can be said that data is the foundation of megatrends in information technology fields relating to Artificial Intelligence (AI) which can strongly support the industry. On the grounds that the number of online activities keeps arising, the amount of data has been increasing exponentially and becomes more various in nature. We obviously know that data may contain underlying information and knowledge. However, data seems useless if it is not processed properly.
Therefore, a big challenge is to extract value from collected raw data for further uses. In order to make use of information from text data, knowledge graphs appeared as a good way of knowledge representation. The term knowledge graph became obviously notable due to the introduction of Google’s Knowledge Graph in 2012 in which the basic motto is to focusing on searching things not strings [1] by utilizing knowledge graphs. Special concerns about this technology led to the fact that many other companies began to develop their own knowledge graphs such as: DBpedia [2], YAGO [3], Wikidata [4] or Freebase [5].
Due to knowledge graph’s powerful capability of knowledge representation and reasoning, it was used as the backbone of artificial intelligence to serve both academic and industrial purposes. To construct a knowledge graph, a set of triples (head, relation, tail) must be produced from text data. Therefore, triple extraction is one of the key basic steps. Until now, methods of triple extraction have been being developed with many different techniques which relate to statistical based and neural network based models.
However, it is supposed to be unclean and insufficient to instantly apply extracted triples through these methods for knowledge graph construction. As a result, the main purpose of this thesis is to make triples extracted from unstructured text more useful and valuable in order that high-quality knowledge graphs can be created from them to support advanced uses.2 The problem and its significance Triple extraction, a subset of information extraction, is an aspect in Natural Language Processing that takes an important role in knowledge graph construction. One of the most well-known information extraction system is Open Information System (OpenIE) [6] which was first introduce in 2007. After that, a wide range of OpenIE system has been developed with different approaches in which modern ones are associated with neural networks so as to enhance performance.
However, by doing experiments with some OpenlE systems such as: OpenIE-5 (including CALMIE [7], BONIE [8], RelNoun [9] and SRLIE [10]), MinIE [11], Supervised-OIE or RnnOIE [12] and IMoJIE [13], it can be found that those systems tend to produce many triples from a sentence to cover all kinds of results. Subjects or objects extracted from those techniques are regularly long sequences of words which includes semantic roles.