Ứng Dụng Knowledge Graph và BERT cho Phân Loại Ba Tuples tại Đại Học Quốc Gia TP.HCM

Khóa luận tốt nghiệp nghiên cứu ứng dụng knowledge graph và BERT trong phân loại ba tuân cho tiếng Việt, mang lại giải pháp hiệu quả.

Chuyên ngành

Information Systems

Người đăng

Ẩn danh

Thể loại

Thesis

2021

85
3
0

Phí lưu trữ

30 Point

Mục lục chi tiết

ACKNOWLEDGEMENTS

1. CHAPTER 1: PROBLEM STATEMENT

1.1. Context

1.2. The problem and its significance

1.3. Motivation

2. CHAPTER 2: RELATED WORKS

2.1. Overview

2.2. KG-BERT: BERT for Knowledge Graph Completion

2.3. Google Knowledge Graph

2.4. Knowledge Graph Base

2.5. Neo4j

2.6. Knowledge Graph Embedding

2.7. Doccano

3. CHAPTER 3

3.1. Triple Classification

3.2. Basics Graph

3.3. The Graph Database Space

3.4. Deeper matters

4. CHAPTER 4: SYSTEM DESIGN AND EVALUATION

4.1. Overview of the SYSTEM

4.2. Model building

4.3. Web application architecture

4.4. Encapsulation and deployment

4.5. Data preparation

6. CHAPTER 6: FUTURE WORKS

LIST OF FIGURES

Tóm tắt

I. Tổng quan về Ứng Dụng Knowledge Graph và BERT tại Đại Học Quốc Gia TP

Trong bối cảnh phát triển công nghệ thông tin, việc ứng dụng Knowledge GraphBERT trong phân loại ba tuples đã trở thành một xu hướng quan trọng. Đại học Quốc gia TP.HCM đã tiên phong trong việc nghiên cứu và ứng dụng các công nghệ này nhằm nâng cao khả năng xử lý ngôn ngữ tự nhiên. Sự kết hợp giữa Knowledge GraphBERT không chỉ giúp cải thiện độ chính xác trong phân loại mà còn mở ra nhiều cơ hội mới cho nghiên cứu và ứng dụng trong lĩnh vực này.

1.1. Khái niệm về Knowledge Graph và BERT

Knowledge Graph là một hệ thống lưu trữ thông tin giúp kết nối các thực thể và mối quan hệ giữa chúng. BERT, viết tắt của Bidirectional Encoder Representations from Transformers, là một mô hình học sâu giúp máy tính hiểu ngôn ngữ tự nhiên một cách hiệu quả hơn.

1.2. Tầm quan trọng của nghiên cứu tại Đại Học Quốc Gia TP.HCM

Nghiên cứu tại Đại học Quốc gia TP.HCM không chỉ tập trung vào lý thuyết mà còn chú trọng đến ứng dụng thực tiễn. Việc áp dụng Knowledge GraphBERT trong phân loại ba tuples sẽ giúp nâng cao chất lượng dữ liệu và cải thiện khả năng tìm kiếm thông tin.

II. Vấn đề và Thách thức trong Phân Loại Ba Tuples

Phân loại ba tuples gặp nhiều thách thức, đặc biệt là trong việc xác định tính chính xác của thông tin. Các vấn đề như độ chính xác của dữ liệu, khả năng hiểu ngữ nghĩa và ngữ cảnh là những yếu tố quan trọng cần được giải quyết. Việc thiếu hụt dữ liệu chất lượng cao cũng là một thách thức lớn trong quá trình phát triển mô hình.

2.1. Độ chính xác của dữ liệu trong phân loại

Độ chính xác của dữ liệu là yếu tố quyết định đến hiệu quả của mô hình phân loại. Việc thu thập và xử lý dữ liệu từ nhiều nguồn khác nhau sẽ giúp cải thiện độ chính xác này.

2.2. Khó khăn trong việc hiểu ngữ nghĩa

Mô hình cần phải hiểu được ngữ nghĩa của các từ và mối quan hệ giữa chúng. Điều này đòi hỏi một lượng lớn dữ liệu được gán nhãn chính xác để huấn luyện mô hình.

III. Phương pháp Ứng Dụng Knowledge Graph và BERT trong Phân Loại

Để giải quyết các thách thức trong phân loại ba tuples, phương pháp kết hợp Knowledge GraphBERT đã được áp dụng. Phương pháp này không chỉ giúp cải thiện độ chính xác mà còn tăng cường khả năng hiểu ngữ nghĩa của mô hình. Việc sử dụng Machine Learning trong quá trình huấn luyện cũng đóng vai trò quan trọng.

3.1. Kết hợp Knowledge Graph với BERT

Sự kết hợp này cho phép mô hình hiểu rõ hơn về mối quan hệ giữa các thực thể, từ đó cải thiện khả năng phân loại ba tuples.

3.2. Ứng dụng Machine Learning trong huấn luyện

Việc áp dụng các thuật toán Machine Learning giúp tối ưu hóa quá trình huấn luyện và nâng cao hiệu suất của mô hình phân loại.

IV. Kết quả Nghiên cứu và Ứng dụng Thực tiễn

Kết quả nghiên cứu cho thấy việc ứng dụng Knowledge GraphBERT đã mang lại những cải tiến đáng kể trong phân loại ba tuples. Mô hình đã đạt được độ chính xác cao trong việc xác định các mối quan hệ và thực thể trong ngữ cảnh tiếng Việt. Điều này mở ra nhiều cơ hội cho các ứng dụng trong lĩnh vực du lịch và thông tin.

4.1. Đánh giá hiệu suất mô hình

Mô hình đã đạt được độ chính xác lên đến 80% trong việc phân loại ba tuples, cho thấy tiềm năng lớn trong việc ứng dụng thực tiễn.

4.2. Ứng dụng trong lĩnh vực du lịch

Việc áp dụng mô hình trong lĩnh vực du lịch giúp cải thiện khả năng tìm kiếm thông tin và nâng cao trải nghiệm người dùng.

V. Kết luận và Tương lai của Nghiên cứu

Nghiên cứu về ứng dụng Knowledge GraphBERT trong phân loại ba tuples tại Đại học Quốc gia TP.HCM đã mở ra nhiều hướng đi mới cho nghiên cứu và ứng dụng trong lĩnh vực xử lý ngôn ngữ tự nhiên. Tương lai của nghiên cứu này hứa hẹn sẽ mang lại nhiều giá trị cho cộng đồng và ngành công nghiệp.

5.1. Hướng đi mới trong nghiên cứu

Nghiên cứu sẽ tiếp tục mở rộng và cải tiến các phương pháp hiện tại để nâng cao hiệu quả và độ chính xác của mô hình.

5.2. Tác động đến ngành công nghiệp

Ứng dụng các công nghệ này trong ngành công nghiệp sẽ giúp cải thiện quy trình làm việc và nâng cao chất lượng dịch vụ.

10/07/2025
Khóa luận tốt nghiệp applying knowledge graph and bert for vietnamese triple classification

Trích đoạn nội dung tài liệu

VIETNAM NATIONAL UNIVERSITY HOCHIMINH CITY UNIVERSITY OF INFORMATION TECHNOLOGY ADVANCED PROGRAM IN INFORMATION SYSTEMS APPLYING KNOWLEDGE GRAPH AND BERT FOR VIETNAMESE TRIPLE CLASSIFICATION BACHELOR OF ENGINEERING IN INFORMATION SYSTEMS HO CHI MINH CITY, 2021 NATIONAL UNIVERSITY HOCHIMINH CITY UNIVERSITY OF INFORMATION TECHNOLOGY ADVANCED PROGRAM IN INFORMATION SYSTEMS PHAM BINH AN - 16520016 NGUYEN HUY CƯỜNG - 16520148 BACHELOR OF ENGINEERING IN INFORMATION SYSTEMS THESIS ADVISOR „ Associate Professor. DO PHÚC HO CHI MINH CITY, 2021 ASSESSMENT COMMITTEE The Assessment Committee is established under the Decision. by Rector of the University of Information Technology. - Member ACKNOWLEDGEMENTS First of all, we would like to express our profoundest appreciativeness to Associate Professor Đỗ Phúc who was very passionate in guiding and supporting us throughout the thesis.

His timely teachings and admonitions have helped us continually change to complete the final one in the best possible circumstances. Wholehearted gratefulness for Mr Lé Hung, with the experience of forerunner, our exceptional older brother has countenanced us to learn and acquire a lot of knowledge and skills from him not only in project work. Our top positive receptions also go to all the members of Faculty of Information Systems as well as everyone of University of Information Technology for their guidance, supports to us with greatest cares. Not the least, we feel in extremely need of showing our gratitude to our family, our friends, and our classmates for every support and love that we have received on our maturity path.

Pham Binh An & Nguyen Huy Cuong — students of aep 2016. TABLE OF CONTENTS œ4Í-Ìk› ACKNOWLEDGEMENTTS. Họ ng i TABLE OF CONTIENTTS. Họ họ h m ii LIST OF FIGURES.

«Ănng iv ABSTRACCT. - sọ vii Chapter 1 PROBLEM STATEMENT.2 The problem and its SIØIÍÍCATCC. tt vn 1 122 1E 11111111 HH re 3 Chapter 2 RELATED WORKS. HH HH HH non ngư, 5 2.1 OVELVIOW mẻ ố ố ố ốố Ố.2 KG-BERT: BERT for Knowledge Graph Completion.3 Google Knowledge Oraphh.

--- -- + tk 112121111 ng vn ngư5 2.6 Knowledge Graph Embedding. ---- --- -< HH TH HO HH HH ngư 8 3.-- - «sành TH TH nh greg 3.4 Transfer learning oo.1 Overview of BERT 3.2 Evolution of BERT 3. BERT architecture and COMPOMNENES.4 BERT pre-training.6 Self-attention mechanism.7 Position-wise Feed-Forward Networks.6 Classification SA oa Ko) ke) | ee 3.2 Precision and Recall sẻ Ea. Chapter 4 SYSTEM DESIGN AND EVALUATION.1 Overview of the SYSf€IM.

Án TH TH TH TH TT TT HH cườ 48 4.2 Model building ve 4.3 Web application arCHI{€CfUIT.- ¿+ %2 1v 12v vn ng gàng 4.5 Encapsulation and deployment " 4.7 Data pr€DATALIOH. Ác TT HT TH HT TT TH Hy 2ñ 5⁄0 on.--- ---- 55-5 nọ nh nh 71 Chapter 6 FUTURE WORRKS. nọ Họ nọ In.- ---- Ăn họ TT nh nh 74 iil LIST OF FIGURES œ4Í-Ìk› Figure 3.1: Illustration of triple classification task .3: An overview of the graph database space from Graph Database 2nd edition by O’ Reilly 02/0800.4: An experiment run between RDBMS and Neo4j from Graph Databases for 2030019100 vo.5: Comparison between SQL and NoSQL by Trilochan Parida.6: Knowledge Graph with three nodes and their properties .7: Knowledge Graph example .8: An example of transfer Ï€aTn1nE. ¿55 5+ xxx E#EeEekEkskskeekrrkrxee 16 Figure 3.9: Types of transfer Ï€aTTITE.- óc c3 121911111 11 1111911 1 1 g1 1g 1x nh me, 17 Figure 3.10: Transfer knowledge from one model to anofher.--- ¿+ ++s+s++s+sc+x+ecss 18 Figure 3.11: Two methods of transfer learnIng.12: The details of fine-tuning and feature ©XfTACfOT.

----- «+ x+x+sx+x+sesxserse 20 Figure 3.13: Transformers architecture from Attention Is All You Need by A.14: BERT Dr€-fTA1nIT. 6 6c c1 191911 11911 1 1911 1111 1 ng nh nh Hư, 24 Figure 3.15: Mask language moeÏIn.16: Input Computing from BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding by J. 5 55 «5s £+sx+exsexsxss 26 Figure 3.17: Illustrating tokenization Of & WOT .18: The 128-dimensional positional encoding for a sentence with the maximum length Of 5.- -- --- +11 2119119111911 1111 91 1 1111 HT TT HH TT nà Hàn Tư 28 Figure 3.19: Attention mechanism from Neural machine translation by jointly learning to align and translate 0n.20: Two sub layers in an €TICOI€T-.21: Query, Key and Value vectors Figure 3.22: Self-attention flow .--- ch 2v TH HH HT TT HH TT HH ngờ 32 Figure 3.23: Set of Queries, Keys and VaÏU€S.-- - -- 5 + tk HT HH HH grư 33 Figure 3.24: The concatenation of aft€TIẨIOTIS.- ¿+ 5+ x3 xxx HH ng rrưy 33 Figure 3.25: Feed-Forwards Networks .ó- Gv HH TH HH HH ng HH Hàng.26: Residual connection in encoding component from Exploring the Depths of Recurrent Neural Networks with Stochastic Residual Learning .27: The covariance shift methOd.- ¿+ 65+ + +*£+k+tEsE+kEeEekeEskekrrkrkrkerkrerre 36 Figure 3.28: Matrix in layer normalization 0. cece - 5 eeeeeeeeeeeseeeeeeseeeesceesasessesseetaseees 36 Figure 3.29: Two layers of neural n€EWOTÍK.---- 6 +3 xxx 1 11119111 910 1 gu gu ng re, 37 Figure 3.30: Classification layer to product the T€SuÏ(.31: The comparison between PyTorch and 'TensorFFÏlow.

--- 5< ss<sx+sx+s 40 Figure 3.32: The number of repositories has been written in PyTorch.33: The advantage Of PPtOTCÌ.34: Doccano interface oo. ------ 6 + S11 121 11 1 11v TT TH TH TH TH nh 44 Figure 3.35: Predictions for binary task with sample size 1s [Ũ.--- eee - - + <2 512% 4121 1 1 3 121 9101111211 HH nh như 49 Figure 4.2: Triple extraction DFOĐTAIH.- -- c2 S2 E2%E3128E1E51 E1 kg ng ng rrưy 50 Figure 4.3: Neo4j platform user interface .- ---¿- - - 5+ + SE +3 E*E*E SE kh vn hư 50 Figure 4.4: An overview of the training phase .5: Architecture of our pÏAfÍOTT. --- ¿6 6+4 *+%E*kE#EEvE+eEeeEekEekrekreksseekerkreerxee 53 Figure 4.6: The schema for SSf€IN.7: The UI of configuration part in Ne@O4]. ó- cà LH TT HH TT HH TH HH ưy 57 Figure 4.9: The interaction between client and docker hoSt.---- 5 + ++s++£+s+sc++ecss 58 Figure 4.10: Triple extraction flOW.

5G 2c 19191911111 1 191111111 g1 ng nh ngư.12: Labeled text 1n DOCCATO. s5 6 1912511 1 919321 1 1 ng ng gàng ưy 61 Figure 4.13: Triple is extracted from description example .14: Train, test and valid set 0.------ cece <4 E13 E1 3 9121 51113 E111 111kg re 64 2008190. 5G 2c 3 2119111911 1191 91 0111111 TH Hàn TT nàn ng.18: The set of parameters use for tFA1TIITE.19: The recording accuracy of training PrOCeSS. ee eeesesseeeeesesescseseeeeeeeeeseraenees 66 Figure 4.20: The recording error rate of training DFOC€SS.--¿ + 2+ + ++s£+x+£sxzx+sreeres 67 Figure 4.21: Display all nodes and their relationShIp.23: Output returns true example .24: Output returns false example vi ABSTRACT Natural language processing 1s one of the problems that 1s focused on research and application in the field of machine learning as well as artificial intelligence.

The shortage of available data is one of the biggest challenges for language processing as single language data in other languages could not be compared to English. By applying Neo4j - the leading graph database platform of the industry that allows modeling, visualization data relationships and BERT - state-of-the-art model in contextual understanding, we focus on improving triple classification solution with Vietnamese tourism domain. Vil Chapter 1 PROBLEM STATEMENT 1.1 Context In an era of digital transformation, data and information gradually becomes an irreplaceable piece in various companies. How to create the connection between information is one of the main tasks that make systems more powerful.

Google has introduced their Knowledge Graph system and as they said this knowledge base can enable everyone to search for everything that Google knows about. In general, Knowledge Graph is about collecting information of objects in the real world which can be constructed from entities and the relation between them. In a Knowledge Graph, each entity can be represented as a node and the relationship is an edge, which is called a triple structure that consists of head, relation, tail. In natural language processing, researchers mainly concentrate on demonstrating their system can understand domain-specific knowledge semantically and syntactically via some following experiments: question answering, sentence classification, semantics analysis and the like.

In our project, we focus on the triple classification problem which means searching in paragraphs whether a triple is correct or not after understanding the context. Our ultimate goal is to minimize the distance between human language and machine. To overcome this barrier, it is vital to make machine learning to understand triple input from the user and respond if it is correct or not. Because of the limited time and resources, the application will be designed for a specific domain.

The domain is about Vietnamese tourism.2 The problem and its significance In many documents, researchers have proved the effect of valuable data in systems. Information accuracy is one of the hardest tasks that promote our group to collect and annotate data that brings the best value for the next steps. Then, we define logics to extract triple from the above result, from an example: “Quang Ninh có danh lam thắng cảnh là Vinh Ha Long, noi đây còn có đặc san chả mực Ha Long”, we can get two triples (Quang Ninh, có danh lam thang cảnh là, Vinh Hạ Long) and (Quang Ninh, có đặc san, cha mực Hạ Long). All of the results lead to constructing a knowledge graph which contains some relevant information and is a value for the next step, model building.

Building a model that can understand the context and give correct answers is one of our main tasks. If we have a paragraph: “Minh đến ngân hàng Sacombank dé gửi tiền, ngân hang này yêu cầu anh ấy phải chuẩn bi những thu tục cần thiết”, two pair words (‘Minh’, ‘anh ay’) and (‘Sacombank’, 'ngân hàng nay’) have the same meaning although it represented as two different words. The second examples: “Minh dén Bitexco xem buổi ra mắt sản phẩm mới, anh mua chiếc điện thoại tại đây”, the triple “Minh, mua chiếc điện thoại, tại Bitexco” is correct or not? If our group can solve two challenges above, the system can achieve a powerful performance. To serve an essential demand, academic and industry researchers remain that it is incredibly difficult to actually utilize any deep learning model in production.

Could model in production handle in an expected time, prevent the latency, does it keep the same performance as it before and so on, usual questions that are not easy to find out the answers and in other cases are the significant challenges of applications. At that time, researchers approached two ways to deploy models: cloud services and build their own server. In the first way, it is costly to maintain so far and our systems are passive by depending on their architecture. So we decided to deploy a model by the second way because of the limited resources.3 Motivation We are living in the world of Artificial Intelligence (AD, machines can support people to communicate with those who are handicapped or are incapable of ordinary language and others.

To assist them, machines must understand human languages in many types of format such as text, voice, gesture and the rest. In our project, textual handling is mentioned in the field of NLP. In many prevalent languages, it can be supported by a large public dataset but this process is harder in Vietnamese. We are building a system that helps a machine understand Vietnamese language at a relatively acceptable level of accuracy.

We generate knowledge graphs with free Vietnamese text that are collected 60% in Wikipedia and the rest in newspapers. We try to visualize all of the relevant information in our knowledge graph that follows the triplet (head, relation, tail). After collecting, our capacity contains 2,5 thousand triples with 442 descriptions respectively. The increment of triple can improve the accuracy of application but it must informatively first.

We organized knowledge from collected data before storing them in Neo4j graph platform.

Nội dung được bảo vệ bản quyền — Tải xuống đầy đủ

Tài liệu "Ứng Dụng Knowledge Graph và BERT trong Phân Loại Ba Tuples tại Đại Học Quốc Gia TP.HCM" trình bày những ứng dụng tiên tiến của công nghệ Knowledge Graph và mô hình BERT trong việc phân loại ba tuples, một khía cạnh quan trọng trong xử lý ngôn ngữ tự nhiên. Bài viết không chỉ giải thích cách thức hoạt động của các công nghệ này mà còn nêu bật những lợi ích mà chúng mang lại, như cải thiện độ chính xác trong việc phân loại thông tin và tối ưu hóa quy trình tìm kiếm dữ liệu. Độc giả sẽ tìm thấy những thông tin hữu ích giúp mở rộng hiểu biết về cách mà các công nghệ hiện đại có thể được áp dụng trong nghiên cứu và phát triển hệ thống thông tin.

Để khám phá thêm về các vấn đề liên quan đến bảo vệ thông tin trong hệ thống tính toán, bạn có thể tham khảo tài liệu Luận văn nghiên cứu một số vấn đề bảo vệ thông tin trong hệ thống tính toán lưới. Ngoài ra, nếu bạn quan tâm đến các phương pháp tổng hợp cảm biến cho robot di động, hãy xem tài liệu Luận văn nghiên cứu phương pháp tổng hợp cảm biến dùng cho kỹ thuật dẫn đường các robot di động. Những tài liệu này sẽ giúp bạn có cái nhìn sâu sắc hơn về các ứng dụng công nghệ trong lĩnh vực nghiên cứu và phát triển.