VIETNAM NATIONAL UNIVERSITY HO CHI MINH CITY HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY PHẠM THANH HỮU QA SYSTEM FOR REAL ESTATE LAW IN VIETNAM Major : Computer Science Major code : 8480101 MASTER’S THESIS HO CHI MINH CITY, July 2023 THIS RESEARCH IS COMPLETED AT HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY – VNU – HCM Supervisor 1: ASSOC. QUAN THANH THO, PhD. NGUYEN TIEN THINH, PhD. TRAN TUAN ANH, PhD.
BUI THANH HUNG, PhD. Master’s thesis is defended at HCM City University of Technology, VNU- HCM City on 13/07/2023 Master’s Thesis Committee: 1. VO THI NGOC CHAU 2. PHAN TRONG NHAN 3.
TRAN TUAN ANH 4. BUI THANH HUNG 5. BUI CONG GIAO Approval of the Chairman of Master’s Thesis Committee and Dean of Faculty of Computer Science and Engineering after the thesis is corrected (If any). CHAIRMAN OF THESIS COMMITTEE DEAN OF FACULTY OF COMPUTER SCIENCE AND ENGINEERING VIETNAM NATIONAL SOCIALIST REPUBLIC OF VIETNAM UNIVERSITY HO CHI MINH CITY Independence – Freedom - Happiness HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY THE TASK SHEET OF MASTER’S THESIS Full name: PHAM THANH HUU Student code: 2171066 Date of birth: 03.1978 Place of birth: QuangNgai Major: Computer Science Major code : 8480101 I.
THESIS TITLE (In Vietnamese): HỆ THỐNG HỎI ĐÁP TỰ ĐỘNG LUẬT BẤT ĐỘNG SẢN VIỆT NAM. THESIS TITLE (In English) : QA SYSTEM FOR REAL ESTATE LAW IN VIETNAM. TASKS AND CONTENTS: Developing a chatbot capable of responding to legal real estate queries. THESIS START DATE : 22.
THESIS COMPLETION DATE: 09. QUAN THANH THO, PhD and DR. NGUYEN TIEN THINH, PhD. INSTRUCTOR INSTRUCTOR HCM City, 09/06/2023 CHAIRMAN OF DEAN OF PROGRAM COMMITTEE COMPUTER SCIENCE AND ENGINEERING i Acknowledgment I would like to express my deepest gratitude to my advisors - Assoc.
Quan Thanh Tho, for his valuable and constructive suggestions during the planning and development of this research work. His willingness to give his time so generously has been very much appreciated. Moreover, his advice on algorithms and his recommendations on solutions when I had to deal with problems during doing this research. Finally, I wish to thank IVS JSC for funding this study.
ii Abstract Intelligent legal services have emerged in recent years due to the application of AI technology to the law industry; however, these have yet to be developed in Vietnam since there is a lack of research into automatic processing in the Vietnamese language. In this thesis, the author proposes to build a chatbot that can effectively and automatically answer legal questions, especially those related to real estate. The most important module of the chatbot is the Legal Statutes Identification (LSI), which identifies the legal statutes relevant to a given description of facts or evidence of a legal document (such as a legal question or a description of a legal fact). To deploy the LSI model, the author has built an LSI dataset including more than 300,000 legal questions and millions of judgments of the Supreme People’s Court of Vietnam.
Three models are presented in this thesis. The first is an ML-based model in which the LSI is performed by the Support Vector Machine after the input questions have been word-embedded with TF-IDF Embedding. The second model, based on deep learning, will implement LSI downstream tasks after using a new model called LegarBERT to construct word embedding for the input question. Finally, the author attempts to build LSI using graph machine learning by encoding legal reasoning as nodes and edges, representing by queries, a legal articles, and legal key word (legal terminology).
TÓM TẮT LUẬN VĂN THẠC SĨ Các dịch vụ pháp lý thông minh đã xuất hiện trong những năm gần đây nhờ sự áp dụng của công nghệ Trí tuệ Nhân tạo vào ngành luật; tuy nhiên, tại Việt Nam, chúng vẫn chưa được phát triển do thiếu nghiên cứu về xử lý tự động trong tiếng Việt. Trong luận văn này, tác giả đề xuất xây dựng một chatbot có khả năng trả lời tự động và hiệu quả các câu hỏi pháp lý, đặc biệt là các câu hỏi liên quan đến bất động sản. Mô-đun quan trọng nhất của chatbot là Hệ thống Xác định Căn cứ Pháp lý (Legal Statutes Identification - LSI), được sử dụng để xác định các căn cứ pháp lý liên quan đến một mô tả cụ thể về sự kiện hoặc bằng chứng từ một văn bản pháp lý (như một câu hỏi pháp lý hoặc một mô tả về sự kiện pháp lý). Để triển khai mô hình LSI, tác giả đã xây dựng một tập dữ liệu LSI gồm hơn 300.000 câu hỏi pháp lý và hàng triệu bản án của Tòa án Nhân dân Tối cao Việt Nam.
Luận văn này trình bày ba mô hình. Mô hình đầu tiên dựa trên máy học (ML), trong đó LSI được thực hiện bằng Máy Vector Hỗ trợ sau khi câu hỏi đầu vào được biểu diễn bằng phương pháp Nhúng TF-IDF. Mô hình thứ hai, dựa trên học sâu, sẽ thực hiện các tác vụ LSI sau khi sử dụng một mô hình mới được gọi là LegarBERT để xây dựng việc nhúng từ cho câu hỏi đầu vào. Cuối cùng, tác giả cố gắng xây dựng LSI bằng cách sử dụng học máy đồ thị bằng cách mã hóa lý luận pháp lý thành các nút và cạnh, biểu thị bằng các truy vấn, các điều khoản pháp lý và thuật ngữ pháp lý.
Keywords: LSI, Law Graph, Intelligence Law Service, Vietnamese Law Questions and Answers, Vietnamese Embedded Word, Law Prediction. iii Declaration of Authenticity I guarantee this research is my own, conducted under the supervision of Assoc. Quan Thanh Tho. The contents and results of this research are legitimate and have not been published in any forms prior to this.
The data and materials used for the analysis and feedback are derived from various resources and which are appropriately listed in the References section. The data and results of several other authors and organizations have been used and have been aptly cited. If there is any plagiarism, I stand by our actions and are to be held responsible for it. Ho Chi Minh City University of Technology is not responsible for any copyright infringement relating to this dis- sertation.
Ho Chi Minh City, June 2023 Author Pham Thanh Huu iv Contents 1 Introduction 1 1.2 Ojectives and Scope .3 Contributions of the Thesis .4 Organization of the Thesis. 5 2 Legal Document Structure and Data 7 2.1 VN-LandLaw-2013 Corpus .2 Formal Structure of a Legal Document .3 Legal Data Sourcing .2 Legal Entity Extration .3 Legal Relation Extraction .4 The TF-IDF Matrix of Vietnam Land Law .5 Legal Data Summary Statistics .1 Legal Data Classification .2 Basic Legal Data Statistics .3 Unbalanced Legal Data .4 Vietnam Land Law Article Semantic Relations Matrix .5 Vietnam Land Law Article Co-occurrence Matrix in LSI Dataset .1 QAS Research in NLP .2 Law-related Global QAS Research .3 Vietnamese Law-related QAS Research. 37 CONTENTS CONTENTS 4 Background 40 4.1 Term Fequency-Inverse Document Frequency (TF-IDF) .2 Support Vector Machine (SVM) .7 Masked Language Modeling (MLM) .8 Fine-tune a Pretrained Model .9 Graph Convolution Neural Network (GCN) and Graph Attention Netowrk(GAT) 46 4.3 Legal Domain Background. 47 5 The Proposed System 51 5.2 Overall System Architecture .3 The Main User Cases .5 The Evaluation/Acceptance Criteria .1 Chatbot System Acceptance Criteria .2 LSI Model Metrics.
62 6 LSI by Linear Support Vector Classification with TF-IDF Embedding 63 6.5 Results and Conclusions. 65 7 LSI by Multi Label Classification with LegarBert 67 7.3 Legar Answering Engine. 69 vi CONTENTS CONTENTS 7.1 LegarBert Training from PhoBert .6 Legal-Masked Strategy .7 Results and Conclusions. 75 8 LegarHKB: A LSI Retrieval Model using Heterogeneous Knowledge Graph for the Viet- namese Law Domain 76 8.5 Results and Conclusions.
85 10 List of Deliverables 87 List of Publications 88 References 103 vii List of Figures 1.1 LSI ChatBot’s response .1 Data collection procedure .2 Data labeling application screenshot - Login Page .3 Data labeling application screenshot - Home Page .4 Data labeling application screenshot - Labeling Page .5 The tree of ”Chủ thể.” Subjects of legal relations .6 The tree of ”Hành vi” or ”Quan hệ pháp lý.” Acts/ Legal relations.7 The TF-IDF Matrix of Vietnam Land Law.8 Legal data categories.9 Supervised LSI training data statistics per legal category.10 Supervised LSI training data statistics per legal category (Distribution).11 Semi-supervised LegarBert (MLM) training data statistics per legal category(From books).12 Semi-supervised LegarBert (MLM) training data statistics per legal category(From books) (Distribution).13 Semi-supervised LegarBert (MLM) training data statistics per legal category(From Supreme People’s Court).14 Semi-supervised LegarBert (MLM) training data statistics per legal category(From Supreme People’s Court) (Distribution).15 Unbalanced legal data phenomenon.16 Heatmap of 212 legal documents’ TF-IDF vectors’ cosine similarity .17 Semantic relations of articles 35-51 of chapter IV(Land use master plans and plans).18 Heatmap of Vietnam Land Law articles co-occurrence in LSI dataset.20 High concurrency, high semantic similarity.21 High concurrency, low semantic similarity.1 Timeline of automated law research .1 Architecture of attention Model .2 Architecture of auto encoder. 43 LIST OF FIGURES LIST OF FIGURES 4.3 Architecture of BiLSTM .4 Architecture of PhoBERT .5 Graph Attention Neural Network.6 IVS JSC overview.7 Vietnamese legal structure.8 Vietnam real estate law structure.2 Overall system architecture.3 The main user cases.4 Main screens of the chatbot.5 Legal quick lookup popup.6 Example of KU calculation.7 Vietnam Land Law 2013 Long-tail dataset.1 LSI by Support Vector Machine with TF-IDF Embedding model.1 The Answering Engine of the Legar System.2 LegarBert embedding model training by MLM tasks.3 LSI by Multi Label Classification with LegarBert.1 LSI by Heterogeneous Knowledge Graph.2 Data transformation process.3 Nodes and Edges Definition.5 Graph Demo with some nodes and edges.1 The directions for future work. 86 ix List of Tables 1.1 NLP techniques used in legal domain.1 S, O, R, TO, T legal question analysis.2 Entity extraction example.3 Top 100 single-word TF-IDF values from 212 Vietnamese Land Law.4 Top 200 single-word TF-IDF values from 212 Vietnamese Land Law.5 Legal Document Sentences and Words Statistics .6 Most-paired articles .1 QAS research in NLP .2 Law-related global QAS research .3 Vietnamese Law-related QAS research .2 Long-tail dataset .1 Train/Val/Test Dataset .3 LSI by Support Vector Machine with TF-IDF Embedding results .1 Hyperparameter of LegarBert training.2 Hyperparameter of LSI by LegarBert.3 Perplexity comparing with MLM task .4 LSI by Multi Label Classification with LegarBert results.5 LSI by Multi Label Classification with LegarBert K-Utility .1 Hyperparameter of LSI by Heterogeneous Knowledge Graph .2 LSI by Heterogeneous Knowledge Graph results .3 LSI by Heterogeneous Knowledge Graph K-Utility .1 Comparing 3 models by Precision/Recall/F1 .2 Comparing 2 models by KU. 85 LIST OF TABLES LIST OF TABLES 9.3 Summary the embedding capability of 3 models.
85 xi Chapter 1 Introduction 1.1 Motivation In most nations, the legal system is overburdened by a backlog of cases, particularly in low-level judiciaries. Though speedy justice acts exist, the process in the legal domain is extremely laborious. The legislation to which businesses and citizens have to abide is growing at a constant rate both in complexity and volume. The data present in legislation is mostly in an unstructured format in legal documents [2].
This makes the task of retrieving information highly inefficient and time- consuming, particularly when there are huge quantities of data involved. Further, the utility of such data differs broadly and relies on its representation and structure. In this scenario, legal professionals and users might find it highly problematic to explore the legal data while investigating a specific case or dealing with particular circumstances, even when the data is accessible [3]. These problems have resulted in the necessity of devising better methods for structuring and searching across huge amounts of legal data [4].