VIETNAM NATIONAL UNIVERSITY, HANOI UNIVERSITY OF ENGINEERING AND TECHNOLOGY VUONG THI HAI YEN MODELING AND LEARNING TEXTUAL AND STRUCTURAL RELATIONS FOR DEEP LEGAL INFORMATION RETRIEVAL DOCTOR DISSERTATION IN INFORMATION SYSTEM Hanoi, 2024 VIETNAM NATIONAL UNIVERSITY, HANOI UNIVERSITY OF ENGINEERING AND TECHNOLOGY VUONG THI HAI YEN MODELING AND LEARNING TEXTUAL AND STRUCTURAL RELATIONS FOR DEEP LEGAL INFORMATION RETRIEVAL Major: Information System Code: 9480104 DOCTOR DISSERTATION IN INFORMATION SYSTEM Supervised by: 1. Phan Xuan Hieu 2. Nguyen Le Minh Ph.D CANDIDATE SUPERVISORS VNU UNIVERSITY OF ENGINEERING AND TECHNOLOGY Hanoi, 2024 Abstract With the recent advances in digitalization and digital transformation, legal profes- sionals can now easily access a huge volume of online legal materials. This is extremely important because judges and lawyers frequently need to find relevant legal information when they are working on a new legal case, performing legal research, case analysis, court preparation, giving legal advice to a client, developing a defense strategy, or mak- ing decision on a current case.
However, the larger a legal database is, the more difficult for them to find relevant materials manually. In addition, legal documents like statutory law, case law or contract are normally lengthy and complex, consisting of multiple parts, chapters, sections, articles, and so on. Therefore, building an intelligent and automated legal information retrieval (IR) system is significant to improve and accelerate their legal process and workflow. Generally, this thesis aims to propose different legal IR methods and solutions based on an in-depth understanding of the nature and characteristics of legal data as well as the complexity of legal IR problems.
Accordingly, two major issues we need to consider carefully in this study are legal materials and legal IR problems. Legal materials are diverse, consisting of many differ- ent types of documents like constitution, statutory law, regulation, decision, case law, court document, contract, legal notice, patent, trademark, and so on. Among them, we focus on two main types of legal texts – statutory law and case law – because working on all types of legal materials is too broad and goes beyond the scope of the thesis. Regard- ing legal IR problems, this study focuses on three major IR tasks: (i) case law retrieval; (ii) statutory – case law retrieval; and (iii) IR–based legal question answering.
The first task locates and returns case law documents from a case law database that relate and entail the decision of an input legal case. The second task retrieves statutory laws from a statutory law database that are relevant to a query case. And the third task seeks and returns statutory law articles that are likely to contain answers to a given legal question. The three legal IR problems stated above are much more challenging than tradi- tional IR for general-domain texts.
The concept of relevancy in these tasks is no longer iii about keyword or topic matching. The similarity between legal texts requires the un- derstanding of legal arguments and logical reasoning that are far beyond the lexical or topical comparison. In addition, while working with legal data, we realized that legal language is rigorous and complicated. Legal documents are normally lengthy and heav- ily rely on domain-specific terminologies, jargons, and linguistic nuances.
Furthermore, there is a complex graphical structure hidden in any legal dataset that results from fre- quent mentions, citations, references within and between legal materials. Also, the style and content of legal documents highly depend on the domain and the legal system of each country. And one more important issue is that annotated data is limited because labeling for legal data requires a lot of human effort and domain expertise. All of these reasons are both the challenges as well as the motivations behind our study.
The main objective of this thesis is to enhance the performance and accuracy of the three legal IR problems by making the most of textual and structural relations in the legal data. First, we propose a supporting model that encodes both the lexical and legal relations at different levels of granularity to deal with the case law retrieval problem. In addition, we introduce a method to automatically create a large weak-labeling dataset to overcome the limitation of labeled data. Second, a heterogeneous legal knowledge graph was defined and constructed to leverage the statutory–case relationships in the statutory – case law retrieval.
Third, the thesis presents a novel approach that builds an article reference network to uncover both local and long-range dependencies between legal articles to enhance the performance of the IR–based legal question answering. Moreover, throughout the thesis, we propose appropriate deep learning architectures to encode the textual and structural characteristics of legal data and combine them with powerful pre-trained language models to enhance the overall performance of the three IR problems. Besides the technical contributions, the literature review, the analysis, and discussions throughout this thesis would provide a deeper and clearer understanding of the nature and the limitations in legal NLP in general and in legal IR in particular. It would also be a potential reference for future studies in the field, particularly for low- resource language like Vietnamese.
Keywords: statutory law, case law, legal case, deep legal information retrieval, legal question answering, case law retrieval, statutory – case law retrieval, IR–based legal question answering, legal case entailment, supporting model, weakly labeled data, relevancy, textual relation, structural relation, legal knowledge graph, article reference network, pre-trained language model. iv Acknowledgements Firstly, I would like to express my sincere gratitude to my thesis advisors Asocc. Phan Xuan Hieu and Prof. Nguyen Le Minh, for the continuous support of my Ph.
study and related researches, for their patience, motivation, and immense knowl- edge. Their guidance helped me in all the time of research. I could not have imagined having a better advisor and mentor for my Ph. Besides my advisors, I would like to acknowledge honorary Assoc.
Ha Quang Thuy and Dr. Nguyen Ha Thanh, for their insightful comments and encourage- ment. Without their precious support, it would not be possible to conduct this research. I also own special thanks to all members of the Data Science and Knowledge Technology Laboratory, The Department of Information Systems (VNU University of Engineering and Technology) and Nguyen’s Laboratory (School of Information Sci- ence, Japan Advanced Institute of Science and Technology), who have been a source of friendships as well as good advice and collaboration.
Finally, with all my love, I would like to thank my family for all their love and encouragement. Thank you! v Declaration I hereby declare that this Doctoral Dissertation was carried out by me for the degree of Doctor of Philosophy under the guidance and supervision of my supervisors. This dissertation is my own work and includes nothing, which is the outcome of work done in collaboration except as specified in the text. It is not substantially the same as any I have submitted for a degree, diploma or other qualification at any other university; and no part has already been, or is currently being submitted for any degree, diploma or other qualification.
Hanoi, May 2024 Author Vuong Thi Hai Yen vi Table of Contents ABSTRACT. vi TABLE OF CONTENTS. vii LIST OF ABBREVIATIONS. x LIST OF LEGAL TERMINOLOGIES.
xii LIST OF FIGURES. xiii LIST OF TABLES .1 Overview of Research Context and Challenges .2 Scope of Research .1 The Legal Data of Interest .2 The Deep Legal Information Retrieval Problems .3 Motivations and Objectives .2 Research Questions and Objectives. 22 2 LITERATURE REVIEW OF PROBLEMS AND METHODS .1 Legal Natural Language Processing .2 Case Law Retrieval .3 Statutory – Case Law Retrieval .4 IR–based Legal Question Answering .5 Representation of Legal Data .1 Textual Representation of Legal Data .2 Structural Representation of Legal Data .6 Information Retrieval Models .1 Traditional Information Retrieval Models .2 Deep Learning–based Retrieval Models. 47 3 SUPPORTING RELATION MODEL FOR CASE LAW RETRIEVAL .1 Case Law Supporting Relation .2 Supporting Relation in Case Law Retrieval .1 The Case Law Task in COLIEE Dataset .2 Weak-labeling Supporting Dataset .5 Case Law Retrieval with Supporting Model .2 Combination of Supporting Model and Lexical Model .3 Case Law Retrieval With Scoring Method .6 Experiments and Results .2 The Case Law Retrieval Results .3 The Case Law Entailment Results.
74 4 KNOWLEDGE GRAPH FOR STATUTORY – CASE LAW RETRIEVAL 75 4.1 Legal Knowledge Graph .2 Vietnamese Legal Case Knowledge Graph Definition .3 Knowledge Graph Construction .3 Knowledge Graph Deployment .4 Statutory – Case Law Retrieval Model .5 Experiments and Results. 86 5 ARTICLE REFERENCE NETWORK FOR IR–BASED LEGAL QUES- TION ANSWERING .1 The Article Reference Relation Network .2 Reference Network for IR-based Legal Question Answering .1 Reference Network Model .2 Trail-threshold Ranking .3 Experiments and Results .3 Vietnamese Legal Question Answering .2 Reference Network Approach .3 Supporting Relation for Automatic Data Enrichment Approach. 123 LIST OF PUBLICATIONS. 127 ix LIST OF ABBREVIATIONS Adam Adaptive Moment Estimation AI Artificial Intelligence ALQAC Automated Legal Question Answering Com- petition BERT Bidirectional Encoder Representations from Transformers biLSTM Bidirectional Long Short-term Memory BM25 BM25 Ranking Algorithm CCC Case-Court-Case CDC Case-Domain-Case CNN Convolutional Neural Network COLIEE The Competition on Legal Information Extrac- tion/Entailment DNN Deep Neural Network DSSM Deep Structured Semantic Model FN False Negative FP False Positive GCNs Graph Convolutional Networks GloVe Global Vectors for Word Representation GRNs Graph Nerual Networks IR Information Retrieval x KB Knowledge-base KG Knowledge Graph L2R Learning to Rank LLMs Large Transformer-based Language Models LSTM Long Short-term Memory MLP Multilayer Perceptron NeuIR Neural Information Retrieval NLP Natural Language Processing P Precision PROLEG PROlog-based LEGal reasoning support sys- tem QA Question Answering R Recall RNN Recurrent Neural Network SGD Stochastic Gradient Descent TF-IDF Term Frequency – Inverse Document Fre- quency TN True Negative TP True Positive VSM Vector-Space Model xi List of Legal Terminologies Terminology Meaning (in Vietnamese) article điều luật case law án lệ (luật dựa trên các lập luận, tiền lệ, phán quyết của các vụ án trước đó; bổ sung cho statutory law) code bộ luật (của luật thành văn – statutory law, statute law) constitution hiến pháp counsel luật sư; cố vấn pháp lý courtroom proceedings thủ tục xét xử defense strategy chiến lược bào chữa dispute tranh chấp; tranh luận judge thẩm phán; phán xét judgement phán xét judicial decision quyết định của toà án; quyết định tư pháp; bản án jurisdiction thẩm quyền tài phán; quyền hạn xét xử jury bồi thẩm đoàn lawmaker nhà lập pháp; người làm luật lawsuit vụ kiện lawyer luật sư legal argument tranh luận pháp lý; lập luận pháp lý legal case vụ án; vụ kiện legal filings hồ sơ pháp lý legal proceedings thủ tục tố tụng pháp lý legislation lập pháp; quá trình xây dựng và ban hành luật litigation kiện tụng rulings phán quyết statutory law (statute law) luật thành văn (được ban hành chính thức bởi chính quyền dưới dạng văn bản) trial phiên toà; phiên xử; việc xét xử; sự xử án xii List of Figures 1.1 Categories and tasks in legal natural language processing .2 A sample of Japanese statutory (civil) law .3 A sample of Vietanmese statutory (marriage and family) law .4 A case law sample from The Federal Court of Canada case law database .5 The logical flow of the case law retrieval problem .6 An example of the legal case entailment phase .7 The dissertation outline .1 The difference between Bi-Encoder and Cross-Encoder .1 Example of supporting component extraction between a query case and a candidate case .2 An example of supporting relation among sentences in the case law paragraph, each sentence in the paragraph is represented as a vertice, edges are semantic similarity between the sentences.
S1 is topic sentence in this example.3 Scoring method pipeline in supporting text-pair recognition task .4 A sample of a case law from The Federal Court of Canada case law database .5 The supporting model architecture.6 Example of supporting relation between a base case and candidate cases.