VIETNAM NATIONAL UNIVERISTY HO CHI MINH CITY HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY NGUYEN VINH KHIEM APPLICATION OF LARGE LANGUAGE MODEL IN TEXT-TO-SQL Major: COMPUTER SCIENCE Major code: 8480101 MASTER’S THESIS HO CHI MINH CITY, June 2024 THIS THESIS IS COMPLETED AT HO CHI MINH UNIVERSITY OF TECHNOLOGY – VNU-HCM Supervisors: Assoc. Huynh Tuong Nguyen Assoc. Quan Thanh Tho Examiner 1: Dr. Le Thanh Van Examiner 2: Dr.
Le Thi Thuy This master’s thesis is defended at Ho Chi Minh City University of Technology (HCMUT) – VNU-HCM on 17th June 2024. Master’s Thesis Committee: 1. Vo Thi Ngoc Chau Chairman 2. Le Thanh Van Examiner 1 3.
Le Thi Thuy Examiner 2 4. Tran Thanh Tung Commissioner 5. Phan Trong Nhan Secretary Approval of the Chairperson of the Master’s Thesis Committee and Dean of Faculty of Computer Science and Engineering after the thesis being corrected (If any). CHAIRPERSON OF DEAN OF FACULTY OF THESIS COMMITTEE COMPUTER SCIENCE AND ENGINEERING VIETNAM NATIONAL UNIVERSITY HO CHI MINH CITY VIETNAM NATIONAL UNIVERSITY - HO CHI MINH CITY SOCIALIST REPUBLIC OF VIETNAM HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY Independence – Freedom - Happiness THE TASK SHEET OF MASTER’S THESIS Full name: Nguyen Vinh Khiem Student ID: 2270162 Date of birth: 11/05/1997 Place of birth: Ho Chi Minh Major: Computer Science Major ID: 8480101 I.
THESIS TITLE (in English): Application of large language models in Text-to-SQL II. THESIS TITLE (in Vietnamese): Ứng dụng mô hình ngôn ngữ lớn trong việc tạo câu truy vấn III. TASKS AND CONTENTS: a. Research and design a model capable of generating SQL queries from text.
Implement, test and evaluate model. THESIS START DAY: 15/01/2024 V. THESIS COMPLETION DAY: 20/05/2024 VI. Huynh Tuong Nguyen 2.
Quan Thanh Tho Ho Chi Minh City, date 05/08/2024 SUPERVISOR SUPERVISOR CHAIRMAN OF PROGRAM (Full name and signature) (Full name and signature) COMMITTEE (Full name and signature) DEAN OF FACULTY OF COMPUTER SCIENCE AND ENGINEERING (Full name and signature) i VIETNAM NATIONAL UNIVERSITY - HO CHI MINH CITY SOCIALIST REPUBLIC OF VIETNAM HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY Independence – Freedom - Happiness ACKNOWLEDGEMENTS This thesis could not have been completed without significant support from various individuals and groups. I am profoundly grateful to my primary advisors, Assoc. Quan Thanh Tho and Assoc. Huynh Tuong Nguyen, who has been a constant source of guidance, providing necessary resources and assistance throughout my research, and offering support whenever I faced challenges.
I wish to express my profound gratitude to the esteemed professors and lecturers of the Department of Computer Science and Engineering, and the Ho Chi Minh City University of Technology at large. The knowledge they imparted is priceless and has been instrumental in the completion of this thesis. I am also thankful to my colleagues at GiaoHangNhanh Company for granting me the chance to engage deeply in research and improve my professional expertise, alongside providing resources essential for training my models. Lastly, I owe a deep gratitude to my family, friends and classmate, all of whom have been supportive, encouraging, and provided the emotional and physical support needed to complete this thesis.
With heartfelt gratitude, I wish good health and all the best to the professors and lecturers of the Department of Computer Science and Engineering at the Ho Chi Minh City University of Technology, National University of Ho Chi Minh City. ii VIETNAM NATIONAL UNIVERSITY - HO CHI MINH CITY SOCIALIST REPUBLIC OF VIETNAM HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY Independence – Freedom - Happiness ABSTRACT The rapid advancements in natural language processing (NLP) have been significantly driven by large language models (LLMs), which have demonstrated impressive capabilities in understanding and generating human-like text. This thesis explores the application of LLMs, specifically the Flan-T5 model, in the context of the Text-to-SQL task, which aims to translate natural language queries into structured SQL commands. This translation is crucial for enhancing data accessibility, allowing users without SQL expertise to interact with relational databases effectively.
The proposed solution utilizes a two-tier architecture comprising a generator and a ranker model. The generator, based on the Flan-T5 model, generates multiple SQL query candidates from natural language inputs. These candidates are then evaluated and ranked by the ranker model to select the most accurate query. The approach leverages a constrained decoding technique guided by SQL grammar rules to ensure the syntactic validity of the generated queries.
iii VIETNAM NATIONAL UNIVERSITY - HO CHI MINH CITY SOCIALIST REPUBLIC OF VIETNAM HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY Independence – Freedom - Happiness TÓM TẮT LUẬN VĂN THẠC SĨ Sự tiến bộ nhanh chóng trong xử lý ngôn ngữ tự nhiên (NLP) đã được thúc đẩy đáng kể bởi các mô hình ngôn ngữ lớn (LLMs), những mô hình đã cho thấy khả năng ấn tượng trong việc hiểu và tạo ra văn bản giống con người. Luận văn này khám phá việc áp dụng các LLMs, cụ thể là mô hình Flan-T5, trong bối cảnh nhiệm vụ tạo câu truy vấn từ câu hỏi, nhằm dịch các truy vấn ngôn ngữ tự nhiên thành các câu lệnh SQL có cấu trúc. Việc sinh câu truy vấn này rất quan trọng để nâng cao khả năng truy cập dữ liệu, cho phép người dùng không có chuyên môn về ngôn ngữ SQL tương tác hiệu quả với các cơ sở dữ liệu quan hệ. Giải pháp đề xuất sử dụng kiến trúc hai tầng bao gồm một mô hình tạo và một mô hình xếp hạng.
Mô hình tạo, dựa trên mô hình Flan-T5, tạo ra nhiều câu truy vấn SQL từ các đầu vào ngôn ngữ tự nhiên. Các câu truy vấn này sau đó được đánh giá và xếp hạng bởi mô hình xếp hạng để chọn ra truy vấn chính xác nhất. Cách tiếp cận này tận dụng kỹ thuật giải mã bị ràng buộc theo quy tắc ngữ pháp SQL để đảm bảo tính chính xác về cú pháp của các truy vấn được tạo ra. iv VIETNAM NATIONAL UNIVERSITY - HO CHI MINH CITY SOCIALIST REPUBLIC OF VIETNAM HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY Independence – Freedom - Happiness DECLARATION OF AUTHORSHIP I solemnly affirm that the thesis titled: APPLICATION OF LARGE LANGUAGE MODEL IN TEXT-TO-SQL is the product of my own research endeavors.
The documentation used in this thesis has been clearly stated in the References section. The data and results presented in this thesis are entirely truthful, and I am fully responsible for any inaccuracies and will accept any discipline set forth by the department and the university SUPERVISOR SUPERVISOR STUDENT (Full name and signature) (Full name and signature) (Full name and signature) v Contents 1 Topic Introduction 1 1.2 Overview about Text-to-SQL .4 Scope of thesis .1 RAT-SQL: Relation-Aware Schema Encoding and Linking for Text- to-SQL Parsers .2 Graphix-T5: Mixing Pre-Trained Transformers with Graph-Aware Layers for Text-to-SQL Parsing .3 T5QL: Taming language models for SQL generation .2 SQL Grammar for constrain decoding .1 Recurrent Neural Networks (RNNs) .2 Feed Forward Network .3 Pre-trained Language Model .1 GPT - Generative Pretrained Transformer. 27 vi Ho Chi Minh University of Technology Faculty of Computer Science and Engineering 3.2 BERT - Bidirectional Encoder Representations from Trans- formers .3 T5/Flan-T5: Text-to-Text Transfer Transformer .2 Issues and Challenges. 44 vii List of Figures 1.1 Percentage of Programming Language used by Professional Devel- opers .2 Text-to-SQL problem[2] .1 The technique taxonomy for text-to-SQL .2 Visualization of RAT-SQL model[5] .3 Relationship between members in schema .4 Visualization of Graphix-T5 model[7] .5 Example of Multi-hop relation between nodes .6 Visualization of No-Match and Bridge Node Mode .7 T5QL model architecture[8] .8 Pseudo code for Constrained Decoding .9 SQL Grammar Rule .1 The architecture of a recurrent neural network layer is represented with shorthand notation (left) and represented with a hidden state (right)[10] .2 Visualization of Transformer architecture[6] .3 Visualization of Scaled Dot-Product Attention .4 Multi-Head Attention consists of several attention layers running in parallel .5 Feed Forward Network .6 Overview of some popular LLMs based on Transformers[11] .7 Architecture of GPT model[12].
28 viii Ho Chi Minh University of Technology Faculty of Computer Science and Engineering 3.8 Input transformations for fine-tuning on different tasks[12] .9 The overview of BERT Architecture [14] .10The overview of BERT Architecture .11Overview of Flan-T5 finetuning data and task[3] .1 Architecture of proposed model. 36 ix List of Tables 4.1 ROUGE Score for 2 circumstances. 40 x Chapter 1 Topic Introduction 1.1 General Introduction The progression in the field of natural language processing has been sig- nificantly accelerated with the advent of large language models (LLMs). Models like GPT-3, BERT, and their successors have drastically improved our proficiency in processing, understanding, and generating text that is remarkably human-like.
These models have been meticulously trained on vast collections of text, which has endowed them with a nuanced understanding of language. This breakthrough has laid a foundation for pioneering applications in several linguistic tasks, repre- senting a formidable leap in technology that has transformed the way we interact with machines. SQL’s role in managing and analyzing data within relational databases is indisputably vital in our modern data-centric world. The ubiquity of SQL across various sectors underscores its importance for organizing and retrieving critical data.According to the yearly survey conducted by StackOverflow[1], SQL main- tains its status as one of the globally dominant languages.
It is observed that among the technologies professionals most frequently utilize, JavaScript, HTM- 1 Ho Chi Minh University of Technology Faculty of Computer Science and Engineering L/CSS, and SQL emerge as the top three, with JavaScript and HTML/CSS nearly reaching parity as the leading languages for coding novices Figure 1.1: Percentage of Programming Language used by Professional Developers[1] The task of converting natural language into SQL commands, known as Text-to-SQL, has gained prominence. It grants non-experts the ability to access database information, significantly broadening the scope of data utility and facil- itating informed decision-making across diverse user groups. While LLMs hold the potential to simplify the interaction between natural language and SQL queries, the task of Text-to-SQL generation comes with dis- tinct challenges. LLMs need to acquire a profound semantic understanding of the queries, efficiently generate SQL commands, and interpret the users’ intent with high accuracy.
The intricacies involved in SQL’s structure and the variable na- ture of natural language queries add layers of complexity to this task. Integrating LLMs into Text-to-SQL systems is a complex endeavor that goes beyond techni- cal challenges. It requires the selection of suitable models, rigorous experimental design, and the development of reliable metrics to gauge performance. In addi- tion, it is imperative to consider the wider implications on user experience and 2 Ho Chi Minh University of Technology Faculty of Computer Science and Engineering database functionality, striving towards a solution that is not only seamless and efficient but also scalable.2 Overview about Text-to-SQL In the current era, data has become a critical asset essential for a wide range of human endeavors, encompassing both commercial activities and scien- tific investigations.
However, the burgeoning volume and escalating intricacy of data present significant challenges in its querying and exploration, even for those with expertise in the field. Present-day data query interfaces are generally bifur- cated into two categories: form-based interfaces, which are user-friendly but offer constrained querying capabilities, and more advanced, low-level tools. These ad- vanced tools permit the synthesis of queries in native database languages, such as SQL, but are primarily designed for a specialized audience, like SQL profes- sionals. To democratize data access and utilization, ensuring that everyone can effectively engage with, comprehend, and extract value from data, it is crucial to remove the technical obstacles that hinder data accessibility and reduce reliance on IT specialists.
Adopting natural language for query expression can democra- tize data accessibility. In this vein, there is a growing scholarly interest in the development of Nat- ural Language (NL) Interfaces for Databases (NLIDBs).