VIETNAM NATIONAL UNIVERSITY HO CHI MINH CITY HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY TANG QUOC THAI APPLICATION OF LARGE LANGUAGE MODELS IN SOFTWARE ERROR DEBUGGING Major: COMPUTER SCIENCE Major code: 8480101 MASTER’S THESIS HO CHI MINH CITY, June 2024 VIETNAM NATIONAL UNIVERSITY HO CHI MINH CITY HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY TANG QUOC THAI APPLICATION OF LARGE LANGUAGE MODELS IN SOFTWARE ERROR DEBUGGING Major: COMPUTER SCIENCE Major code: 8480101 MASTER’S THESIS HO CHI MINH CITY, June 2024 THIS THESIS IS COMPLETED AT HO CHI MINH UNIVERSITY OF TECHNOLOGY – VNU-HCM Supervisors: Assoc. Huynh Tuong Nguyen Assoc. Quan Thanh Tho Examiner 1: Dr. Truong Tuan Anh Examiner 2: Dr.
Tran Thanh Tung This master’s thesis is defended at Ho Chi Minh City University of Technology (HCMUT) – VNU-HCM on 17/06/2024. Master’s Thesis Committee: 1. Vo Thi Ngoc Chau Chairperson 2. Truong Tuan Anh Examiner 1 3.
Tran Thanh Tung Examiner 2 4. Le Thi Thuy Commissioner 5. Phan Trong Nhan Secretary Approval of the Chairperson of the Master’s Thesis Committee and Dean of Faculty of Computer Science and Engineering after the thesis being corrected (If any). CHAIRPERSON OF DEAN OF FACULTY OF THESIS COMMITTEE COMPUTER SCIENCE AND ENGINEERING VIETNAM NATIONAL UNIVERSITY - HO CHI MINH CITY SOCIALIST REPUBLIC OF VIETNAM HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY Independence – Freedom - Happiness THE TASK SHEET OF MASTER’S THESIS Full name: Tang Quoc Thai Student ID: 2270376 Date of birth: 02/12/1991 Place of birth: Dong Nai Major: Computer Science Major ID: 8480101 I.
THESIS TITLE (in English): Application of large language models in software error debugging II. THESIS TITLE (in Vietnamese): Ứng dụng mô hình ngôn ngữ lớn trong việc soát lỗi phần mềm III. TASKS AND CONTENTS: a. Construct a self-correcting code model by combining large language models (LLMs).
Research and propose methods to improve the accuracy of the model. Experiment and evaluate the results of the proposed methods. THESIS START DAY: 15/01/2024 V. THESIS COMPLETION DAY: 20/05/2024 VI.
Huynh Tuong Nguyen 2. Quan Thanh Tho Ho Chi Minh City, date 05/08/2024 SUPERVISOR SUPERVISOR CHAIRPERSON OF (Full name and signature) (Full name and signature) PROGRAM COMMITTEE (Full name and signature) DEAN OF FACULTY OF COMPUTER SCIENCE AND ENGINEERING (Full name and signature) i VIETNAM NATIONAL UNIVERSITY - HO CHI MINH CITY SOCIALIST REPUBLIC OF VIETNAM HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY Independence – Freedom - Happiness ACKNOWLEDGEMENTS The completion of this thesis was made possible by the significant support I received from numerous individuals. I am deeply grateful to my primary supervisors, Assoc. Quan Thanh Tho and Assoc.
Huynh Tuong Nguyen. Their guidance, provision of resources, and continuous monitoring of this topic's progress were invaluable. They were always ready to help when I faced challenges. Above all, their passion for machine learning, deep learning, natural language processing, and various other aspects of Computer Science has been a source of inspiration.
I wish to express my profound gratitude to the esteemed professors and lecturers of the Department of Computer Science and Engineering, and the Ho Chi Minh City University of Technology at large. The knowledge they imparted is priceless and has been instrumental in the completion of this thesis. I would also like to extend my appreciation to the startup PixelML for your generous contribution of AWS credits, which facilitated the use of several LLMs for this project. Lastly, I am thankful for my family, relatives, and friends.
Their concern, encouragement, and support, both emotionally and physically, provided me with the strength and resilience needed to successfully complete this thesis. With heartfelt gratitude, I wish good health and all the best to the professors and lecturers of the Department of Computer Science and Engineering at the Ho Chi Minh City University of Technology, National University of Ho Chi Minh City. ii VIETNAM NATIONAL UNIVERSITY - HO CHI MINH CITY SOCIALIST REPUBLIC OF VIETNAM HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY Independence – Freedom - Happiness ABSTRACT This thesis explores the Application of Large Language Models in Software Error Debugging, focusing on the use of self-correcting Large Language Models (LLMs) to assist in software bug resolution. The study introduces and evaluates a novel method that combines the Chain-of-Thought (CoT) prompting technique and a code evolution framework, specifically applied to Python data science problems.
This approach, referred to as the CoT- SelfEvolve model, is evaluated against several benchmarks and compared with other published methods. The CoT-SelfEvolve model demonstrates superior performance, particularly with Pytorch, Sklearn, Matplotlib, Pandas, Numpy, and Tensorflow tasks, highlighting the potential of LLMs in automated debugging. However, areas for improvement were identified, such as underperformance with Scipy questions and the need for further exploration of LLM parameters' effects on performance. The study concludes that the CoT-SelfEvolve model is a promising step towards more efficient and effective automated debugging, and it paves the way for future research in this domain.
iii VIETNAM NATIONAL UNIVERSITY - HO CHI MINH CITY SOCIALIST REPUBLIC OF VIETNAM HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY Independence – Freedom - Happiness TÓM TẮT LUẬN VĂN THẠC SĨ Luận văn này khám phá Ứng dụng của Các Mô Hình Ngôn Ngữ Lớn trong Soát Lỗi Phần Mềm, tập trung vào việc sử dụng các Mô Hình Ngôn Ngữ Lớn (LLMs) tự sửa lỗi để hỗ trợ giải quyết lỗi phần mềm. Nghiên cứu giới thiệu và đánh giá một phương pháp mới kết hợp kỹ thuật gợi ý Chuỗi Tư Duy (CoT) và khung phát triển mã, được áp dụng cụ thể cho các vấn đề khoa học dữ liệu Python. Phương pháp này, được gọi là mô hình CoT-SelfEvolve, được đánh giá dựa trên nhiều tiêu chuẩn và so sánh với các phương pháp đã được công bố khác. Mô hình CoT-SelfEvolve cho thấy hiệu suất vượt trội, đặc biệt với các tác vụ Pytorch, Sklearn, Matplotlib, Pandas, Numpy và Tensorflow, nhấn mạnh tiềm năng của LLMs trong gỡ lỗi tự động.
Tuy nhiên, các lĩnh vực cần cải thiện đã được xác định, chẳng hạn như hiệu suất kém với các câu hỏi Scipy và cần khám phá thêm về ảnh hưởng của các tham số LLM đến hiệu suất. Nghiên cứu kết luận rằng mô hình CoT-SelfEvolve là một bước tiến đầy hứa hẹn hướng tới gỡ lỗi tự động hiệu quả và hiệu suất hơn, và mở đường cho các nghiên cứu tương lai trong lĩnh vực này. iv VIETNAM NATIONAL UNIVERSITY - HO CHI MINH CITY SOCIALIST REPUBLIC OF VIETNAM HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY Independence – Freedom - Happiness DECLARATION OF AUTHORSHIP I solemnly affirm that the thesis titled: APPLICATION OF LARGE LANGUAGE MODELS IN SOFTWARE ERROR DEBUGGING is the product of my own research endeavors. The data and findings presented in this thesis are entirely accurate to the best of my knowledge, and I accept full responsibility for any inaccuracies.
I am prepared to face any disciplinary actions that may be imposed by the department or the university in case of any discrepancies. SUPERVISOR SUPERVISOR STUDENT (Full name and signature) (Full name and signature) (Full name and signature) v Contents 1 Topic Introduction 1 1.2 Self-Correcting LLMs for Software Development .1 A Taxonomy for Self-Correcting LLMs with Automated Feedback 8 2.2 What is the source of the feedback? .3 What is the format of the feedback? .4 When to correct the model with feedback? .5 How to correct the model with feedback? .2 Training-Time Correction .1 Learning from Human Feedback .2 Learning with Automated Feedback .3 Generation-Time Correction .1 Generate-then-Rank .2 Feedback-Guided Decoding .4 Post-hoc Correction .2 Models/Tools as Feedback .3 Multi-Agent Debate. 26 vi Ho Chi Minh University of Technology Faculty of Computer Science and Engineering 2.5 Direction of Current Research .2 Attention in LLMs .2 Code Evolution Framework .1 Auto-CoT Prompt Generator 1 .2 Auto-CoT Prompt Generator 2. 54 vii Ho Chi Minh University of Technology Faculty of Computer Science and Engineering 5.3 Results and Discussions .a The importance of Auto-CoT prompt generators .b Using larger LLM to guide the code generation process .2 Things to Improve.
60 viii List of Figures 1.1 Distribution of time to completion between treated and control groups [15].2 A prevalent approach for iterative debugging employing a large lan- guage model involves multiple steps [16].1 A conceptual framework for self-correcting LLMs with automated feedback.2 Focus areas of the current research.1 Attention patterns in a causal decoder, non-causal decoder, and encoder-decoder.1 Example of CoT reasoning processes.2 An example problem of Matplotlib.3 A discussion on StackOverflow.4 Architecture of SelfEvolve.5 Architecture of CoT-SelfEvolve with Auto-CoT prompt generators.6 Architecture of Auto-CoT Prompt Generator 1.7 Architecture of Auto-CoT Prompt Generator 2. 55 ix List of Tables 4.1 Pass@1 results on the DS-1000 dataset.1 Pass@5 results on the DS-1000 dataset.2 Pass@5 results on the DS-1000 dataset with or without CoT prompts.3 Pass@5 results on the DS-1000 dataset with different LLM stacks. 57 x Chapter 1 Topic Introduction 1.1 General Introduction In recent years, generative AI models, particularly Large Language Models (LLMs), have garnered significant attention from both the Artificial Intelligence (AI) research community and the general public. These models exhibit a re- markable capacity to address a diverse array of intricate language-based tasks.
Their advancements are propelled by factors such as increased model parameter count, augmented training data volume, and refined training configurations [1]– [4]. Prominent LLMs like LaMDA [5] and GPT-4 [6] demonstrate exceptional proficiency in applications ranging from translation and classification to creative writing and code generation. Such capabilities, previously necessitating task- specific models developed by domain experts using specialized data, are now achieved by these broad, state-of-the-art LLMs. Concurrently, researchers have enhanced the steerability, reliability, and utility of these models through techniques like fine-tuning and reinforcement learning with human feedback [7], [8].
These advancements empower models to better understand user intent, thereby enhancing user-friendliness and prac- 1 Ho Chi Minh University of Technology Faculty of Computer Science and Engineering ticality. Recent studies also showcase LLMs’ potential to program and control other digital tools, such as APIs, search engines, and even fellow generative AI systems [9], [10]. This integration of individual components facilitates improved utility, performance, and generalization. At the forefront of these trends, there lies a prospect where LLMs can potentially execute any task traditionally per- formed at a computer.
While generative AI models have primarily been deployed as modular spe- cialists for tasks like image generation from captions or text transcription from speech, the focus is on viewing LLMs as versatile building blocks for creating additional tools. The development and integration of these tools into systems may necessitate time and significant reconfiguration of existing processes across diverse industries. Nevertheless, early adoption trends are already emerging. De- spite their limitations, LLMs are increasingly finding integration into specialized applications in areas such as writing assistance, coding, and legal research.
These specialized applications enable businesses and individuals to incorporate LLMs into their existing workflows. The emphasis is on the significance of these complementary technologies, particularly because standalone general-purpose LLMs may still exhibit unrelia- bility for certain tasks, attributable to issues such as factual inaccuracies, inher- ent biases, privacy concerns, and risks associated with disinformation [11]–[13]. However, specialized workflows, encompassing tooling, software, or human-in- the-loop systems, can effectively mitigate these shortcomings by incorporating domain-specific expertise. For instance, Casetext provides LLM-based legal re- search tools that furnish lawyers with quicker and more accurate legal research results, utilizing embeddings and summarization to counteract the risk of GPT-4 potentially providing inaccurate details about a legal case or set of documents.
GitHub Copilot, a coding assistant, leverages LLMs to generate code snippets and auto-complete code, allowing users to accept or reject suggestions based on their expertise. In essence, while GPT-4, on its own, might not inherently ‘know what time it is’, incorporating a watch can address this limitation.