VIETNAM NATIONAL UNIVERSITY HO CHI MINH CITY HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY -------------------- NGUYEN NGOC HAI DANG APPLICATION OF MACHINE LEARNING ON AUTOMATIC PROGRAM REPAIR OF SECURITY VULNERABILITIES Major: Computer Science Major code: 8480101 MASTER’S THESIS HO CHI MINH CITY, July 2023 THIS THESIS IS COMPLETED AT HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY – VNU-HCM Supervisor: Assoc. Huynh Tuong Nguyen, Assoc. Quan Thanh Tho Examiner 1: Dr. Truong Tuan Anh Examiner 2: Assoc.
Nguyen Van Vu This master’s thesis is defended at HCM City University of Technology, VNU- HCM City on July 11,2023 Master’s Thesis Committee: (Please write down full name and academic rank of each member of the Master’s Thesis Committee) 1. Le Hong Trang 2. Phan Trong Nhan 3. Truong Tuan Anh 4.
Nguyen Van Vu 5. Nguyen Tuan Dang Approval of the Chairman of Master’s Thesis Committee and Dean of Faculty of Computer Science and Engineering after the thesis being corrected (If any). CHAIRMAN OF THESIS COMMITTEE HEAD OF FACULTY OF COMPUTER SCIENCE AND ENGINEERING VIETNAM NATIONAL UNIVERSITY - HO CHI MINH CITY SOCIALIST REPUBLIC OF VIETNAM HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY Independence – Freedom - Happiness THE TASK SHEET OF MASTER’S THESIS Full name: Nguyen Ngoc Hai Dang Student ID: 1970513 Date of birth: 24/11/1997 Place of birth: Lam Dong Major: Computer Science Major ID: 8480101 I. THESIS TITLE: ỨNG DỤNG HỌC MÁY VÀO CHƯƠNG TRÌNH TỰ ĐỘNG SỬA CHỮA LỖ HỔNG BẢO MẬT - APPLICATION OF MACHINE LEARNING ON AUTOMATIC PROGRAM REPAIR OF SECURITY VULNERABILITIES II.
TASKS AND CONTENTS: - Research and build a system to automatically repair vulnerabilities - Research and propose methods to improve the accuracy of the model. - Experiment and evaluate the results of the proposed methods. THESIS START DAY: 05/09/2022 IV. THESIS COMPLETION DAY: 09/06/2023 V.
Huynh Tuong Nguyen, Assoc. Quan Thanh Tho Ho Chi Minh City, date ……… SUPERVISOR CHAIR OF PROGRAM COMMITTEE (Full name and signature) (Full name and signature) DEAN OF FACULTY OF COMPUTER SCIENCE AND ENGINEERING (Full name and signature) Acknowledgement I would like to acknowledge the people who have helped me with their knowledge, encouragement, and patience during the work of this thesis. The thesis would not have been completed without your help and inspiration. First and foremost, I would like to thank my supervisor at Ho Chi Minh City University of Technology, Professor Quan Thanh Tho.
Thank you for your unwavering support. Your insightful feedback and contributions have pushed and guided me throughout the work of this thesis. I would also like to thank my other supervisor at the Norwegian University of Science and Technology, Professor Nguyen Duc Anh. Thank you for your feedback and help.
Lastly, I would like to thank my friends and family for their endless patience, support, and encouragement. i Abstract We have, as individuals and as a society, become increasingly more dependent on software, thus, the consequences of failing software have also become greater. Identi- fying the failing parts of the software and fixing these parts manually could be time- consuming, expensive, and frustrating. The growing research field of automated code repair aims to tackle this problem, by applying machine learning techniques to be able to repair software in an automated fashion.
With the abundance of data of bugs and patches, research on the use of deep learning in code repairing has been on the rise and proven to be effective with the appearances of many systems [1] [2] with state of the art performance. However, the approach is conditioned on a large dataset to be applicable and this condition can not be met by all types of bugs in applications. One type of such bugs is vulnerability, which is the target of security exploitation of attackers to cause great harm to organizations that use the applications. Therefore, the need to automatically identify and fix vulnerabilities is obvious and can significantly reduce the harm that can be caused to these organizations.
In our work, we focus on the application of deep learning in vulnerability repair- ing and experiment with a solution that can be used to handle the lack of data, which is a requirement for deep learning models to be applied effectively, through the use of embeddings extracted from large language models like CodeBERT [3] and Unix- Coder [4]. Although our results show such an approach does not bring significant improvement, they can be used by other researchers to gain more insights into the proximity between the repairing tasks of different types of bugs. ii Tóm tắt luận văn Chúng ta, với tư cách cá nhân và xã hội, ngày càng trở nên phụ thuộc nhiều hơn vào phần mềm, do đó, hậu quả của việc phần mềm bị lỗi cũng trở nên lớn hơn. Việc xác định các phần bị lỗi của phần mềm và sửa các phần này theo cách thủ công có thể tốn thời gian, tốn kém và gây khó chịu.
Lĩnh vực nghiên cứu sửa chữa mã tự động đang phát triển nhằm mục đích giải quyết vấn đề này, bằng cách áp dụng các kỹ thuật máy học để có thể sửa chữa phần mềm theo cách tự động. Với lượng dữ liệu dồi dào về lỗi và bản vá lỗi, nghiên cứu về việc sử dụng học sâu trong sửa mã ngày càng nhiều và được chứng minh là hiệu quả với sự xuất hiện của nhiều hệ thống [1] [2] với công nghệ tiên tiến nhất biểu diễn nghệ thuật. Tuy nhiên, cách tiếp cận này dựa trên một tập dữ liệu lớn và không phải mọi loại lỗi trong ứng dụng cũng đáp ứng được điều kiện này, một trong số đó là lỗ hổng bảo mật, vốn là mục tiêu khai thác bảo mật của những kẻ tấn công nhằm gây hại cho các tổ chức sử dụng các ứng dụng chứa những lỗ hổng này. Do đó, nhu cầu tự động xác định và sửa các lỗ hổng này là hiển nhiên và có thể đem giảm đáng kể những thiệt hại có thể xảy ra cho các doanh nghiệp Trong luận văn này, chúng tôi tập trung vào việc ứng dụng học sâu trong việc khắc phục lỗ hổng bảo mật và thử nghiệm một giải pháp có thể sử dụng để xử lý tình trạng thiếu dữ liệu, vốn là yêu cầu để các mô hình này trở nên hiệu quả, thông qua việc sử dụng các embeddings được trích xuất từ những mô hình ngôn ngữ lớn như CodeBERT [3] và UnixCoder [4].
Mặc dù kết quả của chúng tôi cho thấy cách tiếp cận như vậy không mang lại sự cải thiện đáng kể, nhưng chúng vẫn có thể được các nhà nghiên cứu khác sử dụng để hiểu rõ hơn về khoảng cách giữa các nhiệm vụ sửa chữa của các loại lỗi khác nhau. iii Declaration I, Nguyen Ngoc Hai Dang, declare this thesis with the Vietnamese title as "Ứng dụng của học máy vào chương trình tự động sửa chữa lỗ hổng bảo mật” and English title as "Application of machine learning on automatic program repair of security vulnerabil- ities”, is my own work and contains no material that has been submitted previously, in whole or in part, for the award of any other academic degree or diploma. Signature Nguyen Ngoc Hai Dang iv CONTENTS CONTENTS Contents 1 Introduction 1 1.1 Background on Neural Network and Deep Learning .1 Recurrent Neural Network (RNN) .2 Vanilla recurrent neural network .3 Long short-term memory network(LSTM) .4 Transformer Neural Network .1 Sequence to Sequence Learning .2 Graphs-based Learning .3 Tree-to-tree Learning .4 Bug Repairing and Vulnerabilities Repairing .5 Source code Representation .2 Byte Pair Encoding .6 Source code embeddings. 25 3 The state of the art program repair approraches 27 3.1 Template-based approach .2 Generative-based approach.
32 4 Proposed Methods 34 5 Experiments and Results 35 5.2 Metrics of performance .3 Preprocessing the code as plain text .4 Extracting embeddings from large language models for code. 40 6 Discussions and Conclustion 43 6.1 Discussions of the results. 44 References 45 Appendix 49 vi LIST OF FIGURES LIST OF FIGURES List of Figures 2.1 The basic architecture of recurrent neural network .2 Recurrent Neural Network design patterns .3 LSTM network with three repeating layers .4 Attention-integrated recurrent network .5 The encoder-decoder architecture of transformer .6 Attention head operations .7 Dataset used for CodeBERT .8 CodeBERT architecture for replaced tokens detection task .9 A Python code with its comment and AST .10 Input for contrastive learning task of UnixCoder .1 Workflow of VuRLE .2 Architecture of SeqTrans .3 Input of SeqTrans .4 Normalized code segment .5 The VRepair pipeline .1 Design of our pipeline .1 Sample of buggy code and its patch .4 Syntax of the output sequence. 38 vii LIST OF TABLES LIST OF TABLES List of Tables 5.1 Experiments replicating the VRepair pipeline .2 Experiments with embeddings as input .1 Complete set of hyperparameters used in our models built by Opennmt- py.
49 viii Page 1 of 49 1 Introduction In the modern society, software systems play a crucial role in almost every aspect of our lives [5]. These systems have become the backbone of our interconnected world, enabling us to communicate, work, learn, and entertain ourselves efficiently and effectively [6]. From mobile applications and social media platforms to e-commerce websites and financial systems, software systems have revolutionized the way we interact, transact, and navigate the digital landscape. They have transformed industries, streamlined processes, and empowered individuals by providing access to information and services at our fingertips.
The importance of software systems lies in their ability to automate tasks, enhance productivity, enable innovation, and foster connectivity on a global scale. They have become indispensable tools for businesses, governments, healthcare, education, and countless other sectors, driving progress, enabling efficiency, and shaping the future. With the increasing reliance on software systems for critical functions, such as communication, finance, healthcare, and infrastructure, ensuring the security of these systems is of paramount importance. Software security involves protecting software applications and data from unauthorized access, breaches, and malicious activities.
The consequences of software security breaches can be severe, ranging from financial loss and reputational damage to compromised privacy and even threats to national security. While detecting software security issues can be done during and after software release, addressing security issues early in the development process saves time and resources. It is generally easier and less costly to fix vulnerabilities during the development stage than after the software has been deployed and is in active use. Software security detection involves using various techniques and tools to identify vulnerabilities, weaknesses, and potential threats within the codebase.
This can include static code analysis, dynamic testing, penetration testing, and security auditing. By actively searching for security issues, developers can uncover and address potential flaws before the software is deployed, reducing the risk of exploitation by malicious actors. Once security issues are detected, code repair comes into play. It involves remediation efforts to fix the identified vulnerabilities and weaknesses.
This may involve patching code, implementing security controls, updating dependencies, or improving the overall design of the software. Code repair is a critical step in Page 2 of 49 mitigating security risks and ensuring that the software meets the necessary security standards. In our research, we will explore code repair for software security issues. We will empirically investigate the application of Deep Learning to create patches of such vulnerabilities in software in an automatic manner.
The contribution of this thesis are three folds: • A literature review on state-of-the-art Deep Learning on code repair for software security • Experiments with different DL approaches for vulnerable code repair 1.1 Motivation In the field of software testing, security vulnerabilities are the type of bugs that are both hard to detect and implement patches as they are not explicitly affecting the software functionalities but they only exposed and cause great harm when exploited intentionally.