VIETNAM NATIONAL UNIVERSITY HO CHI MINH CITY HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY FACULTY OF COMPUTER SCIENCE AND ENGINEERING COMPUTER SCIENCE PROJECT VIDEO-BASED PARKING SPACE DETECTION Major: Computer Science THESIS COMMITTEE: CLC KHMT 2 SUPERVISOR(s): NGUYỄN THANH BÌNH MEMBER SECRETARY: PHAN TRỌNG NHÂN STUDENT: NGUYỄN TẤN TÀI (1852725) HO CHI MINH CITY, 2/2023 (9/1/2023) VIETNAM NATIONAL UNIVERSITY HO CHI MINH CITY HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY FACULTY OF COMPUTER SCIENCE AND ENGINEERING COMPUTER SCIENCE PROJECT VIDEO-BASED PARKING SPACE DETECTION Major: Computer Science THESIS COMMITTEE: CLC KHMT 2 SUPERVISOR(s): NGUYỄN THANH BÌNH MEMBER SECRETARY: PHAN TRỌNG NHÂN STUDENT: NGUYỄN TẤN TÀI (1852725) HO CHI MINH CITY, 2/2023 (9/1/2023) ĐẠI HỌC QUỐC GIA TP.HCM CỘNG HÒA XÃ HỘI CHỦ NGHĨA VIỆT NAM ---------- Độc lập - Tự do - Hạnh phúc TRƯỜNG ĐẠI HỌC BÁCH KHOA KHOA: KH & KT Máy tính ___ NHIỆM VỤ LUẬN ÁN TỐT NGHIỆP BỘ MÔN: HTTT _____________ Chú ý: Sinh viên phải dán tờ này vào trang nhất của bản thuyết trình Họ và tên SV: Nguyễn Tấn Tài – 1852725 Ngành (chuyên ngành): Khoa học Máy Tính 1. Đầu đề luận án: VIDEO-BASED PARKING SPACE DETECTION 2. Nhiệm vụ (yêu cầu về nội dung và số liệu ban đầu): - Tìm hiểu các định dạng ảnh video - Tìm hiểu các tập dữ liệu video bãi gởi xe. - Tìm hiểu các công trình nghiên cứu liên quan và ưu nhược điểm của chúng.
- Nghiên cứu các đặc trưng chỗ đậu xe. - Đề xuất phương pháp thực hiện phân tích, detection. - Hiện thực phương pháp đề xuất và thử trên tập dữ liệu chuẩn. - So sánh với các phương pháp khác.
- Đánh giá giải thuật. Ngày giao nhiệm vụ luận án: 10/08/2022 4. Ngày hoàn thành nhiệm vụ: 15/12/2022 5. Họ tên giảng viên hướng dẫn: PGS.TS Nguyễn Thanh Bình Nội dung và yêu cầu LVTN đã được thông qua Bộ môn.
Ngày …… tháng…… năm 2022 CHỦ NHIỆM BỘ MÔN GIẢNG VIÊN HƯỚNG DẪN CHÍNH (Ký và ghi rõ họ tên) (Ký và ghi rõ họ tên) PGS.TS Trần Minh Quang Nguyễn Thanh Bình PHẦN DÀNH CHO KHOA, BỘ MÔN: Người duyệt (chấm sơ bộ): ________________________ Đơn vị: _______________________________________ Ngày bảo vệ: ___________________________________ Điểm tổng kết: _________________________________ Nơi lưu trữ luận án: _____________________________ Ho Chi Minh University of Technology Faculty of Computer Science and Engineering Declaration of Authenticity I declare that this research is my own work, conducted under the supervision and guidance of Assoc. Nguyen Thanh Binh. The result of my research is legitimate and has not been published in any forms prior to this. All materials used within this research are collected myself by various sources and are appropriately listed in the references section.
In addition, within this research, I also used the results of several other authors and organizations. They have all been aptly referenced. In any case of plagiarism, I stand by my actions and will be responsible for it. Ho Chi Minh city University of Technology therefore are not responsible for any copyright infringements conducted within my research Graduation Thesis (Computer Science), Semester 1, Academic year 2022-2023 Page 1/42 Ho Chi Minh University of Technology Faculty of Computer Science and Engineering Acknowledgement I am using this opportunity to express my gratitude to everyone who supported me during my study and my life.
I am thankful for their guidance, constructive criticism and friendly advice. I offer my sincerest and deepest gratitude to my supervisor, Associate Professor Nguyen Thanh Binh, for his support and guidance. I would like to thank Ho Chi Minh University of Technology for giving me the opportunity to learn great lessons of theory and practical experience. Finally, I recognize that this research would not have been possible without the support from my family and from bottom of my heart, I must acknowledge my parents without whose love, encouragement and sacrifice, I would not have finished this thesis.
Graduation Thesis (Computer Science), Semester 1, Academic year 2022-2023 Page 2/42 Ho Chi Minh University of Technology Faculty of Computer Science and Engineering Contents 1 Introduction 7 1.2 Convolution Neural Network (CNN) .1 CNN-based detection algorithms .2 Two-stage detection algorithms.3 One-stage detection algorithms.2 Model re-parameterization .5 Trainable bag-of-freebies .27 Graduation Thesis (Computer Science), Semester 1, Academic year 2022-2023 Page 3/42 Ho Chi Minh University of Technology Faculty of Computer Science and Engineering 3.1 Intersection over Union (IoU) .2 True Positive, False Positive, False Negative, True Negative .5 Precision - Recall curve .7 Mean Average Precision .1 Hardware and Dataset .2 Dataset selection and annotation .2 Research and Evaluation .3 Testing on images .4 Testing on sample video footage .2 Advantages and Disadvantages .40 Graduation Thesis (Computer Science), Semester 1, Academic year 2022-2023 Page 4/42 Ho Chi Minh University of Technology Faculty of Computer Science and Engineering List of Figures 2.1 A road map of object detection with various milestones [2].2 An example of a CNN structure [7].3 An illustration of the convolution operation [7] .5 Max Pooling operation with 2x2 filters and stride 2 [7] .6 Non-linear operations plots visualization [10].9 Fast R-CNN visualization [13] .10 Faster R-CNN visualization [13] .11 Feature Pyramid Network visualization [13].1 Comparison of YOLOv7 with other real-time object detectors [27].3 Extended efficient layer aggregation networks [27].4 Model scaling for concatenation-based models [27].5 RepConv being used in VGG [27].6 Planned re-parameterized model [27].7 Coarse for auxiliary and fine for lead head label assigner [27].8 Intersection over Union [28] .1 Hardware specifications of Google Colaboratory .32 Graduation Thesis (Computer Science), Semester 1, Academic year 2022-2023 Page 5/42 Ho Chi Minh University of Technology Faculty of Computer Science and Engineering 4.2 Various sample images from the PKLot dataset, annotated.7 Precision Recall (PR) curve.8 Mean average precision with 0.5 IoU value (left) and between 0.12 Video footage test 1 .13 Video footage test 2 .39 Graduation Thesis (Computer Science), Semester 1, Academic year 2022-2023 Page 6/42 Ho Chi Minh University of Technology Faculty of Computer Science and Engineering 1 Introduction 1.1 Problem Statement Getting an available parking spot is a problem faced my many car owners, especially in developing modern cities. Sometimes, and most of the time, there are vacant spots, but the drivers do not have any information about them. It could be, either a free spot far away from them, or is it hidden by some other cars or any other objects big enough to hide the spot. In some cases, parking spaces are managed by people such as security guards who might not have the total view of the next available parking space.
Sometimes the driver themselves has to check for a vacant space by circling around the parking lot, and there is the problem of another driver would come and occupied said slot, thus many loss are generated: time, fuel, and maybe temper. In developing modern cities, urban planning does not follow the quick growth of popula- tion dynamics. It implies that newly brought vehicles between two urban planning imple- mentations, which might not be accommodated in all the existing parking facilities. It leads to bad management of the space by drivers and congestion in the parking lot, especially at peak hours, as drivers are stuck not knowing where to go next.
A study by INRIX found that the average American driver spends 17 hours a year looking for a parking spot. That search costs each driver around $345 in wasted time, gas, and emissions. In larger cities, drivers spend even more time looking for spots.2 Goals Due to such nature of finding a vacant parking slot, a system to detect vacant spaces is desirable to route drivers efficiently to proper empty spots. In order to develop such system, this project’s objectives are the following: • To review theoretical knowledge and related works regarding systems to detect parking space occupancy, applying them to our understanding of the problem and methods to resolve said problem.
• To design a model which can efficiently detect the occupancy status of the parking space Graduation Thesis (Computer Science), Semester 1, Academic year 2022-2023 Page 7/42 Ho Chi Minh University of Technology Faculty of Computer Science and Engineering in an image and a video using our proposed method.3 Limitations In the vision-based approach using machine learning models, the model predicts and checks the image and video based on the information from the training data, which is from the camera point of view at the testing time. Unless a few images of the testing environment are included in the training set, the model may not obtain the same detection result compare to normally. Furthermore, the precision values may change according to the parking space conditions such as lighting, weather and parking space arrangements.4 Thesis Structure This paper is organized as follows: • Chapter 1: Introduction - A brief introduction about the objectives of the thesis. • Chapter 2: Theoretical Background - An introduction of theoretical background and related works as foundation knowledge which are applied in the project.
• Chapter 3: Proposed Methods - Proposed methods and evaluation of said methods to solve the project theoretical problem. • Chapter 4: Project Results - The execution of the selected solution, result and evaluation of said solution. • Chapter 5: Summary - A summary of the final results and future plan.1 Object Detection Object detection refers to the capability of computer and software systems to detect and identify instances of objects of a certain class within an image, or in this case multiple images. In short, object detection answers the most fundamental information needed by computer vision applications: ”What objects are where?”.
Graduation Thesis (Computer Science), Semester 1, Academic year 2022-2023 Page 8/42 Ho Chi Minh University of Technology Faculty of Computer Science and Engineering Different strategies have been proposed to solve the problem of object detection through- out the years. In the past two decades, it is widely accepted that the progress of object detection has generally gone through two historical periods: ”traditional object detection period (before 2014)” and ”deep learning based detection period (after 2014)”, as shown in Figure 2.1: A road map of object detection with various milestones [2]. In 2012, Convolution Neural Network (CNN) was created [3]. Due to its ability to learn robust and high-level feature representations of an image, modern object detection algo- rithms utilized them and started to evolve at an incredible rate, with optimization focused algorithms such as VGGNet [4], GoogLeNet [5] and Deep Residual Learning (ResNet) [6] have been invented over the years.2 Convolution Neural Network (CNN) A Convolutional Neural Networks (CNN) is a subclass of artificial neural networks that specialize in processing data that has a grid-like topology, such as an image.
A digital image is a binary representation of a visual data which contains a series of pixels arranged in a grid that contains its own pixel values to denote how bright and what color each pixel should be. The human brain processes a huge amount of information the second we see an image. Graduation Thesis (Computer Science), Semester 1, Academic year 2022-2023 Page 9/42 Ho Chi Minh University of Technology Faculty of Computer Science and Engineering Each neuron works in its own receptive field and is connected to other neurons in a way that they cover the entire visual field. Just as each neuron responds to stimuli only in the restricted region of the visual field called the receptive field in the biological vision systems, each neuron in a CNN processes data only in its receptive field as well.
The layers are arranged in such a way so that they detect simpler patterns first (lines, curves, etc.) and more complex patterns (faces, objects, etc.3 CNN Layers A typical CNN structure consists of several building blocks, or layers: an input layer, a convolutional layer, an active layer, a pooling layer, a fully connected layer and finally, an output layer. Some types of CNN models might include other layers for different purposes.