VIETNAM NATIONAL UNIVERSITY HO CHI MINH CITY HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY FACULTY OF COMPUTER SCIENCE AND ENGINEERING GRADUATION THESIS DESIGN AN EFFICIENT FPGA-BASED ACCELERATOR FOR REAL-TIME PARKING OCCUPANCY DETECTION Major: COMPUTER ENGINEERING THESIS COMMITTEE : COMPUTER ENGINEERING SUPERVISOR(s) : ASSOC. TRAN NGOC THINH MR. HUYNH PHUC NGHI MEMBER SECRETARY: ASSOC. PHAM QUOC CUONG ---o0o--- STUDENT 1 : NGUYEN VU THANH NGUYEN - 1652437 HO CHI MINH CITY, 01/2023 ĐẠI HỌC QUỐC GIA TP.HCM CỘNG HÒA XÃ HỘI CHỦ NGHĨA VIỆT NAM -----Độc lập - Tự do - Hạnh phúc----- TRƯỜNG ĐẠI HỌC BÁCH KHOA KHOA: KH & KT Máy tính NHIỆM VỤ LUẬN ÁN TỐT NGHIỆP BỘ MÔN: KT Máy tính Chú ý: Sinh viên phải dán tờ này vào trang nhất của bản thuyết trình HỌ VÀ TÊN: Nguyễn Vũ Thành Nguyễn MSSV: 1652437 NGÀNH: Kỹ thuật Máy tính LỚP: 1.
Đầu đề luận án: Design an efficient FPGA-based accelerator for real-time parking occupancy detection 2. Nhiệm vụ (yêu cầu về nội dung và số liệu ban đầu): • Research on a BNN approach in image classification task with CNRPark dataset. • Implementation of an encoder for input images and a set of weight parameters for parking solutions. • Build and evaluate a hardware accelerator run on Ultra96v2 SoC with a pre-encoded 32x32 image set.
Ngày giao nhiệm vụ luận án: 19/09/2022 4. Ngày hoàn thành nhiệm vụ: 09/01/2023 5. Họ tên giảng viên hướng dẫn: Phần hướng dẫn: 1) PGS. TS Trần Ngọc Thịnh 2) KS.
Huỳnh Phúc Nghị Nội dung và yêu cầu LVTN đã được thông qua Bộ môn. Ngày tháng năm 2023 CHỦ NHIỆM BỘ MÔN GIẢNG VIÊN HƯỚNG DẪN CHÍNH (Ký và ghi rõ họ tên) (Ký và ghi rõ họ tên) PHẦN DÀNH CHO KHOA, BỘ MÔN: Người duyệt (chấm sơ bộ): Đơn vị: Ngày bảo vệ: Điểm tổng kết: Nơi lưu trữ luận án: Commitment I hereby declare that I worked on and is the sole author of this bachelor thesis and that I have not used any sources other than those listed in the bibliography and identified as references. Other than that, the work presented is entirely my own. Nguyen Vu Thanh Nguyen 1 Acknowledgements This thesis has completely reached its result thanks to the continuous efforts of myself, the support and encouragement of our lecturers, friends and families.
I would like to express our sincere attitude to those who have helped us throughout the study, research and during working on the thesis. I would first like to express my sincere gratitude to my supervisors, Assoc Professor Tran Ngoc Thinh and BEng Huynh Phuc Nghi have consistently helped me out by providing me with not only the necessary tools to complete the thesis but also with a wealth of information and direction so that I could move in the best route. They never stopped inspiring me and provided me the chance to participate in a really fascinating work that was built using a lot of the information I had acquired during our time at university. They have always been a kind, understanding teacher who encourages me to make changes to the work and this thesis as needed.
Working with them and gaining expertise under their guidance was an honor for me. In addition to thanking the supervisors, I also like to acknowledge the councilors at the thesis defense for their wise criticism and suggestions that helped me improve my work. Once more, I also want to express our gratitude to all of the instructors in the Faculty of Computer Science and Engineering, as well as to all of the instructors at Ho Chi Minh City University of Technology, for their commitment to teaching and helping me learn the fundamentals of engineering. I want to express my gratitude to everyone of my friends for being my mentors and helpers during our time at the institution.
Finally, I would like to thank my parents for always providing a positive environment for our development and for supporting me when I face difficulties in both our academic and personal live. Nguyen Vu Thanh Nguyen 2 Abstracts This thesis proposes, studies and examines an approach on developing an edge-ai smart parking solution, including hardware and software components by implementation of the FracBNN+CNRPark model with hardware acceleration on the Ultra96-V2 board. This approach allows end-users to be able to monitor and detect busy and free parking spaces automatically via security cameras. The image classification model runs entirely on the edge, on the Ultra96-V2 board, without the help of a server workstation.
3 Contents Commitment 2 Acknowledgements 3 Abstracts 4 List of Figures 10 List of Table 11 Terms 12 1 Introduction 1 1.1 Purpose and Motivation .2 Scope and Objectives .3 Structure of Thesis. 4 2 Background knowledge and Terminology 5 2.1 Software - Artificial Intelligence .1 The development of AI .3 Convolutional Neural Network (CNN) .4 Binary Neural Network (BNN) .2 Smart Parking concepts .3 Hardware and constraints .1 FPGA and SoC .4 Tools and Frameworks .2 Vivado and Vivado HLS .3 PYNQ and BNN-PYNQ .1 Recall, Precision & F1-score .2 Average IoU & mean Average Accuracy (mAP) .1 Smart Parking Related works .1 Moscow Parking (Source: https://parking.4 Cisco Smart+Connected City Parking .2 Solutions in Vietnam .3 Smart Parking systems with Image Processing .2 Previous Group’s Thesis Result .1 License Plate Dataset .2 YOLOv3 Object Detection Model .3 Implementing Vien’s approach .4 Comparing YOLOv3 on Ultra96-V2 and JetsonNano 40 3.4 The Original BNN Model .5 Improved BNN Models .1 The proposed solution - Previous thesis group .2 Hardware Accelerator Architecture .1 FracBNN+CNRPark model on Pytorch .1 Training the model .2 Hardware Acceleration on Ultra96-V2 .1 Weights and Bias processing .3 Building model on Vivado HLS and Vivado .4 Inference on the Ultra96-V2 .2 Hardware acceleration result .1 Comparing to Vien’s thesis .2 Trained and implemented FracBNN+CNRPark model on Pytorch, run on GPU .3 Implemented hardware acceleration via Vivado HLS on Ul- tra96v2 .4 Achieve inference result using built thermometer encoder on Ultra96v2 .2 Real-time application. 72 7 List of Figures 2.1 Development timeline of AI, ML and DL [1] .2 a Neural Network with 4 layers .3 Composition of a hidden layer on the CNN .4 Five commonly used activation functions: (a) binary step function, (b) sigmoid function, (c) tanh function, (d) ReLU function, (e) leaky ReLU function.5 Structure of a CNN. The data is run through several convolutional and pooling layers learning features in the image.6 Convolutional filter size 3x3 sweeping through image size 4x4 .7 Output (darkgreen) from a convolutional filter .8 Convolutional filter sweeping with padding size 1, stride of 2 .10 Smart Parking Solution .11 Edge AI Workflow .12 An FPGA block diagram .13 Front view of Ultra96-V2 .14 Block diagram of the Ultra96-V2 .16 Calculation of IoU .1 Vietnamese License Plate Dataset.
36 8 List of Figures 3.2 YOLOv3 performance result .3 YOLOv3 box bounding technique .5 GPU Implementation Workflow .6 Benchmark with threshold IoU 50% .7 Benchmark with threshold IoU 75% .8 A visualization of the sign layer and Straight-Through Estimator (STE). While the real values of the weights are processed by the sign function in the forward pass, the gradient of the binary weights are simply passed through to the real valued weights.9 BNN Training Curve .10 Evolution of BNN Accuracy - Source: FracBNN introduction pre- sentation slides (FPGA2021) .11 Main contributions of FracBNN - Source: FracBNN introduction presentation slides (FPGA2021) .12 Input images need to be first encoded .13 Results of binarizing the input layer using thermometer encoding on CIFAR-10 – ResNet-20 BNN has 0.27 million parameters and 40.14 Improving BNN by computing an additional sparse binary convolu- tion layer.1 Previous group’s proposed solution - Smart Parking System Archi- tecture [4] .2 FracBNN+CNRPark model architecture, based on Resnet20 [5] .3 Basic blocks - green highlights are the difference to ReActNet [6] model .4 Edge solution architecture .5 FracBNN accelerator architecture[3] .6 CNRPark+EXT dataset sample .1 FracBNN Training workflow .2 Generating an FPGA accelerator from trained FracBNN. 61 9 List of Figures 5.3 Thermometer encoder workflow .4 Vivado HLS Utilization Estimate. 65 10 List of Tables 2.1 Energy Consumption when Inferencing CNV Model on Ultra96v2 and VGG Model on Jetson Nano[4] .2 Summary of Vien’s thesis result .3 Summary and Comparison between Ultra96v2 and Jetson Nano .4 A table of major details of the methods presented in this section.5 Comparison of accuracies on the ImageNet dataset from works presented in this section.
Full precision network accuracies are included for comparison as well.1 Resources utilization on Ultra96v2 .2 Accuracy comparison between models of input size 32x32 .3 Inference results of prospective models on Ultra96v2 .4 Power consumption on Ultra96v2 of different models. Application Specific Integrated Circuits BNN. Binary Neural Network BRAM. Block Random Access Memory CNN.
Convolutional Neural Network CPU. Core Processing Unit FF. Flip-flop FPGA. Field Programmable Gate Array GPU.
Graphics Processing Unit MPSoC. Multi-Processors System on Chip LUT. Look Up Table PL. Programmable Logic RGB.
Red Green Blue 1 Introduction 1.1 Purpose and Motivation Machine learning became more well-known at the beginning of the twenty- first century. This resulted from the need to process increasing amounts of data and the availability of less expensive ML-capable hardware, like GPUs and RAM. Due to their inherent design, neural networks heavily rely on parallel computing to be effective, making a GPU with its many cores the ideal tool for the job. Due to this, GPUs have been the norm for machine learning applications for the last few years, but new hardware has recently been introduced.
The Tensor Processing Unit, or TPU, a chip made specifically for machine learning, was unveiled by Google in 2016. As a way to obtain the effectiveness of a dedicated chip without the restrictions of an ASIC, FPGAs have also been demonstrating promise over time. The development of machine learning has also greatly benefited from the current trend of cloud computing. More researchers can conduct their research cost-effectively thanks to the availability of large GPU, TPU, or FPGA centers’ computing power.
[1] Each year, more and more applications for machine learning are released, signaling the field’s rapid expansion. Today, we use applications on a daily basis, whether they are in our cars, phones, medical software, or almost any other hi-tech product. To be accurate and effective, machine learning requires a lot of data, and as machine learning becomes more popular, more data is being sent across various networks. We send enormous amounts of data back and forth, particularly for applications that collect data in the field, send it to a datacenter for processing, and then return the results.
As a result, the system that collects the data needs to further develop its ability to process data at the network’s edge. For instance, this could significantly reduce the amount of data on the networks for a video stream 1 1.1 Purpose and Motivation that counts the number of cars on a highway. The counted number of cars can be sent to a datacenter whenever necessary rather than sending full scale images 30 times per second. Machine learning will permeate more and more aspects of our lives, as it stands today.
There are countless uses for intelligent machines that can assist us in performing tasks without being micromanaged. An application should always perform processing as close to the edge as possible to avoid clogging our communication networks with data.[1] As Vietnam today continues to develop economically, our country has seen a steady increase in the number of cars in traffic, facilitating the needs for more, and bigger parking lots to accommodate it. This has led to many new problems for both drivers and management, one of which is the increasing difficulty in identifying and finding empty parking space. Smart Parking emerged as a viable solution to these issues, but this system still can be improved.
To further optimize and enhance its performance, we propose the implementation of hardware acceleration.