Thiết Kế Bộ Tăng Tốc FPGA Hiệu Quả Cho Phát Hiện Tình Trạng Đỗ Xe Thực Thời

Tài liệu nghiên cứu Design an efficient fpga based accelerator for real time parking occupancy detection, tổng hợp lý thuyết và thực hành, cung cấp kiến thức chuyên sâu về .

Trường đại học

Đại Học Quốc Gia TP.HCM

Chuyên ngành

Kỹ thuật Máy tính

Người đăng

Ẩn danh

Thể loại

luận án tốt nghiệp

2023

91
5
0

Phí lưu trữ

35 Point

Mục lục chi tiết

Commitment

Acknowledgements

Abstracts

Contents

1. Introduction

1.1. Purpose and Motivation

1.2. Scope and Objectives

1.3. Structure of Thesis

2. Background knowledge and Terminology

2.1. Software - Artificial Intelligence

2.1.1. The development of AI

Tóm tắt

I. Tổng Quan Về Thiết Kế Bộ Tăng Tốc FPGA Cho Phát Hiện Tình Trạng Đỗ Xe

Thiết kế bộ tăng tốc FPGA cho phát hiện tình trạng đỗ xe thực thời là một lĩnh vực đang phát triển mạnh mẽ. Công nghệ này không chỉ giúp cải thiện hiệu suất phát hiện mà còn giảm thiểu chi phí và thời gian xử lý. Việc sử dụng FPGA cho phép tối ưu hóa các thuật toán học máy, mang lại kết quả chính xác và nhanh chóng hơn. Hệ thống này có thể được áp dụng trong nhiều lĩnh vực khác nhau, từ quản lý bãi đỗ xe đến các ứng dụng thông minh trong đô thị.

1.1. Ứng Dụng Của FPGA Trong Phát Hiện Tình Trạng Đỗ Xe

FPGA được sử dụng để xử lý hình ảnh từ camera giám sát, giúp phát hiện tình trạng đỗ xe một cách nhanh chóng và hiệu quả. Hệ thống này có thể nhận diện các vị trí đỗ xe trống và đầy, từ đó cung cấp thông tin cho người dùng.

1.2. Lợi Ích Của Việc Sử Dụng Bộ Tăng Tốc FPGA

Sử dụng bộ tăng tốc FPGA giúp giảm thiểu độ trễ trong việc xử lý dữ liệu, đồng thời tiết kiệm năng lượng. Điều này rất quan trọng trong các ứng dụng yêu cầu xử lý thời gian thực như phát hiện tình trạng đỗ xe.

II. Vấn Đề Và Thách Thức Trong Phát Hiện Tình Trạng Đỗ Xe

Mặc dù công nghệ phát hiện tình trạng đỗ xe đã có những bước tiến đáng kể, nhưng vẫn còn nhiều thách thức cần phải vượt qua. Độ chính xác của các thuật toán hiện tại vẫn chưa đạt yêu cầu, đặc biệt trong các điều kiện ánh sáng khác nhau. Ngoài ra, việc xử lý dữ liệu lớn từ nhiều camera cũng là một vấn đề cần được giải quyết.

2.1. Độ Chính Xác Của Hệ Thống Phát Hiện

Độ chính xác của hệ thống phát hiện tình trạng đỗ xe phụ thuộc vào chất lượng hình ảnh và thuật toán được sử dụng. Các mô hình học sâu cần được tối ưu hóa để cải thiện khả năng nhận diện.

2.2. Khả Năng Xử Lý Dữ Liệu Lớn

Hệ thống cần có khả năng xử lý dữ liệu từ nhiều camera cùng lúc mà không làm giảm hiệu suất. Việc này đòi hỏi một kiến trúc phần cứng mạnh mẽ và hiệu quả.

III. Phương Pháp Thiết Kế Bộ Tăng Tốc FPGA Hiệu Quả

Để thiết kế một bộ tăng tốc FPGA hiệu quả cho phát hiện tình trạng đỗ xe, cần áp dụng các phương pháp tối ưu hóa thuật toán và phần cứng. Việc sử dụng các mô hình học máy như BNN có thể giúp giảm thiểu tài nguyên cần thiết mà vẫn đảm bảo hiệu suất.

3.1. Tối Ưu Hóa Thuật Toán Học Máy

Các thuật toán học máy cần được tối ưu hóa để phù hợp với kiến trúc FPGA. Việc này bao gồm việc giảm số lượng tham số và cải thiện tốc độ xử lý.

3.2. Thiết Kế Kiến Trúc Phần Cứng

Kiến trúc phần cứng cần được thiết kế sao cho tối ưu hóa việc sử dụng tài nguyên FPGA, từ đó nâng cao hiệu suất xử lý và giảm thiểu độ trễ.

IV. Ứng Dụng Thực Tiễn Của Hệ Thống Phát Hiện Tình Trạng Đỗ Xe

Hệ thống phát hiện tình trạng đỗ xe có thể được áp dụng trong nhiều lĩnh vực khác nhau, từ quản lý bãi đỗ xe công cộng đến các ứng dụng trong các khu đô thị thông minh. Việc này không chỉ giúp tiết kiệm thời gian cho người lái xe mà còn tối ưu hóa việc sử dụng không gian đỗ xe.

4.1. Quản Lý Bãi Đỗ Xe Công Cộng

Hệ thống có thể giúp quản lý bãi đỗ xe công cộng hiệu quả hơn, từ đó giảm thiểu tình trạng ùn tắc giao thông và tiết kiệm thời gian cho người lái xe.

4.2. Ứng Dụng Trong Đô Thị Thông Minh

Hệ thống phát hiện tình trạng đỗ xe có thể tích hợp vào các giải pháp đô thị thông minh, giúp cải thiện chất lượng cuộc sống cho cư dân.

V. Kết Luận Và Tương Lai Của Công Nghệ Phát Hiện Tình Trạng Đỗ Xe

Công nghệ phát hiện tình trạng đỗ xe đang trên đà phát triển mạnh mẽ với nhiều tiềm năng ứng dụng trong tương lai. Việc cải thiện độ chính xác và hiệu suất của hệ thống sẽ là mục tiêu hàng đầu trong nghiên cứu tiếp theo.

5.1. Tiềm Năng Phát Triển Trong Tương Lai

Công nghệ này có thể mở ra nhiều cơ hội mới trong việc quản lý giao thông và phát triển đô thị thông minh.

5.2. Hướng Nghiên Cứu Tiếp Theo

Nghiên cứu tiếp theo sẽ tập trung vào việc cải thiện các thuật toán học máy và tối ưu hóa phần cứng để nâng cao hiệu suất của hệ thống.

08/07/2025

Trích đoạn nội dung tài liệu

VIETNAM NATIONAL UNIVERSITY HO CHI MINH CITY HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY FACULTY OF COMPUTER SCIENCE AND ENGINEERING GRADUATION THESIS DESIGN AN EFFICIENT FPGA-BASED ACCELERATOR FOR REAL-TIME PARKING OCCUPANCY DETECTION Major: COMPUTER ENGINEERING THESIS COMMITTEE : COMPUTER ENGINEERING SUPERVISOR(s) : ASSOC. TRAN NGOC THINH MR. HUYNH PHUC NGHI MEMBER SECRETARY: ASSOC. PHAM QUOC CUONG ---o0o--- STUDENT 1 : NGUYEN VU THANH NGUYEN - 1652437 HO CHI MINH CITY, 01/2023 ĐẠI HỌC QUỐC GIA TP.HCM CỘNG HÒA XÃ HỘI CHỦ NGHĨA VIỆT NAM -----Độc lập - Tự do - Hạnh phúc----- TRƯỜNG ĐẠI HỌC BÁCH KHOA KHOA: KH & KT Máy tính NHIỆM VỤ LUẬN ÁN TỐT NGHIỆP BỘ MÔN: KT Máy tính Chú ý: Sinh viên phải dán tờ này vào trang nhất của bản thuyết trình HỌ VÀ TÊN: Nguyễn Vũ Thành Nguyễn MSSV: 1652437 NGÀNH: Kỹ thuật Máy tính LỚP: 1.

Đầu đề luận án: Design an efficient FPGA-based accelerator for real-time parking occupancy detection 2. Nhiệm vụ (yêu cầu về nội dung và số liệu ban đầu): • Research on a BNN approach in image classification task with CNRPark dataset. • Implementation of an encoder for input images and a set of weight parameters for parking solutions. • Build and evaluate a hardware accelerator run on Ultra96v2 SoC with a pre-encoded 32x32 image set.

Ngày giao nhiệm vụ luận án: 19/09/2022 4. Ngày hoàn thành nhiệm vụ: 09/01/2023 5. Họ tên giảng viên hướng dẫn: Phần hướng dẫn: 1) PGS. TS Trần Ngọc Thịnh 2) KS.

Huỳnh Phúc Nghị Nội dung và yêu cầu LVTN đã được thông qua Bộ môn. Ngày tháng năm 2023 CHỦ NHIỆM BỘ MÔN GIẢNG VIÊN HƯỚNG DẪN CHÍNH (Ký và ghi rõ họ tên) (Ký và ghi rõ họ tên) PHẦN DÀNH CHO KHOA, BỘ MÔN: Người duyệt (chấm sơ bộ): Đơn vị: Ngày bảo vệ: Điểm tổng kết: Nơi lưu trữ luận án: Commitment I hereby declare that I worked on and is the sole author of this bachelor thesis and that I have not used any sources other than those listed in the bibliography and identified as references. Other than that, the work presented is entirely my own. Nguyen Vu Thanh Nguyen 1 Acknowledgements This thesis has completely reached its result thanks to the continuous efforts of myself, the support and encouragement of our lecturers, friends and families.

I would like to express our sincere attitude to those who have helped us throughout the study, research and during working on the thesis. I would first like to express my sincere gratitude to my supervisors, Assoc Professor Tran Ngoc Thinh and BEng Huynh Phuc Nghi have consistently helped me out by providing me with not only the necessary tools to complete the thesis but also with a wealth of information and direction so that I could move in the best route. They never stopped inspiring me and provided me the chance to participate in a really fascinating work that was built using a lot of the information I had acquired during our time at university. They have always been a kind, understanding teacher who encourages me to make changes to the work and this thesis as needed.

Working with them and gaining expertise under their guidance was an honor for me. In addition to thanking the supervisors, I also like to acknowledge the councilors at the thesis defense for their wise criticism and suggestions that helped me improve my work. Once more, I also want to express our gratitude to all of the instructors in the Faculty of Computer Science and Engineering, as well as to all of the instructors at Ho Chi Minh City University of Technology, for their commitment to teaching and helping me learn the fundamentals of engineering. I want to express my gratitude to everyone of my friends for being my mentors and helpers during our time at the institution.

Finally, I would like to thank my parents for always providing a positive environment for our development and for supporting me when I face difficulties in both our academic and personal live. Nguyen Vu Thanh Nguyen 2 Abstracts This thesis proposes, studies and examines an approach on developing an edge-ai smart parking solution, including hardware and software components by implementation of the FracBNN+CNRPark model with hardware acceleration on the Ultra96-V2 board. This approach allows end-users to be able to monitor and detect busy and free parking spaces automatically via security cameras. The image classification model runs entirely on the edge, on the Ultra96-V2 board, without the help of a server workstation.

3 Contents Commitment 2 Acknowledgements 3 Abstracts 4 List of Figures 10 List of Table 11 Terms 12 1 Introduction 1 1.1 Purpose and Motivation .2 Scope and Objectives .3 Structure of Thesis. 4 2 Background knowledge and Terminology 5 2.1 Software - Artificial Intelligence .1 The development of AI .3 Convolutional Neural Network (CNN) .4 Binary Neural Network (BNN) .2 Smart Parking concepts .3 Hardware and constraints .1 FPGA and SoC .4 Tools and Frameworks .2 Vivado and Vivado HLS .3 PYNQ and BNN-PYNQ .1 Recall, Precision & F1-score .2 Average IoU & mean Average Accuracy (mAP) .1 Smart Parking Related works .1 Moscow Parking (Source: https://parking.4 Cisco Smart+Connected City Parking .2 Solutions in Vietnam .3 Smart Parking systems with Image Processing .2 Previous Group’s Thesis Result .1 License Plate Dataset .2 YOLOv3 Object Detection Model .3 Implementing Vien’s approach .4 Comparing YOLOv3 on Ultra96-V2 and JetsonNano 40 3.4 The Original BNN Model .5 Improved BNN Models .1 The proposed solution - Previous thesis group .2 Hardware Accelerator Architecture .1 FracBNN+CNRPark model on Pytorch .1 Training the model .2 Hardware Acceleration on Ultra96-V2 .1 Weights and Bias processing .3 Building model on Vivado HLS and Vivado .4 Inference on the Ultra96-V2 .2 Hardware acceleration result .1 Comparing to Vien’s thesis .2 Trained and implemented FracBNN+CNRPark model on Pytorch, run on GPU .3 Implemented hardware acceleration via Vivado HLS on Ul- tra96v2 .4 Achieve inference result using built thermometer encoder on Ultra96v2 .2 Real-time application. 72 7 List of Figures 2.1 Development timeline of AI, ML and DL [1] .2 a Neural Network with 4 layers .3 Composition of a hidden layer on the CNN .4 Five commonly used activation functions: (a) binary step function, (b) sigmoid function, (c) tanh function, (d) ReLU function, (e) leaky ReLU function.5 Structure of a CNN. The data is run through several convolutional and pooling layers learning features in the image.6 Convolutional filter size 3x3 sweeping through image size 4x4 .7 Output (darkgreen) from a convolutional filter .8 Convolutional filter sweeping with padding size 1, stride of 2 .10 Smart Parking Solution .11 Edge AI Workflow .12 An FPGA block diagram .13 Front view of Ultra96-V2 .14 Block diagram of the Ultra96-V2 .16 Calculation of IoU .1 Vietnamese License Plate Dataset.

36 8 List of Figures 3.2 YOLOv3 performance result .3 YOLOv3 box bounding technique .5 GPU Implementation Workflow .6 Benchmark with threshold IoU 50% .7 Benchmark with threshold IoU 75% .8 A visualization of the sign layer and Straight-Through Estimator (STE). While the real values of the weights are processed by the sign function in the forward pass, the gradient of the binary weights are simply passed through to the real valued weights.9 BNN Training Curve .10 Evolution of BNN Accuracy - Source: FracBNN introduction pre- sentation slides (FPGA2021) .11 Main contributions of FracBNN - Source: FracBNN introduction presentation slides (FPGA2021) .12 Input images need to be first encoded .13 Results of binarizing the input layer using thermometer encoding on CIFAR-10 – ResNet-20 BNN has 0.27 million parameters and 40.14 Improving BNN by computing an additional sparse binary convolu- tion layer.1 Previous group’s proposed solution - Smart Parking System Archi- tecture [4] .2 FracBNN+CNRPark model architecture, based on Resnet20 [5] .3 Basic blocks - green highlights are the difference to ReActNet [6] model .4 Edge solution architecture .5 FracBNN accelerator architecture[3] .6 CNRPark+EXT dataset sample .1 FracBNN Training workflow .2 Generating an FPGA accelerator from trained FracBNN. 61 9 List of Figures 5.3 Thermometer encoder workflow .4 Vivado HLS Utilization Estimate. 65 10 List of Tables 2.1 Energy Consumption when Inferencing CNV Model on Ultra96v2 and VGG Model on Jetson Nano[4] .2 Summary of Vien’s thesis result .3 Summary and Comparison between Ultra96v2 and Jetson Nano .4 A table of major details of the methods presented in this section.5 Comparison of accuracies on the ImageNet dataset from works presented in this section.

Full precision network accuracies are included for comparison as well.1 Resources utilization on Ultra96v2 .2 Accuracy comparison between models of input size 32x32 .3 Inference results of prospective models on Ultra96v2 .4 Power consumption on Ultra96v2 of different models. Application Specific Integrated Circuits BNN. Binary Neural Network BRAM. Block Random Access Memory CNN.

Convolutional Neural Network CPU. Core Processing Unit FF. Flip-flop FPGA. Field Programmable Gate Array GPU.

Graphics Processing Unit MPSoC. Multi-Processors System on Chip LUT. Look Up Table PL. Programmable Logic RGB.

Red Green Blue 1 Introduction 1.1 Purpose and Motivation Machine learning became more well-known at the beginning of the twenty- first century. This resulted from the need to process increasing amounts of data and the availability of less expensive ML-capable hardware, like GPUs and RAM. Due to their inherent design, neural networks heavily rely on parallel computing to be effective, making a GPU with its many cores the ideal tool for the job. Due to this, GPUs have been the norm for machine learning applications for the last few years, but new hardware has recently been introduced.

The Tensor Processing Unit, or TPU, a chip made specifically for machine learning, was unveiled by Google in 2016. As a way to obtain the effectiveness of a dedicated chip without the restrictions of an ASIC, FPGAs have also been demonstrating promise over time. The development of machine learning has also greatly benefited from the current trend of cloud computing. More researchers can conduct their research cost-effectively thanks to the availability of large GPU, TPU, or FPGA centers’ computing power.

[1] Each year, more and more applications for machine learning are released, signaling the field’s rapid expansion. Today, we use applications on a daily basis, whether they are in our cars, phones, medical software, or almost any other hi-tech product. To be accurate and effective, machine learning requires a lot of data, and as machine learning becomes more popular, more data is being sent across various networks. We send enormous amounts of data back and forth, particularly for applications that collect data in the field, send it to a datacenter for processing, and then return the results.

As a result, the system that collects the data needs to further develop its ability to process data at the network’s edge. For instance, this could significantly reduce the amount of data on the networks for a video stream 1 1.1 Purpose and Motivation that counts the number of cars on a highway. The counted number of cars can be sent to a datacenter whenever necessary rather than sending full scale images 30 times per second. Machine learning will permeate more and more aspects of our lives, as it stands today.

There are countless uses for intelligent machines that can assist us in performing tasks without being micromanaged. An application should always perform processing as close to the edge as possible to avoid clogging our communication networks with data.[1] As Vietnam today continues to develop economically, our country has seen a steady increase in the number of cars in traffic, facilitating the needs for more, and bigger parking lots to accommodate it. This has led to many new problems for both drivers and management, one of which is the increasing difficulty in identifying and finding empty parking space. Smart Parking emerged as a viable solution to these issues, but this system still can be improved.

To further optimize and enhance its performance, we propose the implementation of hardware acceleration.

Nội dung được bảo vệ bản quyền — Tải xuống đầy đủ

Tài liệu có tiêu đề Thiết Kế Bộ Tăng Tốc FPGA Hiệu Quả Cho Phát Hiện Tình Trạng Đỗ Xe Thực Thời trình bày một giải pháp công nghệ tiên tiến nhằm cải thiện khả năng phát hiện tình trạng đỗ xe trong thời gian thực thông qua việc sử dụng FPGA. Bài viết nhấn mạnh các phương pháp thiết kế tối ưu, giúp tăng tốc độ xử lý và độ chính xác trong việc nhận diện các vị trí đỗ xe. Điều này không chỉ mang lại lợi ích cho các hệ thống quản lý bãi đỗ xe mà còn mở ra cơ hội cho việc phát triển các ứng dụng thông minh hơn trong lĩnh vực giao thông.

Để hiểu rõ hơn về các ứng dụng liên quan, bạn có thể tham khảo tài liệu Luận văn thạc sĩ công nghệ thông tin xây dựng ứng dụng dự đoán biển số xe bằng thuật toán nhận dạng đối tượng, nơi trình bày cách sử dụng AI trong việc nhận diện biển số xe. Ngoài ra, tài liệu Nghiên cứu các thuật toán xử lý ảnh ứng dụng trong nhận dạng biển kiểm soát phương tiện giao thông sẽ cung cấp cái nhìn sâu sắc về các thuật toán xử lý ảnh, hỗ trợ cho việc phát triển các hệ thống nhận diện biển số hiệu quả hơn. Những tài liệu này sẽ giúp bạn mở rộng kiến thức và khám phá thêm nhiều khía cạnh thú vị trong lĩnh vực công nghệ giao thông.