Đánh Giá Hiệu Quả Của Kiến Trúc CNN-Transformers Trong Việc Phân Loại Tình Trạng Hư Hỏng Đường Bộ

Khóa luận đánh giá hiệu quả kiến trúc CNN Transformers trong phân loại tình trạng hư hỏng đường bộ, góp phần nâng cao chất lượng giao thông.

Trường đại học

University of Information Technology - VNU-HCM

Chuyên ngành

Computer Science

Người đăng

Ẩn danh

Thể loại

dissertation

2024

75
3
0

Phí lưu trữ

30 Point

Mục lục chi tiết

ACKNOWLEDGEMENTS

1. CHAPTER 1: Objectives of the Thesis

1.1. Subjects and Scope of Research

1.1.1. Research Subjects

1.1.2. Scope of Research

2. CHAPTER 2: Overview of Artificial Intelligence, Machine Learning, Deep Learning

2.1. Typical Architectures in Deep Learning

2.2. Convolution Neural Network

2.3. Image classification problem

2.4. Some studies on the problem of detecting damaged paths

3. CHAPTER 3: Overview of the Model

3.1. Zero shot learning

4. CHAPTER 4: EXPERIMENTS AND EVALUATION

4.1. Experimental results

5. CHAPTER 5: SUMMARY AND DEVELOPMENT DIRECTIONS

5.1. The achieved results

5.2. Development Directions

REFERENCES

Tóm tắt

I. Tổng Quan Về Kiến Trúc CNN Transformers Trong Phân Loại Hư Hỏng Đường Bộ

Kiến trúc CNN-Transformers đang trở thành một trong những giải pháp tiên tiến trong việc phân loại tình trạng hư hỏng đường bộ. Sự kết hợp giữa mạng nơ-ron tích chập (CNN) và transformers mang lại khả năng phân tích hình ảnh mạnh mẽ, giúp cải thiện độ chính xác trong việc phát hiện các vấn đề trên bề mặt đường. Nghiên cứu này sẽ khám phá cách mà kiến trúc này hoạt động và lý do tại sao nó lại hiệu quả trong lĩnh vực này.

1.1. Định Nghĩa Kiến Trúc CNN và Transformers

Kiến trúc CNN là một loại mạng nơ-ron được thiết kế đặc biệt để xử lý dữ liệu hình ảnh. Trong khi đó, transformers sử dụng cơ chế attention để xử lý thông tin theo cách toàn cục, giúp phát hiện các đặc điểm ẩn mà CNN có thể bỏ qua.

1.2. Lợi Ích Của Việc Kết Hợp CNN và Transformers

Sự kết hợp này không chỉ cải thiện độ chính xác mà còn tăng tốc độ xử lý hình ảnh. Việc sử dụng transformers giúp mô hình hiểu rõ hơn về ngữ cảnh của hình ảnh, từ đó nâng cao khả năng phân loại tình trạng hư hỏng.

II. Thách Thức Trong Phân Loại Tình Trạng Hư Hỏng Đường Bộ

Phân loại tình trạng hư hỏng đường bộ gặp nhiều thách thức, từ việc thu thập dữ liệu đến việc phát triển mô hình chính xác. Các yếu tố như điều kiện ánh sáng, góc chụp và chất lượng hình ảnh đều ảnh hưởng đến kết quả phân loại. Nghiên cứu này sẽ chỉ ra những khó khăn chính và cách mà kiến trúc CNN-Transformers có thể giải quyết.

2.1. Khó Khăn Trong Việc Thu Thập Dữ Liệu

Việc thu thập dữ liệu hình ảnh chất lượng cao là một thách thức lớn. Dữ liệu không đồng nhất có thể dẫn đến việc mô hình học không chính xác, ảnh hưởng đến khả năng phân loại.

2.2. Ảnh Hưởng Của Điều Kiện Môi Trường

Điều kiện môi trường như ánh sáng và thời tiết có thể làm giảm chất lượng hình ảnh, từ đó ảnh hưởng đến độ chính xác của mô hình. Việc phát triển các giải pháp để xử lý những yếu tố này là rất cần thiết.

III. Phương Pháp Nghiên Cứu Kiến Trúc CNN Transformers

Nghiên cứu này áp dụng phương pháp kết hợp giữa CNN và transformers để phân loại tình trạng hư hỏng đường bộ. Các bước thực hiện bao gồm thu thập dữ liệu, tiền xử lý, xây dựng mô hình và đánh giá hiệu quả. Mô hình được thiết kế để tối ưu hóa khả năng nhận diện và phân loại các tình trạng hư hỏng khác nhau.

3.1. Quy Trình Thu Thập và Tiền Xử Lý Dữ Liệu

Quy trình này bao gồm việc thu thập hình ảnh từ nhiều nguồn khác nhau và thực hiện các bước tiền xử lý như cắt, làm sạch và chuẩn hóa dữ liệu để đảm bảo chất lượng đầu vào cho mô hình.

3.2. Xây Dựng Mô Hình CNN Transformers

Mô hình được xây dựng bằng cách kết hợp các lớp CNN để trích xuất đặc trưng và các lớp transformers để xử lý thông tin. Điều này giúp mô hình có khả năng nhận diện các đặc điểm phức tạp trong hình ảnh.

IV. Kết Quả Nghiên Cứu và Ứng Dụng Thực Tiễn

Kết quả nghiên cứu cho thấy mô hình CNN-Transformers đạt được độ chính xác cao trong việc phân loại tình trạng hư hỏng đường bộ. Các ứng dụng thực tiễn của mô hình này có thể giúp các cơ quan quản lý giao thông cải thiện hiệu quả trong việc duy trì và bảo trì hạ tầng giao thông.

4.1. Đánh Giá Hiệu Quả Mô Hình

Mô hình đã được đánh giá qua nhiều bộ dữ liệu khác nhau và cho thấy khả năng phân loại chính xác lên đến 95%. Điều này chứng tỏ tính khả thi của việc áp dụng mô hình trong thực tế.

4.2. Ứng Dụng Trong Quản Lý Giao Thông

Mô hình có thể được tích hợp vào các hệ thống giám sát giao thông để tự động phát hiện và phân loại tình trạng hư hỏng, từ đó giúp các cơ quan chức năng có biện pháp xử lý kịp thời.

V. Kết Luận và Hướng Phát Triển Tương Lai

Nghiên cứu này đã chỉ ra rằng kiến trúc CNN-Transformers có tiềm năng lớn trong việc phân loại tình trạng hư hỏng đường bộ. Hướng phát triển tương lai có thể bao gồm việc tối ưu hóa mô hình cho các thiết bị di động và mở rộng khả năng áp dụng cho các loại hình giao thông khác.

5.1. Tóm Tắt Kết Quả Nghiên Cứu

Kết quả nghiên cứu đã chứng minh rằng việc kết hợp CNN và transformers mang lại hiệu quả cao trong phân loại tình trạng hư hỏng đường bộ, mở ra hướng đi mới cho các nghiên cứu tiếp theo.

5.2. Hướng Phát Triển Mới

Nghiên cứu có thể được mở rộng để áp dụng cho các loại hình giao thông khác nhau, cũng như phát triển các mô hình nhẹ hơn để có thể chạy trên các thiết bị di động.

10/07/2025
Khóa luận tốt nghiệp khoa học máy tính đánh giá hiệu quả của kiến trúc cnn transformers trong việc phân loại tình trạng hư hỏng đường bộ

Trích đoạn nội dung tài liệu

NGUYEN MINH LY- 20521592 KHOA LUAN TOT NGHIEP ĐÁNH GIA HIEU QUA CUA KIÊN TRÚC CNN- TRANSFORMERS TRONG VIEC PHAN LOAI TINH TRANG HU HONG DUONG BO (EVALUATING THE EFFECTIVENESS OF CNN-TRANSFORMERS ARCHITECTURE IN CLASSIFYING ROAD DAMAGE CONDITIONS) CU NHAN NGANH KHOA HOC MAY TINH GIANG VIEN HUONG DAN TS. LE KIM HUNG TP. HO CHI MINH, 2024 Dissertation Defense Committee List The dissertation defense committee, established according to Decision No. 154/QD-DHCNTT dated March 1, 2023, by the Rector of the University of Information Technology.

Acknowledgements First of all, I would like to express my sincere gratitude to all the professors and teachers working and teaching at the University of Information Technology - VNU- HCM for the knowledge, lessons, and valuable experiences that I have acquired during my recent journey. I wish the Department of Computer Science in particular, and the University of Information Technology - VNU-HCM in general, continued brilliant success in the field of education, always training talents for the country, and remaining a firm and attractive destination for future generations of students. I would also like to extend my heartfelt thanks to Dr. Le Kim Hung.

Thanks to his experiences, lessons, care, assistance, and guidance, I have overcome the difficulties and challenges in the process of completing this graduation thesis. Next, I would like to express my sincere thanks to my family for always believing in and encouraging me throughout my studies at the University of Information Technology - VNU-HCM, giving me additional motivation to strive for development and achieve the success I have today. Finally, I would like to thank my fellow students at the University of Information Technology - VNU-HCM for their companionship, assistance, enthusiasm in sharing opinions, and suggestions to help improve and perfect my graduation thesis. Ho Chi Minh, 2024 Table of Contents Acknowledgements 4 ABSTRACT 11 CHAPTER 1.

Objectives of the Thesis 14 1. Subjects and Scope of Research 14 1. Scope of Research 15 CHAPTER 2. Overview of Artificial Intelligence, Machine Learning, Deep Learning 16 2.

Typical Architectures in Deep Learning 21 2. Convolution Neural Network 25 2. Image classification problem 34 2. Some studies on the problem of detecting damaged paths 35 CHAPTER 3.

Overview of the Model 36 3. Zero shot learning 45 CHAPTER 4. EXPERIMENTS AND EVALUATION 50 4. Experimental results 59 CHAPTER 5.

SUMMARY AND DEVELOPMENT DIRECTIONS 69 5. The achieved results 69 5. Development Directions 70 REFERENCES 72 List of images Figure 1. Investigate the overall management theory, applications, and classifications in the fields of AI, ML, and DL.The performance differences among various AI and ML model groups.

Transformers arChIf€CfUT€.Proposed model architecture. The skip connection architecture was invented in resnet. Architecture of the ResNet50 model.---- -5- -sc+c+csereersee 40 Figure 7. Ways to scale up neural network architectures.

The baseline architecture of the EfficientNet model. Compare the accuracy of the models on the ImageNet dataset. The architecture of the EfficientNetB3 model. Image classification with Zero-shot learning.

CLIP training process on 400 million image-caption data pairs. The process of predicting the label of an image using the clip model Figure 14. Distribute data in train, valid and test S€fs. The ratio of images between the classes in the dataset; 0: non-damaged road class; 1: damaged road CÏaSS.- - -- s91 ng ng c 52 Figure 16.

Distribution of images for each country in the train da(a. Distribution of images for each country in the valid data. Distribution of images for each country in the test dafa. Some images from the dataset.

a, b: Images in the case of undamaged roads; c, d: Images in the case of damaged roads. The loss and accuracy scores of the model during training. The GradCAM results display the important regions on an image that the model uses to make Dr€dICfIOTIS.- 5 E111 19112 1231 11 HH HH ni, 65 List of Tables Table 1. Details the number of images in the đafa.

Values of TP, FN, FP, TN in the confusion matr1X. Result for Zero-shot learning for each cÌaSS€S. The overview result for zero-shot learnIng.- --‹--- «<< «++<ex++es 60 Table 7. The number of parameters for each model.

The detailed accuracy metrics of the proposed model using various TT€8SUT€ITRIES.G- 6 5162112001855 111112 T91 nh HT TH HH TH nh nhờ 64 Table 9. The results of comparing the accuracy of our model ("Ours") with the CNN models 1177777.- óc + 2111911 91 931 9 1 9311 v1 ng nhiệt 68 List of Abbreviations Artificial Intelligence 10 ABSTRACT Detecting road damage is an essential and crucial task for ensuring road infrastructure, traffic safety, and maintaining the supply chain for the economy. With the rapid technological development in recent years, especially in automated image processing techniques in the field of artificial intelligence, this thesis researches solutions for automatically detecting road damages. The aim is to improve accuracy and consider reasonable computational resources for practical deployment, meeting the significant global demand.

In this thesis, I propose a new model architecture that utilizes a combination of convolutional neural network (CNN) and transformers for image classification problems. This architecture leverages the strength of CNNs in extracting features from intermediate to advanced levels of input images, and then processes these through transformer encoder blocks. This process uncovers hidden features in the feature vector that CNNs may not detect, thereby enhancing accuracy for image classification. Moreover, we compile various related datasets to create a comprehensive dataset for evaluating the effectiveness of the proposed solution compared to the best available solutions.

Thesis Title Evaluating the effectiveness of CNN-Transformers architecture in classifying road damage conditions. Problem Statement In the current context, the automatic assessment of road damage is not only a technical issue but also an urgent need for society. Cities and traffic management authorities globally are increasingly recognizing the importance of maintaining road infrastructure. This is crucial not only for traffic safety but also significantly contributes to improving transportation efficiency, especially in the context where road transport accounts for a predominant share of freight transportation in many countries.

A reliable statistic shows that the proportion of goods transported by road in European countries accounted for 77.3% in the year 2021. However, current solutions such as human monitoring through video or image analysis from unmanned aerial vehicles, while partly addressing the issue, face many limitations regarding labor and equipment costs, raising questions about their widespread applicability. Additionally, management agencies often lack technological expertise in deploying and maintaining these advanced and modern systems. Recently, new solutions, although applying advanced technologies like machine learning and automatic image processing, still cannot fully meet these needs.

For example, the crack detection model on roads developed by the University of Tokyo, while achieving high accuracy, requires up to 1500ms to process a single image when experimented on a smartphone. This clearly does not meet the real-time requirements in practical applications [2]. Another solution is the work of Yachao Yuan and colleagues, 12 who attempted to address this issue by combining various simple image processing techniques to increase processing speed [3]. Although the model achieved impressive accuracy and prediction speed, the dataset used was not large and diverse enough to conclusively demonstrate the solution's effectiveness across different road types in various countries.

This highlights a significant limitation in developing models that can generalize effectively across various road types worldwide. Machine Learning (ML) and Artificial Intelligence (AI) have opened a new approach to addressing this issue. The combination of AI with image processing techniques holds great potential in automatically and accurately detecting and classifying road damages. However, developing AI models that can generalize well across various environmental and road conditions remains a significant challenge.

Additionally, processing and analyzing the large volume of collected data is a considerable challenge, requiring a smooth integration of advanced image processing techniques and machine learning. In the field of deep learning in recent years, the emergence of the transformer architecture has revolutionized the field of artificial intelligence, particularly in natural language processing. There have been many studies utilizing the power of transformers to improve accuracy in image processing tasks, achieving significant advancements. The superiority of this approach is due to the attention mechanism, which allows the model to view the entire context rather than just focusing on local areas, as is the case with traditional CNNs.

Therefore, a current trend is to research models that combine both CNNs and transformers to create significant technological advancements. The research and development of new solutions are necessary both technically and socially. Advanced research is needed to develop AI models capable of quickly, accurately, and efficiently classifying road damages at a reasonable cost. This will play a crucial role in improving the safety and efficiency of the global road traffic system.

13 In this context, the topic of this thesis aims not only to solve a technical problem but also to contribute to the sustainable development of global road infrastructure. The research and development of AI-based automatic road damage classification solutions are not only a significant step forward in the technical field but also a practical contribution to improving the quality of life and safety of the community. Objectives of the Thesis e Explore the fields of artificial intelligence, machine learning, deep learning, and their applications. e Research several existing studies and methods in the problem of classifying road damage conditions.

e Focus on designing and testing architectures that combine Convolutional Neural Networks (CNN) and Transformers. The goal is to create an optimal solution for the road damage classification problem, leveraging the advantages of both architectures: the powerful image feature recognition capability of CNNs and the ability to process sequential, complex data of Transformers. e Compare and evaluate the performance of the proposed model against models that use only the CNN architecture. e Evaluate the model's ability to deploy the proposed model on an edge device.

Subjects and Scope of Research 1. Research Subjects This thesis focuses on: 14 Modern network architectures in machine learning and artificial intelligence, including ANN (Artificial Neural Networks), CNN (Convolutional Neural Networks), and Transformers. The problem of image classification for the detection of damaged road surfaces in photographs. The feasibility of deploying these models in practical applications, particularly on devices with limited hardware capabilities.

Scope of Research The scope of research includes: Research on existing solutions to address the problem of road damage classification. Research on effective models in the field of image processing, specifically for image classification problems. Overview of Artificial Intelligence, Machine Learning, Deep Learning @ — — — — — ARTIFICIAL INTELLIGENCE =< A technique which enables machines Artificial Intelligence _ -“ to mimic human behaviour kh Machine Learning MACHINE LEARNING max" ”——————^T—T Subset of AI technique which use \ | statistical methods to enable machines x to improve with experience ~~ DEEP LEARNING ~— — — — — — Subset of ML which make the computation of multi-layer neural network feasible Figure 1. Investigate the overall management theory, applications, and classifications in the fields of AI, ML, and DL.

Artificial Intelligence: In the field of computer science, artificial intelligence, or AI, sometimes referred to as synthetic intelligence, is intelligence exhibited by machines, contrasting with the natural intelligence of humans. Typically, the term "artificial intelligence" is used to describe server (or computer) systems capable of mimicking "cognitive" functions often associated with the human mind, such as "learning" and "problem-solving." [4] Artificial intelligence systems are classified based on their ability to replicate human characteristics and are broadly divided into three main types as follows: 16 e Narrow Artificial Intelligence (ANI): At this level, artificial intelligence can only solve problems in a specialized domain, such as image classification or spam email filtering. e Artificial General Intelligence (AGI): This type of intelligence is similar to human capabilities, meaning it can perform tasks that humans can do and can be considered a miniature representation of human intelligence. e Artificial Super Intelligence (ASI): At this level, artificial intelligence surpasses human intelligence in its capabilities.

Nội dung được bảo vệ bản quyền — Tải xuống đầy đủ

Tài liệu "Đánh Giá Hiệu Quả Kiến Trúc CNN-Transformers Trong Phân Loại Tình Trạng Hư Hỏng Đường Bộ" cung cấp cái nhìn sâu sắc về việc áp dụng các kiến trúc học sâu hiện đại, cụ thể là CNN và Transformers, trong việc phân loại tình trạng hư hỏng của đường bộ. Bài viết không chỉ phân tích hiệu quả của các mô hình này mà còn chỉ ra những lợi ích mà chúng mang lại cho việc cải thiện an toàn giao thông và bảo trì cơ sở hạ tầng. Độc giả sẽ tìm thấy thông tin hữu ích về cách mà công nghệ có thể hỗ trợ trong việc phát hiện và phân loại các vấn đề trên đường, từ đó giúp nâng cao chất lượng và độ bền của hệ thống giao thông.

Nếu bạn muốn mở rộng kiến thức của mình về các ứng dụng công nghệ trong lĩnh vực giao thông, hãy tham khảo thêm tài liệu Khóa luận tốt nghiệp kỹ thuật máy tính nghiên cứu và xây dựng mô hình xử lý đồng thời đa tác vụ cho bài toán xe tự hành, nơi bạn sẽ tìm thấy thông tin về các mô hình xử lý đa tác vụ trong xe tự hành. Bên cạnh đó, tài liệu Khóa luận tốt nghiệp khoa học máy tính ứng dụng các thuật toán multiple object tracking dựa trên kỹ thuật học sâu ước tính tốc độ phương tiện giao thông sẽ giúp bạn hiểu rõ hơn về việc theo dõi nhiều đối tượng trong giao thông. Cuối cùng, tài liệu Khóa luận tốt nghiệp truyền thông và mạng máy tính tìm hiểu và đánh giá các kỹ thuật học máy và học sâu sử dụng để nhận diện biển số xe sẽ cung cấp cái nhìn sâu sắc về các kỹ thuật nhận diện biển số xe, một phần quan trọng trong việc quản lý giao thông. Những tài liệu này sẽ giúp bạn có cái nhìn toàn diện hơn về ứng dụng công nghệ trong lĩnh vực giao thông.