Ứng dụng Deep Learning trong nhận diện sâu bệnh trên lá cà phê – Luận văn Đại học Công nghệ TP.HCM

Ứng dụng học sâu trong việc phát hiện sâu bệnh và bệnh lý trên lá cà phê, nâng cao hiệu quả sản xuất nông nghiệp và bảo vệ cây trồng.

Chuyên ngành

Khoa học Máy tính

Người đăng

Ẩn danh

Thể loại

Luận văn

2024

54
1
0

Phí lưu trữ

30 Point

Tóm tắt

I. Giới thiệu về Deep Learning trong Phát hiện Bệnh Lá Cà phê

Deep learning đã trở thành giải pháp tiên tiến trong lĩnh vực phát hiện bệnh trên cây trồng, đặc biệt là bệnh lá cà phê. Cà phê là sản phẩm nông nghiệp toàn cầu, được tiêu thụ bởi 30-40% dân số thế giới với hơn 178,5 triệu bao 60 kg được tiêu thụ trong năm 2023/2024. Tuy nhiên, cây cà phê dễ bị tấn công bởi các loài sâu bệnh khác nhau, gây hại đáng kể đến năng suất và hiệu suất kinh tế. Công nghệ học sâu (deep learning) cung cấp khả năng phát hiện sớm các dấu hiệu bệnh thông qua phân tích hình ảnh, giúp nông dân thực hiện các biện pháp phòng chống kịp thời. Ứng dụng này không chỉ nâng cao hiệu quả canh tác mà còn giảm thiểu tổn thất kinh tế cho người nông dân.

1.1. Tầm quan trọng của Phát hiện Sâu Bệnh Sớm

Phát hiện sâu bệnh ở giai đoạn sớm là chìa khóa để bảo vệ năng suất cà phê. Khi các bệnh như Miner (sâu xỉng) và Brown eye spot (đốm nâu mắt) được phát hiện sớm, nông dân có thể áp dụng các biện pháp quản lý hiệu quả, từ đó giảm thiểu tổn hại. Công nghệ phân tích hình ảnh qua deep learning giúp xác định chính xác các loại bệnh trên lá, giảm chi phí điều trị và nâng cao chất lượng sản phẩm cuối cùng.

1.2. Vai trò của Nông nghiệp Thông minh

Nông nghiệp hiện đại đòi hỏi các giải pháp công nghệ tiên tiến để tối ưu hóa sản lượng. Deep learningartificial intelligence (AI) là những công cụ quan trọng trong nông nghiệp thông minh, cho phép nông dân giám sát sức khỏe cây trồng một cách hiệu quả. Các ứng dụng di động tích hợp mô hình học sâu giúp nông dân nhanh chóng xác định vấn đề và đưa ra quyết định canh tác đúng đắn.

II. Kiến trúc Deep Learning và Mô hình YOLOv10

Mô hình YOLOv10 (You Only Look Once version 10) là một trong những kiến trúc deep learning tiên tiến nhất hiện nay, được thiết kế để phát hiện đối tượng trong hình ảnh với độ chính xác cao. Khác với các mô hình trước đó, YOLOv10 sử dụng kiến trúc mạng nơron sâu với nhiều lớp tích chập và cơ chế chú ý để cải thiện độ nhạy. Dual Label Assignment là một tính năng nổi bật của YOLOv10, giúp tối ưu hóa quá trình huấn luyện mô hình. Kiến trúc này cho phép phát hiện các dấu hiệu bệnh trên lá cà phê với tốc độ nhanh và độ chính xác cao, rất phù hợp cho ứng dụng di động trên nông trại.

2.1. Thành phần Cơ bản của Mạng Nơron Sâu

Mạng nơron sâu bao gồm nhiều thành phần quan trọng: Perceptron là khối xây dựng cơ bản, trong khi các hàm kích hoạt như Logistic functionHyperbolic tangent giúp mạng học các mối quan hệ phi tuyến tính. Các lớp tích chập (Convolutional layers) xử lý dữ liệu hình ảnh, và pooling layers giảm độ phức tạp tính toán mà vẫn giữ các đặc trưng quan trọng.

2.2. Quá trình Huấn luyện và Tối ưu hóa Mô hình

Quá trình huấn luyện mô hình YOLOv10 yêu cầu bộ dữ liệu cân bằng và đủ lớn. Các kỹ thuật data augmentation như điều chỉnh độ sáng (+40%), thêm nhiễu (±10%) được áp dụng để cải thiện khả năng tổng quát hóa. Sau khi huấn luyện, mô hình cần được prune (cắt tỉa) để giảm kích thước file và tăng tốc độ xử lý trên điện thoại di động.

III. Tập Dữ liệu và Kỹ thuật Tiền xử lý

Chất lượng và kích thước của tập dữ liệu là yếu tố quan trọng quyết định hiệu suất của mô hình deep learning trong phát hiện bệnh lá cà phê. Tập dữ liệu ban đầu bao gồm hai lớp chính: Miner (sâu xỉng) được đánh dấu màu tím và Brown eye spot (đốm nâu mắt) được đánh dấu màu vàng. Ngoài ra, lớp Normal (bình thường) cũng được bao gồm để giúp mô hình phân biệt các lá khỏe mạnh. Các kỹ thuật tiền xử lý dữ liệu bao gồm chuẩn hóa kích thước hình ảnh, data augmentation để tăng sự đa dạng của tập huấn luyện, và normalization để cải thiện hiệu suất hội tụ.

3.1. Phân loại và Cân bằng Dữ liệu

Phân loại dữ liệu là bước đầu tiên trong xây dựng bộ dữ liệu hiệu quả. Bộ dữ liệu tFlite được tổ chức thành các thư mục theo từng lớp bệnh. Sự cân bằng giữa các lớp rất quan trọng - nếu một lớp có quá nhiều mẫu so với lớp khác, mô hình sẽ bị bias (thiên lệch). Các kỹ thuật resamplingsynthetic data generation có thể được sử dụng để cân bằng phân bố dữ liệu.

3.2. Kỹ thuật Augmentation và Chuẩn hóa Dữ liệu

Data augmentation bao gồm các phép biến đổi như xoay ảnh, thay đổi độ sáng, thêm nhiễu, và flip ảnh để tăng sự đa dạng của tập dữ liệu. Các phép biến đổi này giúp mô hình generalize tốt hơn trên dữ liệu mới. Chuẩn hóa dữ liệu đảm bảo rằng các giá trị pixel được chuẩn hóa về khoảng [0,1] hoặc [-1,1], giúp quá trình huấn luyện hội tụ nhanh hơn.

IV. Ứng dụng Di động và Triển khai Thực tế

Ứng dụng di động tích hợp mô hình deep learning giúp nông dân có thể phát hiện bệnh lá cà phê trực tiếp trên nông trại bằng điện thoại thông minh. Ứng dụng cung cấp giao diện thân thiện với người dùng bao gồm các tính năng như chụp hình và phân tích, lịch sử phát hiện, và cảnh báo sâu bệnh. Mô hình được tối ưu hóa thành định dạng TensorFlow Lite (tFlite) để có thể chạy trên các thiết bị di động với hiệu suất cao và tiêu thụ pin thấp. Giao diện bao gồm các trang Home, Profile, History cho phép người dùng quản lý và theo dõi các phát hiện bệnh theo thời gian.

4.1. Tính năng Chính của Ứng dụng

Ứng dụng cung cấp nhiều tính năng hữu ích: Chữ ký xác thực cho phép nông dân đăng nhập an toàn, chụp và phân tích hình ảnh giúp phát hiện bệnh ngay lập tức, lịch sử hoạt động lưu trữ các phát hiện trước đó. Trang Profile cho phép cập nhật thông tin cá nhân, Change password bảo vệ tài khoản, và trang History cung cấp báo cáo chi tiết về các bệnh được phát hiện.

4.2. Triển khai Mô hình trên Thiết bị Di động

TensorFlow Lite là công cụ quan trọng cho việc triển khai mô hình deep learning trên điện thoại di động. Mô hình được chuyển đổi từ định dạng huấn luyện sang định dạng .tflite nhẹ hơn, đảm bảo tốc độ xử lý nhanh. Confusion matrix được sử dụng để đánh giá độ chính xác của mô hình, giúp xác nhận rằng ứng dụng có thể phát hiện bệnh một cách đáng tin cậy trước khi triển khai thực tế.

11/12/2025
Luận văn application of deep learning in detecting pests and diseases on coffee leaves

Trích đoạn nội dung tài liệu

Declaration of Authenticity I declare that this research is our own, carried out under the supervision of Assoc. Le Hong Trang and Mr. Nguyen Quang Duc. The results of our study are credible and have not yet been made public.

All materials utilized in this research were gathered by myself from various sources and are properly cited in the reference section. Furthermore, all of the research results are properly referenced and unrelated to the initial data. In any event, I stand by my actions and accept responsibility for any plagiarism. Thus, any copyright violations resulting from our research are not the responsibility of Univer- sity of Technology-Vietnam National University Ho Chi Minh City.

Ho Chi Minh City, Dec 2024 Project Author Ly Kim Phong 1 Acknowledgment First and foremost, I want to show my heartfelt gratitude to my parents and friends, particularly my parents, who have always supported and encouraged me on my path. They are the ones who have always been there when I need them the most, they are my source of inspiration, drive, and strength to go through all the hardships in my life. I am very grateful to have them by my side. Second, I want to express the appreciation to my advisers, Assoc.

Le Hong Trang and Mr. Nguyen Quang Duc. Without their assistance and support, I would not have been able to finish my paper effectively; they have been incredibly gracious and patient in guiding me through problems. Third, I also want to express my deepest thank to Dr.

Nguyen Duc Dung, my re- viewer. He has pointed out my mistakes and shortcomings and provides few direct ad- vices to improve my work. In closing, I would like to thank all the teachers, TA and the Department of Com- puter Science for their assistance and support in getting me ready for this project; their opinions and assessment have been invaluable. Without their help, I could not have com- pleted this job.

They provided guidance for the path of my studies. One more time, I would like to express my gratitude and admiration to everyone who has helped and inspired me. Thank you to everyone. 2 Abstract In order to develop a basic mobile application for farmers to identify diseases on cof- fee leaves in their early stages, the topic has been researched.

The goal of this study is to develop deep learning tools / AI for the identification and early warning of pest in- festations in coffee plants. The theory of disease detection systems in the modern world has been summarized in this dissertation, which also reviews various popular methods of early detection employing deep learning models, machine learning, computer vision, and hardware monitoring. We provide a one-stage model of YOLOv10 based on ref- erence research. In order to optimize the model and further prune the models, further enhancements for this project require obtaining more balanced datasets.4 Structure of this project .1 Deep Learning Neural Network .2 Basic components of a neural network .3 How does a Deep Learning Neural Network work.

32 4 Mobile app integration 35 4. 41 4 5 Conclusion 46 6 Future work 47 5 List of Figures 2.1 The basic building block of Deep Learning models - Perceptron .3 Hyperbolic tangent function [7] .4 Logistic-curve function [15] .6 Deep Neural Network architecture [4] .10 YOLOv10 architecture Dual Label Assignment .11 Example of YOLOv10 architecture .12 Example of PR curve .1 Original dataset with 2 classes: Miner (purple) and Brown eye spot (yel- lowish) .3 Example of tflite dataset with normal class .2 tflite data distribution .4 Original and +40% brightness applied .5 Original and 10% noise applied .7 Confusion matrix normalized .1 Use-case diagram .5 Login/Sign up .7 Home & Recent activities & Profile pages .9 Update info & Change password & History pages. 45 6 List of Tables 3.1 tflite class distribution .2 List of augments .1 Motivation Coffee is consumed daily by 30-40% of the global population, with more than 178.5 million 60-kilogram bags reported to be consumed in 2023/2024. However, coffee plants are highly susceptible to various pests and diseases, which pose a significant threat to crop health and yield.

Early detection and effective management of these challenges are crucial for sustaining productivity. Our research aims to develop an innovative application capable of identifying pests in coffee plants through image analysis of roots, stems, and leaves. This tool is designed to provide farmers with timely alerts, enabling them to implement preventive measures and mitigate potential losses. The motivation for this project stems from several key factors.

Coffee is a vital agri- cultural product, cultivated extensively in numerous countries and contributing signifi- cantly to global economies. Pests and diseases, if left unchecked, can drastically reduce yields and profitability. In addition, there is an urgent need for efficient and accessible solutions to facilitate early threat detection for farmers. We are confident that this application will create a transformative impact on the coffee industry by enhancing productivity, protecting farmers’ livelihoods, and contributing to a sustainable coffee supply chain that benefits all stakeholders.2 Problem statement Diseases such as coffee rust, coffee berry disease, and pink disease pose significant threats to coffee plants by infecting leaves and impairing the ripening process, ultimately reducing coffee bean yields.

Early detection of these diseases is critical, and this can be achieved through automated systems capable of identifying symptoms at their initial stages. Currently, most leaf disease detection systems rely on convolutional neural networks (CNNs) and their variants, including R-CNN, F-CNN, Faster-CNN, SSD, and YOLO, to identify disease-induced damage. While these methods are effective, they heavily de- pend on accurately labeled datasets and face limitations in adaptability. Specifically, the addition of new disease variants often necessitates retraining the model from the begin- ning, which is both time-consuming and resource-intensive.

To address these limitations, first we propose a one-stage approach using the YOLOv10x general object detection model to directly identify leaf diseases. This method also cal- culates the probability of various diseases, such as rust and miner infestations. Our ap- proach enables seamless incorporation of new disease classes by simply adding a clas- 8 sification label, eliminating the need for complete model retraining. Then we develop a mobile application that takes input (coffee leaves) images from farmers, detects the diseases of that images and provides the details and recommends treatments for that diseases.

This report presents the experiments conducted and the findings that inform the de- velopment of a robust, scalable model for our application.3 Scope In this project, my aim is to achieve these goals: 1. Assessing current document matching methods and making the necessary adjust- ments to determine which ones best meet our needs. Using the latest model to get a more accurate model architecture. Developing a basic mobile application to demonstrate our model.4 Structure of this project The rest of this paper is organized as follows.1, we recall some back- grounds on deep neural network, how it works, and CNN introduction.2 is a brief introduction of relevant models.4 are the training losses and eval- uating metrics used to evaluate the model’s performance.

Chapter 3 is about the dataset being used in this project, our training and evaluating results. Chapter 4 discusses about what we have done to develop the application. Chapter 5 and 6 discusses the results and future improvements that can be made.1 Deep Learning Neural Network Before going to what we have done on this project, we need some basic knowledge about Deep Learning Neutral Network. First of all, Machine Learning is a field of study in artificial intelligence that en- ables algorithms to uncover hidden patterns within datasets, allowing them to make predictions on new, similar data without explicit programming for each task.[14] And Deep Learning, conceptualized by Geoffrey Hinton in the 1980s, is a subset of Machine Learning that uses some functions to map input into output.

These functions will form a relationship between the input and the output by extracting essential information from the input data. This is called learning, and the process of learning is training.[5] Next, about Neural Networks, also known as Artificial Neural Networks, it was also created by Hinton, which is a Deep Learning algorithm structured similar to the orga- nization of Neurons in the brain. Hinton took this approach because the human brain is arguably the most powerful computational engine known today. This network consists of 3 layers of perceptron: input, hidden and output layer.

Before going to output layer, from hidden layer, the calculations of inter- connected nodes must be carried out and between those calculations, in order to avoid overfitting problem when training, weights, biases and activation functions are added. In the next section, we will introduce more details about the basic components of a neural network, how does a Deep Learning Neural Network work, what are CNN and why we use CNN for our problem.2 Basic components of a neural network First, let us have a look at a perceptron or neuron: Figure 2.1: The basic building block of Deep Learning models - Perceptron 10 Inputs: They are passed on to a neural network to make predictions, they are presented as features of a dataset. Weights: They are important real values associated with the inputs that tell the signif- icant of the feature passed in. Bias: Its mission is to shift the activation function across the plane towards either left or right.

More information will be explained later. Sum: It is a function to add up the product of the weight and the input with bias. Layers: Layers in a deep learning model form the fundamental components of its ar- chitecture. They process data sequentially, where each layer takes input from the previous one and passes the output to the next.[22] There are several types of layer we want to declare in this project: • Dense layer: also called a fully connected layer, uses a linear operation to mainly transform the dimensionality of the input to fit the desired output (e., classification probabilities in the final layer), but sometimes it is used to aggregate and process information.

Below is its mathematical operation. For an input vector x: y = σ (W x + b) where σ is the activation function, W is the weight matrix, b is the bias, and y is the output.[19] • Pooling layer: Used for scaling down the input. • Normalization layer: Normalization layers are components in neural networks that help to stabilize and hasten the training process by normalizing the inputs to a layer.[16] They are particularly useful in deep learning architectures, including con- volutional neural networks (CNNs) and fully connected networks. • Convolutional layer: A convolutional layer is a fundamental component of con- volutional neural networks (CNNs), primarily used for processing and analyzing visual data.

It applies a mathematical operation called convolution to the input data, utilizing filters (or kernels) to extract features such as edges, textures, and patterns. This process helps the network learn hierarchical representations of the input data.2: Convolution layer [17] Activation function: It is used to add non-linearity to the model. Here are some acti- vation functions that are commonly used: • Tanh function: Tanh or Hyperbolic Tangent function commonly denoted as ( tanh(x)) is a mathematical function that is widely used as an activation function in neural networks. It is defined as the ratio of the hyperbolic sine and hyperbolic cosine functions[9]: sinh(x) ex − e−x tanh(x) = = cosh(x) ex + e−x Figure 2.3: Hyperbolic tangent function [7] 12 • Sigmoid / Logistic function: The sigmoid function, also known as the logistic func- tion, is a mathematical function that maps any real-valued number into a value be- tween 0 and 1.

It is defined by the following formula[15]: 1 σ (x) = 1 + e−x Figure 2.4: Logistic-curve function [15] • ReLU function: The Rectified Linear Unit (ReLU) function is a widely used acti- vation function in deep neural networks.

Nội dung được bảo vệ bản quyền — Tải xuống đầy đủ