VIETNAM NATIONAL UNIVERSITY, HO CHI MINH CITY UNIVERSITY OF INFORMATION TECHNOLOGY COMPUTER SCIENCES NGA PHAM THI - 21521168 GRADUATE THESIS APPLYING MACHINE LEARNING FOR CHILI PEPPER PHENOTYPING AND FEATURE EXTRACTION BACHELOR OF COMPUTER SCIENCE LECTURE PhD. DUNG MAI TIEN HO CHI MINH CITY, 2024 ACKNOWLEDGMENT No one achieves anything great without the help of those around them, whether directly or indirectly. To complete this thesis, I was fortunate to re- ceive much help and support from teachers, colleagues, friends, and family. I would like to dedicate these first pages to express my gratitude to everyone who has accompanied our group during this period.
First of all, I would like to extend my deepest thanks to all the teachers at the University of Information Technology in general and the teachers of the Department of Computer Science in particular. Thanks to the valuable knowl- edge imparted by the teachers, as well as their dedicated support throughout the process, our group was able to complete the thesis and achieve commendable results. I especially want to thank Ph. Dung Mai Tien and Ph.
Tuan Thai Thanh, who inspired, meticulously guided, and provided extensive knowledge, creating a favorable environment for me to learn and exchange ideas with the seniors and peers in the research group. These are invaluable insights and experiences, beneficial not only for this graduation thesis but also for the future work ahead. Finally, I express my heartfelt gratitude to my family and loved ones, who have always been a strong support and consistently backed every decision our group made. Despite having put in a lot of effort to perfect this thesis, it is hard to avoid mistakes and limitations.
I hope to receive sympathy and constructive feedback from the teachers and friends. Ho Chi Minh city, June 24, 2024 Nga Pham Thi, Student 1H Contents Thông tin hội đồng cham khóa luận tốt nghiệp ACKNOWLEDGMENT iii Contents iv List of Figures Vii List of Tables ix List of Abbreviations xi Provisional glossary xiii ABSTRACT 1 INTRODUCTION 1.2 The Objectives and Scope 1. 11 2 PROBLEM FORMULATION 13 iv 2.2 Problem in the Field of Blology.3 Perspective of Computer Vision. ee ee CHILI PEPPER IMAGES AND CHILI PEPPER DATASET 19 3.2 Chili Pepper Cultivation and Sample Selection.
ee ee ee 22 3.4 Chili pepper Dataset for Step] .5 Statistics for the Chili Pepper Dataset. 27 CHILI PEPPERS AND SEEDS DETECTION PROBLEM 29 4.000 eee eee ne 29 4.2 Object Detection Problem. eeee 32 44 YOLO nn*+táồẳ. eee eee ee 44 45.
Q Q Q Q Q Q và và và 44 4.6 Results and Evaluation .02 eee ee ee 47 4. eee ee ee 49 FEATURE EXTRACTING PROBLEM 5.000 eee ee eee 5.2 Pre-PrOC©SSINE.2 Ratio convert Pixel to Millimeter Block .3 Width and Length of Bounding_box Chilipepper .5 Average Width and Length. Q Q Q Q Q HQ ng va 5.7 Degree of Redness .8 Wrinkles of Chili peppersedge.1 By Angles Formed by Three Consecutive Vertices.2 By Smoothness ofContour.3 Using Contour Over a Defined Segment. 0p eee ee ee ee 6 EVALUATION FEATURES EXTRACTING 67 6.00000 00000] 73 7 CONCLUSION AND FUTURE RESEARCH 75 7.Ặ c Q Q Q ee ee ee 75 7.Ặ Q SH Q Q2 77 REFERENCE 81 List of Figures 1.1 Input and Output.1 Unique species code (IT name of each pepper variety).3 Camera and environment setup forimage capture .4 QR barcode to calculated mm/pixel.5 Cropping Chili pepperlmages.6 Chili Pepper DatasetLabel.7 Structured of Chili pepper Dataset.9 Distribution of ObJectCounts.1 Input and Output ofStepl.3 Overview of YOLO.4 Diagram of YOLO architectire.6 The network architecture of Yolov5.
It consists of three parts: (1) Backbone: CSPDarknet, (2) Neck: PANet, and (3) Head: Yolo Layer. The data are first inputted to CSPDarknet for feature ex- traction and then fed to PANet for feature fusion. Finally, YOLO Layer outputs detection results (class, score, location, size). 000002 eee ee ee 42 4.10 Illustration of how to calculate Precision and Recall.11 Illustration of loU Metrics .1 Input of Step2.3 Illustration of Chili pepper mask with only Threshold .4 Illustration of Chili pepper mask after using Closing method .5 Refined Chili Pepper segmentation.6 Illustration of bb_x 2.1 Examples of 330 Consumer.2 Histogram of Feature Extracting .3 Scatter Matrix of Feature Extracting .4 Correlation Matrix of Feature Extracting .5 PCA Feature Extracting.
74 List of Tables 3.1 Distribution of ObJectCounts.2 Evaluation with classSeed .3 Evaluation Overall Model Evaluation.1 Statistical Summary of Feature Extracting. 69 1X List of Abbreviations Ph. Doctor of Philosophy XI Provisional ølossary Machine Learning Học máy Deep Learning Học sâu xiii ABSTRACT In the current agricultural sector, identifying phenotypes and accurately de- scribing the morphology of chili peppers involve manual inspections and mea- surements performed by trained personnel. This process is labor-intensive, time-consuming, and prone to errors due to subjective biases and human mis- takes.
With the rapid advancements in computer vision and machine learning, we propose a method that utilizes machine learning and computer vision to au- tomate the process of phenotype identification and feature extraction of chili peppers. Additionally, we aim to establish a dataset for information retrieval regarding various chili pepper varieties. This study is supported by a secure dataset provided through the collaboration between the Department of Com- puter Science and the BIO-RESOURCE COMPUTING RESEARCH CENTER of Jeju National University, South Korea. Our approach involves using computer vision and machine learning tech- niques to automatically extract features from images to store these characteris- tics for each chili pepper variety.
This is highly beneficial for managing pepper breeding and reproduction in the biological field, catering to the expansive mar- ket for chili peppers today. To address the outlined problem, we will divide it into smaller sub-problems for step-by-step resolution, including localization, image processing, and feature extraction. Subsequently, we will address the application and implementation of the initial objectives. In summary, this thesis accomplishes the following: 1.
Constructing a dataset from images of chili peppers cultivated by biolo- 1 gists at the research institute. For the localization and seed detection of chili peppers in images, to meet real-time conditions, we propose using the one-stage object detec- tion model YOLO [11]. For feature extraction post-chili identification, we will utilize image pro- cessing techniques within the computer vision domain, which will be elaborated on later. Chapter 1 INTRODUCTION In this chapter, we will provide an overview of the problem of APPLYING MA- CHINE LEARNING FOR CHILI PEPPER PHENOTYPING AND FEA- TURE EXTRACTION, along with the challenges encountered during the im- plementation of this project.
Subsequently, we will summarize the subjects, scope, and research objectives of this thesis. At the end of the chapter, we will present the accomplished work and the main structure of the thesis.1 Problem statement Chili peppers, cultivated worldwide and used for thousands of years, are spicy fruits belonging to the Solanaceae family. They are highly valued for their unique flavor, nutritional properties, and medicinal benefits. Chili peppers are rich in various vitamins, including vitamins E, C, A, and B complex, as well as minerals such as thiamine, folate, molybdenum, manganese, potassium, calcium, and iron.
[2] Additionally, they contain polyphenols (mainly luteolin), flavonoids, and quercetin. In many regions, chili peppers play a crucial role in local cuisine, pro- viding unique flavors and adding depth to traditional dishes. Beyond culinary applications, chili peppers are used in various industries, including pharmaceu- ticals, cosmetics, and even self-defense products, due to their capsaicinoid [12] content—the compound responsible for their characteristic spiciness. Problem statement The chili pepper market has seen significant growth, driven by increasing consumer preference for diverse and authentic flavors, as well as the recogni- tion of the potential health benefits associated with capsaicinoids [12].
Their widespread use as a spice and functional food ingredient has increased global demand for both fresh and processed chili products, creating opportunities for growers, processors, and traders. Furthermore, chili peppers have become an important component in many industrial applications. Their popularity extends beyond culinary use, as they are utilized in pharmaceuticals, cosmetics, and even self-defense products due to the presence of capsaicinoids [12]. The vast diversity of chili pepper varieties presents numerous challenges.
With many types of chili peppers available, each having distinct phenotypic traits in terms of shape, size, color, spiciness, and flavor, it brings opportuni- ties and challenges for breeding programs and variety management. Accurate and efficient characterization of these phenotypic traits is crucial for unlock- ing the full potential of chili pepper varieties and promoting targeted breeding efforts. Traditional methods of phenotyping and accurately describing chili pep- per morphology involve manual inspections and measurements by trained per- sonnel, which are labor-intensive, time-consuming, and prone to human error and subjectivity. Moreover, these manual methods often lack the precision and consistency required for comprehensive analysis and comparison of varieties.
Digitizing phenotypic traits through advanced imaging techniques and computer- assisted analysis offers a transformative solution to these challenges. Researchers can quantify and extract numerical features with unprecedented accuracy and objectivity by capturing high-resolution images of chili peppers and leverag- ing machine learning algorithms. This digital approach facilitates precise mea- surements of traits such as fruit size, seed count, color parameters, as well as other relevant morphological and biochemical characteristics. The obtained dig- ital data enables detailed variety profiling and supports data-driven decision- making in breeding programs.
Moreover, digitizing phenotypic traits allows the creation of comprehensive databases, enabling efficient storage, retrieval, and analysis of varietal information. This data-driven approach allows breeders to 4 1. Problem statement identify desirable traits, assess genetic diversity, and make informed choices for developing new varieties that meet market demands or specific environmental conditions. By adopting phenotypic digitization, the chili pepper industry can unlock new avenues for variety management, accelerate breeding cycles, and foster the development of improved varieties to meet the growing demands of consumers and stakeholders.
For these reasons, we were motivated to undertake the project “Applying Machine Learning for Chili Pepper Phenotyping and Feature Extraction” The project is divided into several sub-problems, which we will discuss later. First, we need to describe this project: ¢ We will first define the phenotypic traits that need to be extracted. ¢ Input: Images of chili peppers from which we want to extract informa- tion, including QR barcode(Figure 3. ¢ Output: The phenotypic characteristics of chili peppers digitized into : €sv File telude 4 column | - ( —"me of Seeds ¥ Width b_bo: “BE b “Ss Avg with Chili tr enl: K Y Degree of Rerlne:Sie 004 _IT158286_1_1.1: Input and Output Based on our limited understanding during the execution of this thesis, we realized that there are no scientific papers on the image processing of chili pep- 5 1.
Problem statement pers for information extraction. Our team decided to define this problem by breaking it down into sequential sub-problems. These include the following tasks: ¢ Identifying chili peppers and their seeds, which lays the foundation for subsequent information extraction steps. ¢ Defining the extractable information fields, specifically, we can extract the following eight pieces of information: 1.
Width and Length of the chili pepper’s bounding box 2. Average Width and Length of the Chili pepper 3. Area of the Chili pepper 4. Degree of Redness of the Chili pepper 5.
Number of seeds in a Chili pepper 6. Wrinkle of the Chili Pepper’s Edge e Image processing to extract phenotypic characteristics. Analyzing these sub-problems allows us to find solutions to object detection problems easily. For the object detection problem, many models have already been developed to solve similar issues for other types of fruits.
Specifically, for the feature extraction problem, we will analyze features and find ways to extract these types of information from images. Therefore, what we need to do is an- alyze algorithms and pattern similarities to apply existing solutions to our sub- problems.