HỒ SỸ TRỌNG KIÊN CROSS DOMAIN (FEW SHOT) OBJECT DETECTION BENCHMARK LUẬN VĂN THẠC SĨ NGÀNH ĐIỆN TỬ, NĂNG LƯỢNG ĐIỆN, TỰ ĐỘNG HÓA CHUYÊN NGÀNH KỸ THUẬT TRUYỀN THÔNG VÀ DỮ LIỆU NGƯỜI HƯỚNG DẪN KHOA HỌC GS.TS Anissa Mokraoui GS.TS Pierre Duhamel TS. Nguyễn Hồng Thịnh PARIS, NĂM 2023 Cross Domain (Few Shot) Object Detection Benchmark Master thesis of Paris-Saclay University and VNU University of Engineering and Technology Specialization: M2 Data and Communication Engineering Research unit: Le Centre Borelli, École Normale Supérieure, Université Paris-Saclay Thesis presented at Gif-sur-Yvette, on 21 December 2023 HO Sy Trong Kien Committee Arnaud BOURNEL Paris-Saclay University Chairman Pierre DUHAMEL CNRS, CentraleSupelec, Paris-Saclay University Rapporteur Jocelyn FIORINA CNRS, CentraleSupelec, Paris-Saclay University Examniner Anissa MOKRAOUI L2TI, Sorbonne Paris Nord University Examiner NGUYEN Linh Trung VNU University of Engineering and Technology Examiner Thesis Supervision Master Thesis Anissa MOKRAOUI L2TI, Sorbonne Paris Nord University Supervisor. Pierre DUHAMEL CNRS, CentraleSupelec, Paris-Saclay University Co-supervisor. Hong Thinh NGUYEN VNU University of Engineering and Technology Co-supervisor.
Laurent OUDRE École Normale Supérieure, Paris-Saclay University Co-supervisor. Acknowledgements I express my heartfelt gratitude to the professors, students, staff, and institutions whose pivotal roles contributed to the success of my Paris Saclay M2 Master program at the University of Engineering and Technology and Université Paris-Saclay. Foremost, I extend my deepest appreciation to my esteemed professors: Prof. Pierre DUHAMEL at Centrale Supélec, Prof.
NGUYEN Linh Trung, and Dr. NGUYEN Hong Thinh at VNU - University of Engineering and Technology, Prof. Anissa MOKRAOUI at Sorbonne Paris Nord University, Prof. Laurent Oudre at École Normale Supérieure, and Prof.
Arnauld Bournel at Université Paris-Saclay. Their kindness, patience, and guidance have been invaluable throughout the program, culminating in the successful completion of the learning semester and this memorable internship in France. I am honored to have been supervised by Prof. Pierre DUHAMEL, Prof.
Anissa MOKRAOUI, and Dr. NGUYEN Hong Thinh for this master’s thesis. Without their guidance and support, the accomplishment of this thesis would not have been possible. I also wish to express my gratitude to Prof.
Laurent Oudre and the Centre Borelli at École Normale Supérieure for providing the necessary facilities and a conducive working environment during the internship program. A special thank you to my dear mother and friends for your unwavering support; I deeply appreciate your help during this time, and I am always grateful for your support. Authorship I solemnly declare that my thesis, titled ’Cross Domain (Few Shot) Object Detection Benchmark’, is my own research work conducted under the guidance of Prof. Anissa MOKRAOUI, Prof.
Pierre DUHAMEL, Prof. Hong Thinh NGUYEN and Prof. The sources used in the thesis are explicitly mentioned in the reference section, with proper citations. The data and results presented in the thesis are entirely truthful, and there is no copying from the works of others.
If any discrepancies are found, I take full responsibility and am subject to any disciplinary actions imposed by the university., 2023 Student Ho Sy Trong Kien Contents 1 Introduction 6 2 Background and related works 8 2.1 Two-stage detectors .2 One-stage detectors .2 Classification and Regression Criteria. 15 3 Cross Domain (Few shot) Benchmark for FCOS 18 3.2 Source Domain Dataset .3 Evaluation results for DeepFruits .1 Mango dataset Evaluation .1 The Training Phase .2 Apple Dataset Evaluation .1 The training phase .3 Orange dataset evaluation .1 The Training Phase .4 Avocado dataset evaluation .1 The Training Phase .5 Capsicum dataset evaluation .1 The Training Phase .4 DIOR Dataset Evaluation .2 The Training Phase .3 The Testing Result .5 Oktoberfest Dataset Evaluation .2 The Training Phase .3 The Testing Result .6 ClipArt1k Dataset Evaluation .2 The Trainging Phase .3 The Test Results .7 ImageNet Dataset Evaluation .2 The Trainging Phase .3 The Test Results .8 IWildCam Dataset Evaluation .2 The Trainging Phase. 42 5 Conclusion and future works 43 5.1 FCOS Architecture Details. 46 2 List of Figures 2.1 Two architectures comparsion - taken from [27] .2 Yolo v1 - taken from [16] .3 Retinanet Architecture - taken from [14] .4 FCOS architecture - taken from [21] .5 The lateral connection and the top-down pathway - taken from [13] .6 Ground Truth Box - taken from [13] .7 Moduling factor influence - taken from [14] .1 Fine-tuning process on source and target dataset .1 Source Dataset: MS-COCO - Target Dataset: DeepFruits .2 Source Dataset: MS-COCO - Target Dataset: DeepFruits .3 Source Dataset: MS-COCO - Target Dataset: DeepFruits .4 Source Dataset: MS-COCO - Target Dataset: DeepFruits .5 Source Dataset: MS-COCO - Target Dataset: DeepFruits .6 Source Dataset: MS-COCO - Target Dataset: DIOR .7 Source Dataset: MS-COCO - Target Dataset: Oktoberfest .8 Source Dataset: MS-COCO - Target Dataset: Clipart1k .9 Source Dataset: MS-COCO - Target Dataset: ImageNet .1 Resnet50 Architecture - taken from [23] .2 Residual block 1 - Kahe et al [25] .3 FCOS head - taken from [21].
54 3 List of Tables 3.1 Domain Dataset Information .1 Distribution of train/set data .2 F1 score for fruit detect .3 mAP of 5 fruits at diffrent IoU threshold .4 F1-score of 5 fruits at diffrent IoU threshold .5 F1 score comparison between FCOS and Faster R-CNN .6 Object Instances in 20 Categories .7 Overall AP Metrics .8 Average Precision (AP) for Different Categories .9 Training Data Class Statistics .10 Overall AP Metrics .11 Per-Category AP Metrics .13 Per-category AP metric .14 Training dataset per-category instances .15 Average Precision Scores .16 AP Scores for Different Categories .17 Mean Average Precision Scores .18 AP Scores for Different Categories. 42 4 Notations and Abbreviations Notations R Set of real numbers H Height W Width C Number of categories Fi Feature maps at level i s Total stride TP True Positive TN True Negative FP False Positive FN False Negative AP Average Precision AR Average Recall mAP mean Average Precision mAP50 mean Average Precision at IoU threshold of 0.5 mAP75 mean Average Precision at IoU threshold of 0.75 mAP(small) mean Average Precion for small objects with size less than 32*32 mAP(medium) mean Average Precision for medium objects with size greater than 32 *32 and smaller than 96*96 mAP(large) mean Average Precision for large objects with size > 96*96 5 Chapter 1 Introduction Object detection, an essential and challenging task in computer vision, involves a two-step process: localizing in- stances within an image using bounding boxes and classify them among a predefined set of categories. With the ever-expanding landscape of technology and the emergence of mainstream models like Faster R-CNN [17], YOLO [16], and RetinaNet [14], object detection has undeniably proven its worth in addressing real-world challenges. While object detection models serve as powerful tools, they are not without their limitations and specific requirements.
In fact, the effectiveness of an object detector in providing accurate predictions for instance types in a given en- vironment does not ensure comparable accuracy in diverse settings, marked by variations in lighting, viewpoints, and weather conditions. To achieve optimal model efficiency, it usually demands learning from huge dataset with various scenarios to extract general features and make predictions on unseen targets. In real-world scenarios, these limitations may arise in tasks associated with specific domains experiencing data scarcity or when the costs associ- ated with labeling become excessively high. Cross-domain methods in machine learning is about the creation of new models or modification of optimal ar- chitectures to improve the generalization capability across diverse domains or datasets.
The goal is to ensure that a well-trained model on a source domain will perform well on a target domain, even with differences between the characteristics of the two domains. This becomes especially critical when there is an abundance of labeled data in one domain but a scarcity of such data in another [10]. Large labeled datasets such as MS-COCO [12], ImageNet [4], or PascalVOC [5] have served as the cornerstone for numerous renowned and state-of-the-art architectures. No- table examples include AlexNet [11], VGG16 [19], ResNet, EfficientNet [20], RetinaNet [14], and Mask R-CNN [8].
The method we propose in this study is using a model pre-trained on large dataset :MS-COCO 2017 [12] as the source domain, then use a transfer learning approach: fine-tuned the model last layer using data of the target domain. Transfer learning serves as a method to address the problems associated with cross-domain adaptation. Based 6 upon the concept of utilizing knowledge gleaned from a pre-trained model on a sizable labeled dataset (specifically, MS-COCO [12] in our study), the approach seeks to enhance model performance when applied to target domains. Fine-tuning last layer, a specific strategy within transfer learning, refines this process by maintaining fixed param- eters learned from previous layers, while training only the weights of the last layer on a subset of target datasets.
This nuanced approach allows the model to tailor its understanding to the unique characteristics of the new target domain [24], facilitating improved performance on specific tasks within that domain and building upon the founda- tional knowledge acquired during pre-training. The base of our will be upon FCOS model (Full Convolution One Stage), which has been introduced by Zhi Tian and his colleagues in 2019 [21]. FCOS [21] is an upgraded version of Retinanet [14] with some modifications to improve the mean Average Preicion. Notably, FCOS [21] take each spatial location on the image as training sample instead of anchor box like RetinaNet [14], transforming it into an anchor-free model.
This modification suggests increased adaptability to varying instance sizes across diverse domains. In addition, FCOS [21] also inherits the implementation of focal loss from RetinaNet [14], a strategy previously validated as effective in handling imbalanced datasets. By incorporating the last layer fine-tuning method, our study seeks to examine the performance of FCOS [21] across datasets originating from distinct domains. In particular, we aim to understand how variations in characteristics between source and target domains impact the overall performance.
The remainder of this report will be organized as follows: • Chapter 2 will provide an overview of related work and background knowledge. • Chapter 3 will detail our work approach and benchmark criterion • Chapter 4 we will present our experimental setup and results. • Chapter 5 will conclude the paper and discuss potential plans for future research. 7 Chapter 2 Background and related works 2.1 Related works Object detection has made remarkable success recently.
However, it originated from rudimentary concepts, notably the sliding windows paradigm, where a kernel is systematically applied to a dense image grid [14]. A pivotal leap was when LeCun et al.[25] successfully applied Convolutional Neural Networks (CNNs) to hand-written digit recogni- tion.To the present, the rapid evolution of Deep Learning and CNNs has resulted in the construction of many optimal architectures and proving their worth in solving different tasks. Object detection models, in terms of architecture, can be categorized into two main types: one-stage detectors and two-stage detectors [27].1: Two architectures comparsion - taken from [27] 8 2.1 Two-stage detectors R-CNN family architecture As shown in Figure 2.1 [27], two-stage detection models proceed with an input image in two distinct phases. The first phase involves the generation of region proposals, where the model identifies areas it believes may contain objects, while disregarding background regions without objects.
Subsequently, the second stage is responsible for two functions: classifying and regressing bounding boxes for region proposal. One of the most renowned two-stage architectures is R-CNN [7] (Region-based Convolution Neural Network), intro- duced by Ross Girshick and colleagues in 2014. In the first stage, an "exhaustive search" algorithm is employed to generate 2000 candidate regions per image. These regions are then forwarded to the second stage, where a Con- volutional Neural Network (CNN) processes them to produce a fixed-length feature vector.
This feature vector is subsequently input into a Support Vector Machine (SVM) for region classification and prediction of bounding box co- ordinates [27]. The successor to R-CNN, Fast R-CNN [6], brought significant advancements in model training speed by eliminating the need for multiple forward pipelines for each region proposal.