Hand detection, segmentation and tracking from egocentric vision by Van-Tien Pham Submitted to the School of Information Technology and Communication in partial fulfillment of the requirements for the degree of Master of Science in Information System and Communication at the HANOI UNIVERSITY OF SCIENCE AND TECHNOLOGY October 2020 © Hanoi University of Science and Technology 2020. All rights reserved. School of Information Technology and Communication October 10, 2020 Certified by. Thi-Thanh-Hai Tran Associate Professor Thesis Supervisor Accepted by.
Chairman Chairman, Department Committee on Graduate Theses 17061132203221000000 Hand detection, segmentation and tracking from egocentric vision by Van-Tien Pham Submitted to the School of Information Technology and Communication on October 10, 2020, in partial fulfillment of the requirements for the degree of Master of Science in Information System and Communication Abstract Multiple object tracking is the process of assigning unique and consistent identities to objects throughout a video sequence. A popular approach to multiple object tracking is to use a method called tracking by detection. Tracking by detection is a two-stage procedure: an object detection or segmentation algorithm first detects objects in a given frame, these detected objects are then associated with already tracked objects in a second step by a tracking algorithm. Egocentric vision is an emerging field of computer vision that is characterized by the acquisition of images and video from the first-person perspective.
In egocentric view, the two human hands are essential in the execution of actions and characterizing their movements and trajectories are the principal cues to define and recognize actions. One of the main concerns of this thesis is to develop an automatic tracking by de- tection algorithm that extracts hands positions and identities in consequence frames from egocentric surveillance video. The proposed framework consists of state-of-the- art detectors from RCNN and YOLO family models combined with the SORT or DeepSORT for object tracking task. The thesis aims to explore how the stand-alone performance of the object detection algorithm correlates with overall performance of a tracking-by-detection system.
Finally, the thesis investigates how the use of visual descriptors of DeepSORT in the tracking stage of a tracking-by-detection system ef- fects performance. Results presented in this thesis suggest that the capacity of the object detection al- gorithm is highly indicative of the overall performance of the tracking-by detection system. Further, this thesis also shows how the use of visual descriptors in the track- ing stage can reduce the number of identity switches and thereby increase performance of the whole system. This thesis also presents a new egocentric hand tracking dataset Micand32 for future researches.
Thesis Supervisor: Thi-Thanh-Hai Tran Title: Associate Professor Acknowledgments First of all, I might want to offer my special thanks to my supervisor, Assoc. Tran Thi Thanh Hai. I’d really appreciate everything she’ve guided me all through this thesis. I would like to thank my colleagues at Viettel High Technology Industries Corpo- ration for supporting me in technical issues.
Also, I might want to express gratitude toward Assoc. Vu Hai and alumni at MICA Institute, Hanoi University of Science and Technology for giving me significant suggestions. Deep inside my heart, I wish to show my gratefulness to my family for always inspiring and trusting me in every of my steps.1 Overview of object recognition and tracking from video .2 Context and scope of the thesis .2 Background project and motivation .1 Video object recognition and tracking challenges .2 Hand gestures recognition related works .4 Problem formulation and assumptions. 23 2 Methodology and Datasets 25 2.1 Tracking by detection approach .2 Object detection and segmentation algorithms .1 RCNN model family .2 YOLO model family .3 Object tracking algorithms .4 Egocentric vision datasets .1 GTEA family datatsets .1 Proposed framework: tracking by detection .1 Training detection and segmentation models .2 Training deep appearance descriptor for DeepSORT .1 Object detection evaluation metrics .2 Object tracking evaluation metrics .1 Egocentric hand detection and segmentation result .2 Egocentric hand tracking result .1 Object detection: tradeoff between accuracy and speed .2 The superiority of DeepSORT over SORT .3 Impact of detection method over tracking result .4 Complexity of 4 types of patients’s actions .1 Short-term tracking results on Micand32S .2 Long-term tracking results on Micand32E.
83 7 THIS PAGE INTENTIONALLY LEFT BLANK 8 List of Figures 2-1 Schematic of the R-CNN pipeline [1]. 28 2-2 The architecture of Fast RCNN [2]. 29 2-3 An illustration of Faster RCNN model [3]. 31 2-4 The Mask RCNN framework for instance segmentation [4].
33 2-5 The network architecture of YOLO [5]. 34 2-6 Upshots from GTEA family datasets. 40 2-7 Hand masks after post-processing EGTEA Gaze+. 41 2-8 Visualizations of EgoHand dataset [6].
43 2-9 Randomly selected actions 5 6 7 8, from left to right respectively. Left: labels statistical visualization. Right: labels correlogram. 45 2-11 Visualization of groundtruth tracklets of patient’s hands practicing with cylinders.
Frames extracted from GH010358_8_8000_8547, or- dered from left to right, up to down: 1, 31, 61, 91, 121, 151, 181, 211, 241, 271, 301, 311, 361, 391, 421, 451. 47 2-12 The workflow of EHTA. 48 3-1 Overview of the proposed framework: D2D. The x-axis represents the time flow of 4 stages.
The y-axis shows the degree of abstraction levels of stages. 50 3-2 Workflow of the training stage. 51 3-3 Pictorial of data augmentations in training batch. 54 3-4 FasterRCNN_R_50_FPN_3x losses visualization during training time.
54 3-5 Images from the self-generated egocentric hand re-identification dataset. Images in the same row have the same identity. 55 3-6 Left: training configuration of DeepSORT’s appearance descriptor. Right: curve of total loss and top1-error during training loop.
56 9 3-7 Workflow of inference stage. 57 4-1 Overall MOTA and IDF1 metric on Micand32E. 70 4-2 Illustration of hand occluded by an obstacle. Pay attention to the patient’s right hand.
Frames extracted from GH010373_5_1284_2724 using FasterRCNN+SORT, ordered from left to right, up to down: (200, 210, 220, 230, 240, 250), (256, 257, 258, 259, 260, 261), (267, 270, 273, 276, 279, 282). 73 4-3 Motion blur phenomenon due to hand’s rapid movement. Frames extracted from GH010354_5_17718_19366 using Yolov3+SORT, or- dered from left to right: 119, 125, 130, 131, 137 and 143. 73 4-4 Shape changing illustration.
Frames extracted from GH010358_6_10208_11900 using MaskRCNN+DeepSORT, ordered from left to right: 1582, 1592, 1602, 1612, 1622 and 1630. 74 4-5 The unclear "hand" definition illustration. Pay attention to the pa- tient’s left hand. Frames extracted from GH010373_5_1284_2724 using FasterRCNN+SORT, ordered from left to right, up to down: (1123, 1173, 1223, 1273, 1323, 1406), (1407, 1408, 1409, 1410, 1411, 1412).
74 5-1 Schematic illustration of an online annotation pipeline. 77 B-1 Y3S detail results on Micand32E. 84 B-2 Y4S detail results on Micand32E. 84 B-3 FS detail results on Micand32E.
85 B-4 MS detail results on Micand32E. 85 B-5 GS detail results on Micand32E. 86 B-6 Y3DS detail results on Micand32E. 86 B-7 Y4DS detail results on Micand32E.
87 B-8 FDS detail results on Micand32E. 87 B-9 MDS detail results on Micand32E. 88 10 B-10 RDS detail results on Micand32E. 88 B-11 GDS detail results on Micand32E.
89 B-12 Y3S detail results on Micand32E. 89 B-13 Y4S detail results on Micand32E. 89 B-14 FS detail results on Micand32E. 89 B-15 MS detail results on Micand32E.
90 B-16 GS detail results on Micand32E. 90 B-17 Y3DS detail results on Micand32E. 90 B-18 Y4DS detail results on Micand32E. 90 B-19 FDS detail results on Micand32E.
90 B-20 MDS detail results on Micand32E. 90 B-21 RDS detail results on Micand32E. 91 B-22 GDS detail results on Micand32E. 91 B-23 Illustration of YOLO’s data format.
91 B-24 Illustration of YOLO’s training process. 92 11 THIS PAGE INTENTIONALLY LEFT BLANK 12 List of Tables 1.1 Some definitions for calculating FastRCNN losses.2 Some definitions for calculating FasterRCNN losses.3 Overview of the CNN architecture [7]. The final batch and L2 normal- ization projects onto the unit hyper-sphere.4 Detailed enumeration of the Micand32 dataset.1 Data format for both the inference result and ground-truth annotation of tracking.1 Object detection and segmentation Average Precision following the COCO standard.2 Object detection and segmentation Average Recall following the COCO standard.3 Object detection and segmentation average training and inference time, speed and machine requirements.4 Notation of the 11 approaches mentioned and used in the experiments.5 Short-term tracking overall result on Micand32S following the MOT16 evaluation protocol.6 Long-term tracking overrall results on Micand32E following the MOT16 protocol.7 GDS detail results on Micand32E.8 RDS detail results on Micand32E.1 FasterRCNN_R_50_FPN3x training configuration file. Detail infor- mation field is explained at the Detectron2’s application program- ming interface (API) documentation.
The main difference of Faster- RCNN and MaskRCNN in term of configuration is MaskRCNN’s op- tion MASK_ON value.2 YOLOv4x training configuration file. Detail information field is ex- plained at Ultralytic API.3 Detectron2’s custom dataset format.1 Overview of object recognition and tracking from video 1.1 Object recognition Object recognition aims at detecting the presence of an object in image and giving it a label that is the category to which it belongs. Objects of interest can be face, vehicle, hand, people, tree, tumors depending on applications such as face id, in- telligent traffic system and autonomous vehicles, human machine interaction, health care, security or bio diversity, etc. Segmentation is fundamental that go further than object detection and classification by give a label to a pixel, not a bounding box.
There are two types of segmentation: semantic segmentation and instance segmenta- tion. Semantic segmentation classifies all pixels of an image into meaningful object categories. These categories are "semantically interpretable" and correspond to the classes in the real world. Semantic segmentation gives an unique label to two objects of the same category.
This is called dense prediction because it can predict the mean- ing of each pixel. Instance segmentation otherwise gives every pixel belonging to an object instance a label. Since the past decade, object detection and recognition as well as segmentation has achieved impressive performance thanks to the significant advances of AI and deep learning. Deep learning can learn patterns in visual input in order to predict not only the categories of objects in image but also the ones of pixels.
Deep learning architectures used for object detection / recognition / segmentation are convolutional neural networks (CNN) or specific CNN frameworks such as AlexNet, VGG, Inception and ResNet.2 Object tracking Tracking objects in a video stream involves association of moving objects in consecu- tive video frames. In order to track an object, the target object requires to be firstly detected manually (detection free trackers) or automatically by detection algorithms (detection based tracker). Then tracking algorithm will associate detected objects to existing tracks by putting constraints on distance function between movement and appearance of the object with its previous instances. Besides, giving a complete tra- jectory of a moving object during time is very important for video analysis.
Tracking is more helpful to solve some common challenges (e. as lighting changes, motion blur, zoom ratio changes, occlusion when the target is partially or completely hidden by another object in the video for a period of time, poor image quality) from that simple object detection often suffers from. Almost proposed trackers until now based on Siamese network or Correlation Fil- ter (CF), combined with effective appearance models (CNN, HOG). In challenges on object tracking task, most of the highest performance obtained with CF trackers.
Their performances is better than Siam tracker. In term of computational time, Siam tracker’s performance is better than CF. Depending on the context and application, tracking could be single object tracking (SOT) and multiple object tracking (MOT). MOT is more challenging than SOT because ID switching is difficult to avoid, espe- cially in crowded videos, the nature and number of objects in each frame are unknown.
Therefore, MOT algorithms strongly rely on detection algorithms. Unfortunately, de- tection algorithms itself are not perfect. A popular object tracking method is to use a method called tracking by detection. It first apply object detection algorithms to detect objects in current frame.