MINISTRY OF EDUCATION AND TRAINING HANOI UNIVERSITY OF SCIENCE AND TECHNOLOGY NGUYEN THUY BINH PERSON RE-IDENTIFICATION IN A SURVEILLANCE CAMERA NETWORK DOCTORAL DISSERTATION OF ELECTRONICS ENGINEERING Hanoi−2020 luan an MINISTRY OF EDUCATION AND TRAINING HANOI UNIVERSITY OF SCIENCE AND TECHNOLOGY NGUYEN THUY BINH PERSON RE-IDENTIFICATION IN A SURVEILLANCE CAMERA NETWORK Major: Electronics Engineering Code: 9520203 DOCTORAL DISSERTATION OF ELECTRONICS ENGINEERING SUPERVISORS: 1. Pham Ngoc Nam 2. Le Thi Lan Hanoi−2020 luan an DECLARATION OF AUTHORSHIP I, Nguyen Thuy Binh, declare that the thesis titled "Person re-identification in a surveillance camera network" has been entirely composed by myself. I assure some points as follows: This work was done wholly or mainly while in candidature for a Ph.
research degree at Hanoi University of Science and Technology. The work has not be submitted for any other degree or qualifications at Hanoi University of Science and Technology or any other institutions. Appropriate acknowledge has been given within this thesis where reference has been made to the published work of others. The thesis submitted is my own, except where work in the collaboration has been included.
The collaborative contributions have been clearly indicated. Hanoi, 24/11/ 2020 PhD Student SUPERVISORS i luan an ACKNOWLEDGEMENT This dissertation was written during my doctoral course at School of Electronics and Telecommunications (SET) and International Research Institute of Multimedia, Infor- mation, Communication and Applications (MICA), Hanoi University of Science and Technology (HUST). I am so grateful for all people who always support and encourage me for completing this study. First, I would like to express my sincere gratitude to my advisors Assoc.
Pham Ngoc Nam and Assoc. Le Thi Lan for their effective guidance, their patience, continuous support and encouragement, and their immense knowledge. I would like to express my gratitude to Dr. Vo Le Cuong and Dr.
Ha thi Thu Lan for their help. I would like to thank to all member of School of Electronics and Telecom- munications, International Research Institute of Multimedia, Information, Communi- cations and Applications (MICA), Hanoi University of Science and Technology (HUST) as well as all of my colleagues in Faculty of Electrical-Electronic Engineering, University of Transport and Communications (UTC). They have always helped me on research process and given helpful advises for me to overcome my own difficulties. Moreover, the attention at scientific conferences has always been a great experience for me to receive many the useful comments.
During my PhD course, I have received many supports from the Management Board of School of Electronics and Telecommunications, MICA Institute, and Faculty of Electrical-Electronic Engineering. My sincere thank to Assoc. Nguyen Huu Thanh, Dr. Nguyen Viet Son and Assoc.
Nguyen Thanh Hai who gave me a lot of support and help. Without their precious support, it has been impossible to conduct this research. Thanks to my employer, University of Transport and Communications (UTC) for all necessary support and encouragement during my PhD journey. I am also grateful to Vietnam’s Program 911, HUST and UTC projects for their generous financial support.
Special thanks to my family and relatives, particularly, my beloved husband and our children, for their never-ending support and sacrifice. Student ii luan an CONTENTS DECLARATION OF AUTHORSHIP. vi LIST OF TABLES. x LIST OF FIGURES.
Person ReID classifications. Single-shot versus Multi-shot. Closed-set versus Open-set person ReID. Supervised and unsupervised person ReID.
Datasets and evaluation metrics. Hand-designed features. Deep-learned features. Metric learning and person matching.
Fusion schemes for person ReID. Representative frame selection. Fully automated person ReID systems. Research on person ReID in Vietnam.
MULTI-SHOT PERSON RE-ID THROUGH REPRESEN- TATIVE FRAMES SELECTION AND TEMPORAL FEATURE POOLING 36 2. Representative image selection. 37 iii luan an 2. Image-level feature extraction.
Temporal feature pooling. Evaluation of representative frame extraction and temporal feature pooling schemes. Quantitative evaluation of the trade-off between the accuracy and compu- tational time. Comparison with state-of-the-art methods.
Conclusions and Future work. PERSON RE-ID PERFORMANCE IMPROVEMENT BASED ON FUSION SCHEMES. Fusion schemes for the first setting of person ReID. Image-to-images person ReID.
Images-to-images person ReID. Obtained results on the first setting. Fusion schemes for the second setting of person ReID. The proposed method.
Obtained results on the second setting. QUANTITATIVE EVALUATION OF AN END-TO-END PERSON REID PIPELINE. An end-to-end person ReID pipeline. GOG descriptor re-implementation.
Comparison the performance of two implementations. Analyze the effect of GOG parameters. Evaluation performance of an end-to-end person ReID pipeline. The effect of human detection and segmentation on person ReID in single- shot scenario.
102 iv luan an 4. The effect of human detection and segmentation on person ReID in multi- shot scenario. Conclusions and Future work. 113 v luan an ABBREVIATIONS No.
Abbreviation Meaning 1 ACF Aggregate Channel Features 2 AIT Austrian Institute of Technology 3 AMOC Accumulative Motion Context 4 BOW Bag of Words 5 CAR Learning Compact Appearance Representation 6 CIE The International Commission on Illumination 7 CFFM Comprehensive Feature Fusion Mechanism 8 CMC Cummulative Matching Characteristic 9 CNN Convolutional Neural Network 10 CPM Convolutional Pose Machines 11 CVPDL Cross-view Projective Dictionary Learning 12 CVPR Conference on Computer Vision and Pattern Recognition 13 DDLM Discriminative Dictionary Learning Method 14 DDN Deep Decompositional Network 15 DeepSORT Deep learning Simple Online and Realtime Tracking 16 DFGP Deep Feature Guided Pooling 17 DGM Dynamic Graph Matching 18 DPM Deformable Part-Based Model 19 ECCV European Conference on Computer Vision 20 FAST 3D Fast Adaptive Spatio-Temporal 3D 21 FEP Flow Energy Profile 22 FNN Feature Fusion Network 23 FPNN Filter Pairing Neural Network 24 GOG Gaussian of Gaussian 25 GRU Gated Recurrent Unit 26 HOG Histogram of Oriented Gradients 27 HUST Hanoi University of Science and Technology 28 IBP Indian Buffet Process 29 ICCV International Conference on Computer Vision 30 ICIP International Conference on Image Processing vi luan an 31 IDE ID-Discriminative Embedding 32 iLIDS-VID Imagery Library for Intelligent Detection Systems 33 ILSVRC ImageNet Large Scale Visual Recognition Competition 34 ISR TIterative Spare Ranking 35 KCF Kernelized Correlation Filter 36 KDES Kenel DEScriptor 37 KISSME Keep It Simple and Straightforward MEtric 38 kNN k-Nearest Neighbour 39 KXQDA Kernel Cross-view Quadratic Discriminative Analysis 40 LADF Locally-Adaptive Decision Functions 41 LBP Local Binary Pattern 42 LDA LinearDiscriminantAnalysis 43 LDFV Local Descriptor and coded by Feature Vector 44 LMNN Large Margin Nearest Neighbor 45 LMNN-R Large Margin Nearest Neighbor with Rejection 46 LOMO LOcal Maximal Occurrence 47 LSTM Long-Short Term Memory 48 LSTMC Long Short-Term Memory network with a Coupled gate 49 mAP mean Average Precision 50 MAPR Multimedia Analysis and Pattern Recognition 51 Mask R-CNN Mask Region with CNN 52 MCT Multi -Camera Tracking 53 MCCNN Multi-Channel CNN 54 MCML Maximally Collapsing Metric Learning 55 MGCAM Mask-Guided Contrastive Attention Model 56 ML Machine Learning 57 MLAPG Metric Learning by Accelerated Proximal Gradient 58 MLR Metric Learning to Rank 59 MOT Multiple Object Tracking 60 MSCR Maximal Stable Color Region 61 MSVF Maximally Stable Video Frame 62 MTMCT Multi-Target Multi-Camera Tracking 63 Person ReID Person Re -Identification 64 Pedparsing Pedestrian Parsing 65 PPN Pose Prediction Network vii luan an 66 PRW Person Re-identification in the Wild 67 QDA Quadratic Discriminative Analysis 68 RAiD Re-Identification Across indoor-outdoor Dataset 69 RAP Richly Annotated Pedestrian 70 ResNet Residual Neural Network 71 RHSP Recurrent High-Structured Patches 72 RKHS Reproducing Kernel Hilbert Space 73 RNN Recurrent Neural Network 74 ROIs Region of Interests 75 SDALF Symmetry Driven Accumulation of Local Feature 76 SCNCD Salient Color Names based Color Descriptor 77 SCNN Siamese Convolutional Neural Network 78 SIFT Scale-Invariant Feature Transform 79 SILTP Scale Invariant Local Ternary Pattern 80 SPD Symmetric Positive Definite 81 SMP Stepwise Metric Promotion 82 SORT Simple Online and Realtime Tracking 83 SPIC Signal Processing: Image Communication 84 SVM Support Vector Machine 85 TAPR Temporally Aligned Pooling Representation 86 TAUDL Tracklet Association Unsupervised Deep Learning 87 TCSVT Transactions on Circuits and Systems for Video Technology 88 TII Transactions on Industrial Informatics 89 TPAMI Transactions on Pattern Analysis and Machine Intelligence 90 TPDL Top-push Distance Learning 91 Two-stream MR Two-stream Multirate Recurrent Neural Network 92 UIT University of Information Technology 93 UTAL Tracklet Association Unsupervised Deep Learning 94 VIPeR View-point Invariant Pedestrian Recognition 95 VNU-HCM Vietnam National University - Ho Chi Minh City 96 WH Weighted color Histogram 97 WHOS Weighted Histograms of Overlapping Stripes 98 WSC Weight-based Sparse Coding 99 XQDA Cross-view Quadratic Discriminative Analysis 100 YOLO You Only Look One viii luan an LIST OF TABLES 1.1 Benchmark datasets used in the thesis.1 The matching rates (%) when applying different pooling methods on different color spaces in case of using four key frames on PRID 2011 dataset. The two best results for each case are in bold.2 The matching rates (%) when applying different pooling methods on different color spaces in case of using frames within a walking cycle on PRID 2011 dataset. The two best results for each case are in bold.3 The matching rates (%) when applying different pooling methods on different color spaces in case of using all frames on PRID 2011 dataset. The two best results for each case are in bold.4 The matching rates (%) when applying different pooling methods on different color spaces in case of using four key frames on iLIDS-VID dataset.
The two best results for each case are in bold.5 The matching rates (%) when applying different pooling methods on different color spaces in case of using frames within a walking cycle on iLIDS-VID dataset. The two best results for each case are in bold.6 The matching rates (%) when applying different pooling methods on different color spaces in case of using all frames on iLIDS-VID dataset. The two best results for each case are in bold.7 Matching rates (%) in several important ranks when using four key frames, four random frames, and one random frame in PRID-2011 and iLIDS-VID datasets.8 Comparison of the three representative frame selection schemes in term of accuracy at rank-1, computational time, and memory requirement on PRID 2011 dataset.9 Comparison between the proposed method and existing works on PRID 2011 and iLIDS-VID datasets. Two best results are in bold.2 Matching rates (%) in case of images-to-images person ReID on the RAiD dataset.3 Comparison the best matching rates at rank-1 in image-to-images case and those of in images-to-images one.
80 ix luan an 3.4 Comparison of images-to-images and image-to-images schemes at rank- 1. (*) means the obtained results by applying the proposed strategies over 10 random trials in case A of CAVIAR4REID.5 Comparison between the proposed method and existing works on PRID 2011 and iLIDS-VID datasets.Two best results are in bold.1 Comparison of the proposed method with state of the art methods for PRID 2011 (the two best results are in bold). 107 x luan an LIST OF FIGURES 1 The ranked list of gallery person corresponding to the given query based on the similarities between the query and each of gallery ones. 2 2 An example for challenges caused by variations in a) illumination b) view-point.
3 3 A person has multiple images captured in different camera-views. 4 4 A fully-automatic person ReID system consisting of three main stages: human detection, tracking and re-identification.1 Some important milestones for person ReID problem [8]. Several ap- proaches related to this thesis are bounded by red blocks.2 An example for a) single-shot (image-based) and b)multi-shot person (video-based) ReID approaches.3 The differences between a) Closed-set and b) Open-set person ReID. In closed-set person ReID, an individual appears on at least two camera- views.
Inversely, in open-set person ReID, a pedestrian might appear on only one camera-view.4 Two popular settings for person ReID problem: a) The testing persons have appeared in the training set (represented by the same colors) b) Persons in the training and testing sets are absolutely different.5 Camera layout for PRID-2011 dataset [33].6 iLIDS-VID is captured by five non-overlapping cameras [36].