MINISTRY OF EDUCATION AND TRAINING HANOI UNIVERSITY OF SCIENCE AND TECHNOLOGY TT ĐNqNV“1L Tuan Dung LE NIL DNQIIL ONOQILL II IMPROVING MULTI-VIEW ITUMAN ACTION RECOGNITION WITII SPATIAL-TEMPORAL POOLING AND VIEW SIIFTING TECIINIQUES MASTER OF SCIENCE THESIS IN INFORMATION SYSTEM 8102-2102 Hanoi 2018 MINISTRY OF EDUCATION AND TRAINING HANOT UNIVERSITY OF SCIENCE AND TECHNOLOGY Tuan Dung LE IMPROVING MULTI-VIEW HUMAN ACTION RECOGNITION WITH SPATIAL-TEMPORAL POOLING AND VIEW SHIFTING TECHNIQUES Speciality: Information System MASTER OF SCIENCE THESIS IN INFORMATION SYSTEM SUPERVISOR 1. Thi Oanh NGUYEN Hanoi 2015 Master aavieni: Tuan Bung LB - CBCI7016 Page 2 ACKNOWLEDGEMENT First of all, I sincerely thank the teachers in the School of Information and Communication Technology as well as all the teachers at the Hanoi University of Technology has taught me the knowledge and valuable experience during the pasl 5 years 1 would like to thank the two supervisors, Dr. Nguyen ‘Thi Oanh - lecturer in Tnformation Systems and = Commmumication, Tustitule of Tnformation and Communication Technology, Hanoi University of Technology and Dr. Tran Thi Thanh Hai, MICA Research Tnslilute has guided me Lo complets this master thesis.
T have learned a lot. from them, nol only the knowledge of the field of computer vision but alsa working and studying skills such as writing papers, preparing slides and presenting to the crowd Finally, T would like (o send ay thanks to my lamily, friends and people who have always supported me in the process of studying and researching this thesis Hanoi, March 2018 Master student ‘Tuan Dung LE Master steedend: Tuan Bung LB - CBCI7016 Page 3 Fipurc 3. 9 IIlustration of view shifting on N-UCLA dataset.đT Master aavieni: Tuan Bung LB - CBCI7016 Page 7 LIST OF ABBREVIATIONS AND DEFINITIONS OF TERMS Index | Abbreviation Full name 1 MHI ‘Motion History Image 2 MEI Motion Energy Image 3 LMEL Localized Motion Energy Image 4 SIP Spatio-lemporal Interest Point 5 SSM Self-Similarities Matrix 6 HOG Histogram of Oriented Gradient 7 HOF Hislogram of Optical Flow 8 TNMAS INRIA Xmas Acquisition Sequences 9 BoW Bag-ol-Words 10 ROIs Region of Interest Master steedend: Tuan Bung LB - CBCI7016 Page 9 LIST OF FIGURES Figure 1. | a) human body in frame, b) binary silhouttes, c) 3D Human Pose (visual hull), d} motion history volume, e) Motion Context, f) Gaussian blob human body model, g) cylindrical/cHipsoid human body model |1|.
2 Construct HOG-HOF descriptive vector bassd ơn SSM matrix[6]. 3 a) Original video of walking action with viewpoints 0° and 45°, their volumes and silhouettes, b) epipolar geometry in case of extracted actor body silhouettes, o) epipolar geometry in case of dynamic scene with dynamic actor and static background without extracting silhơuettes[9]. 4 MHI (middle row) and MEI (lagt rơø) templats [ISI. 5 Lustration of spatio-temporal interest point detected in a people clapping’s video [16].
6 Three ways lo combine multiple 2D views information in the BoW model ũM. | Proposed framework 24 Figure2. 2 Dividing space domain based on bounding box and eentroid 26 Figure 2. 3 lllustration of T-BoW modeL.
4 Illustration of view shifting in testing phase. t Tlustralion of 12 action classes im the WVU Multi-view actions dalasel - - 31 Figure 3. 2 Cameras setup for capturing WVU dataset - 31 Figure 3. 3 Ilustration of 10 action classes in the N-UCLA Multi-view Actions 3D -- TA.
4 Cameras1 setup for capturing | N-UCLA dataset. 5 IHustration of confusion natrix. 6 Confusion matrix: a) Basic BoW mods! with codebook D3 accuracy 70,83%; b) $-BoW model with 4 spatial parts codebook D3, accuracy 82,41%. 7 Confusion matrices: a) S-BoW model with 6 spatial parts, codebook D3, accuracy 78,24%; b) S-BoW model with 6 spatial parts and view shifting, codebook D3, accuracy 96,6790.
vcsessssssenssnssisseeeseesieeteetsoeetentn sn BB Figure 3. 8 Confusion matrices: a) Basic BOW model, codebook D3, accuracy 59,57%; b) S-BoW mofel with6 spatial parts, codebook 13, accuracy 63,40%.41 Master steedend: Tuan Bung LB - CBCI7016 Page 6 LIST OF TABLES ‘Table 3. 1 Accuracy (%) of basic BoW model on WVL datasel.36 Table 3 2 Accuracy (%) of T-BoW model on WVU dataset 36 Table 3. 3 Accuracy (%) of S-BoW model on WVU dalasel 38 ‘Table 3.
4 Accuracy (%) of 8-BoW model with (w) and without (w/o) view shifting, technique on WVU dataset. 5 Comparison wHhi olters metbiods on WVU DataseL. 6 Accuracy (%) of basic model on N-UCLA dataset.7 Accuracy (%) of T-BoW model on N-UCLA dataset 40 Table 3. 8 Accuracy (%) of the combination of S-BoW model and view shifting on N-UCLA dataset.
seer AL Table 3. 9 Accuracy (%) of 5-BoW model with (w) and without (w/o) view shifting technique on N-UCLA dataset. eeesieeesssieeeiseenicaeriensieane AZ, Master steedend: Tuan Bung LB - CBCI7016 Page 8 TABLE OF CONTENT ACKNOWLEDGEMENT. 3 ‘TABLE OF CONTEN' 4 LIST OF FIGURULS.
6 LIST OF TAILIES. kheierireooiỂ LIST OF ABBREVIATIONS AND DEFINITIONS OF TERMS. HUMAN ACTION RECOGNITION APPROACHES .2 Baseline method: combination of multiple 21D views in the Rag-ol-Words model - - - 20 CHAPTER 2.2 Combination of spatial/lemporal mfortalion and Bag-of-Words model 25 2.1 Combination of spatial information and Bag-ol-Words model (S-BoW).2 Combination of temporal information and Bag-of-Words model (T-BoW) 26 2.3 View shilting technique - - 27 CHAPTER 3, EXPERIMENTS.2 Setup XE rrrrerierariierrrrrererearo.1 Westem Virginia University Multi-view Action Recognition Dataset (WVU).2 Northwestem-UCLA Multiview Action 3D (N-UCLA). cseesscesseseessessteesitesinesinstesomvstestinssies eet BS 3.
- - - - 40 CONCLUSION & FUTURE WORK - - 43 REFERENCES. - 44 Master aavieni: Tuan Bung LB - CBCI7016 Page 4 LIST OF ABBREVIATIONS AND DEFINITIONS OF TERMS Index | Abbreviation Full name 1 MHI ‘Motion History Image 2 MEI Motion Energy Image 3 LMEL Localized Motion Energy Image 4 SIP Spatio-lemporal Interest Point 5 SSM Self-Similarities Matrix 6 HOG Histogram of Oriented Gradient 7 HOF Hislogram of Optical Flow 8 TNMAS INRIA Xmas Acquisition Sequences 9 BoW Bag-ol-Words 10 ROIs Region of Interest Master steedend: Tuan Bung LB - CBCI7016 Page 9 LIST OF TABLES ‘Table 3. 1 Accuracy (%) of basic BoW model on WVL datasel.36 Table 3 2 Accuracy (%) of T-BoW model on WVU dataset 36 Table 3. 3 Accuracy (%) of S-BoW model on WVU dalasel 38 ‘Table 3.
4 Accuracy (%) of 8-BoW model with (w) and without (w/o) view shifting, technique on WVU dataset. 5 Comparison wHhi olters metbiods on WVU DataseL. 6 Accuracy (%) of basic model on N-UCLA dataset.7 Accuracy (%) of T-BoW model on N-UCLA dataset 40 Table 3. 8 Accuracy (%) of the combination of S-BoW model and view shifting on N-UCLA dataset.
seer AL Table 3. 9 Accuracy (%) of 5-BoW model with (w) and without (w/o) view shifting technique on N-UCLA dataset. eeesieeesssieeeiseenicaeriensieane AZ, Master steedend: Tuan Bung LB - CBCI7016 Page 8 Fipurc 3. 9 IIlustration of view shifting on N-UCLA dataset.đT Master aavieni: Tuan Bung LB - CBCI7016 Page 7 LIST OF FIGURES Figure 1.
| a) human body in frame, b) binary silhouttes, c) 3D Human Pose (visual hull), d} motion history volume, e) Motion Context, f) Gaussian blob human body model, g) cylindrical/cHipsoid human body model |1|. 2 Construct HOG-HOF descriptive vector bassd ơn SSM matrix[6]. 3 a) Original video of walking action with viewpoints 0° and 45°, their volumes and silhouettes, b) epipolar geometry in case of extracted actor body silhouettes, o) epipolar geometry in case of dynamic scene with dynamic actor and static background without extracting silhơuettes[9]. 4 MHI (middle row) and MEI (lagt rơø) templats [ISI.
5 Lustration of spatio-temporal interest point detected in a people clapping’s video [16]. 6 Three ways lo combine multiple 2D views information in the BoW model ũM. | Proposed framework 24 Figure2. 2 Dividing space domain based on bounding box and eentroid 26 Figure 2.
3 lllustration of T-BoW modeL. 4 Illustration of view shifting in testing phase. t Tlustralion of 12 action classes im the WVU Multi-view actions dalasel - - 31 Figure 3. 2 Cameras setup for capturing WVU dataset - 31 Figure 3.
3 Ilustration of 10 action classes in the N-UCLA Multi-view Actions 3D -- TA. 4 Cameras1 setup for capturing | N-UCLA dataset. 5 IHustration of confusion natrix. 6 Confusion matrix: a) Basic BoW mods! with codebook D3 accuracy 70,83%; b) $-BoW model with 4 spatial parts codebook D3, accuracy 82,41%.
7 Confusion matrices: a) S-BoW model with 6 spatial parts, codebook D3, accuracy 78,24%; b) S-BoW model with 6 spatial parts and view shifting, codebook D3, accuracy 96,6790. vcsessssssenssnssisseeeseesieeteetsoeetentn sn BB Figure 3. 8 Confusion matrices: a) Basic BOW model, codebook D3, accuracy 59,57%; b) S-BoW mofel with6 spatial parts, codebook 13, accuracy 63,40%.41 Master steedend: Tuan Bung LB - CBCI7016 Page 6 Fipurc 3. 9 IIlustration of view shifting on N-UCLA dataset.đT Master aavieni: Tuan Bung LB - CBCI7016 Page 7 LIST OF FIGURES Figure 1.
| a) human body in frame, b) binary silhouttes, c) 3D Human Pose (visual hull), d} motion history volume, e) Motion Context, f) Gaussian blob human body model, g) cylindrical/cHipsoid human body model |1|. 2 Construct HOG-HOF descriptive vector bassd ơn SSM matrix[6]. 3 a) Original video of walking action with viewpoints 0° and 45°, their volumes and silhouettes, b) epipolar geometry in case of extracted actor body silhouettes, o) epipolar geometry in case of dynamic scene with dynamic actor and static background without extracting silhơuettes[9]. 4 MHI (middle row) and MEI (lagt rơø) templats [ISI.
5 Lustration of spatio-temporal interest point detected in a people clapping’s video [16]. 6 Three ways lo combine multiple 2D views information in the BoW model ũM. | Proposed framework 24 Figure2. 2 Dividing space domain based on bounding box and eentroid 26 Figure 2.
3 lllustration of T-BoW modeL. 4 Illustration of view shifting in testing phase. t Tlustralion of 12 action classes im the WVU Multi-view actions dalasel - - 31 Figure 3. 2 Cameras setup for capturing WVU dataset - 31 Figure 3.
3 Ilustration of 10 action classes in the N-UCLA Multi-view Actions 3D -- TA. 4 Cameras1 setup for capturing | N-UCLA dataset. 5 IHustration of confusion natrix. 6 Confusion matrix: a) Basic BoW mods! with codebook D3 accuracy 70,83%; b) $-BoW model with 4 spatial parts codebook D3, accuracy 82,41%.
7 Confusion matrices: a) S-BoW model with 6 spatial parts, codebook D3, accuracy 78,24%; b) S-BoW model with 6 spatial parts and view shifting, codebook D3, accuracy 96,6790. vcsessssssenssnssisseeeseesieeteetsoeetentn sn BB Figure 3. 8 Confusion matrices: a) Basic BOW model, codebook D3, accuracy 59,57%; b) S-BoW mofel with6 spatial parts, codebook 13, accuracy 63,40%.41 Master steedend: Tuan Bung LB - CBCI7016 Page 6 APPENDIX 1. Master steedend: Tuan Bung LB - CBCI7016 Page 5 INTRODUCTION In the growing social scene from the 3.0 era (automation of information technology and electronic production) to the new 4.0 (a new convergence of technologies such as the Internet ‘Things - Intemet, collaboration robots, 31) printing and cloud computing, and the emergence of new business models), automatically collecting and processing information by the compuler is very necessary.
This leads to highor demands on the interaction between humans and machines both in precision and speed. Thus, the problems of object recognition, motion recognition, speech yTecogmition. are now allracling a lot of interest of scientists and comparies around the world. Nowadays, video data is easily generated by devices such as digital cameras, laptops, mobile phones, and video-sharing websites.
Human action yeeogtition in the video, contributing to the automtaled exploitation of the resources of this rich data source. Applications related to human action recognition problems such as: Security and traditional monitoring systems include networks of cameras and are monitored by humans. With the increase in the number of cameras as well as these systems being deployed in multiple locations, the supervisor's efficiency and accuracy issues are required ta cover the entire system. The task of computer vision is to find a solution that can replace or assist the supervisor.
Aulomatic recognition of abmormalilics from surveillance systems is a matter that attracts a lot of research.