MINISTRY OF EDUCATION AND TRAINING HANOI UNIVERSITY OF SCIENCE AND TECHNOLOGY PHAM DINH TAN A STUDY ON DEEP LEARNING TECHNIQUES FOR HUMAN ACTION REPRESENTATION AND RECOGNITION WITH SKELETON DATA DOCTORAL DISSERTATION IN COMPUTER ENGINEERING Hanoi−2022 MINISTRY OF EDUCATION AND TRAINING HANOI UNIVERSITY OF SCIENCE AND TECHNOLOGY PHAM DINH TAN A STUDY ON DEEP LEARNING TECHNIQUES FOR HUMAN ACTION REPRESENTATION AND RECOGNITION WITH SKELETON DATA Major: Computer Engineering Code: 9480106 DOCTORAL DISSERTATION IN COMPUTER ENGINEERING SUPERVISORS: 1. Le Thi Lan Hanoi−2022 DECLARATION OF AUTHORSHIP I, Pham Dinh Tan, declare that the dissertation titled "A study on deep learning techniques for human action representation and recognition with skeleton data" has been entirely composed by myself. I assure some points as follows: This work was done wholly or mainly while in candidature for a Ph. research degree at Hanoi University of Science and Technology.
The work has not been submitted for any other degree or qualifications at Hanoi University of Science and Technology or any other institution. Appropriate acknowledgment has been given within this dissertation, where ref- erence has been made to the published work of others. The dissertation submitted is my own, except where work in the collaboration has been included. The collaborative contributions have been indicated.
Hanoi, March 08, 2022 Ph. Le Thi Lan i ACKNOWLEDGEMENT This dissertation is composed during my Ph. at the Computer Vision Department, MICA Institute, Hanoi University of Science and Technology. I am grateful to all people who contribute in different ways to my Ph.
First, I would like to express sincere thanks to my supervisors Assoc. Vu Hai and Assoc. Le Thi Lan for their guidance and support. I would like to thank all MICA members for their help during my Ph.
My sincere thank to Dr. Nguyen Viet Son, Assoc. Dao Trung Kien, and Assoc. Tran Thi Thanh Hai for giving me a lot of support and valuable advice.
Many thanks to Dr. Nguyen Thuy Binh, Nguyen Hong Quan, Hoang Van Nam, Nguyen Tien Nam, and Pham Quang Tien for their support. I would like to thank colleagues at Hanoi University of Mining and Geology for all support during my Ph. Special thanks to my family for understanding my hours glued to the computer screen.
Hanoi, March 08, 2022 Ph. Student ii ABSTRACT Human action recognition (HAR) from color and depth sensors (RGB-D), especially derived information such as skeleton data, is receiving the research community’s at- tention due to its wide range of applications. HAR has many practical applications such as abnormal event detection from camera surveillance, gaming, human-machine interaction, elderly monitoring, and virtual/augmented reality. In addition to the ad- vantages in fast computation, low storage, and immutability with human appearance, skeleton data have shortcomings.
The shortcomings include pose estimation errors, skeleton noise in complex actions, and incompleteness due to occlusion. Moreover, action recognition remains challenging due to the diversity of human actions, intra- class variations, and inter-class similarities. The dissertation focuses on methods to improve the performances of action recognition using the skeleton data. The proposed methods are evaluated using public skeleton datasets collected by RGB-D sensors.
Es- pecially, they consist of MSR-Action3D/MICA-Action3D - datasets with high-quality skeleton data, CMDFALL - a challenging dataset with noise in skeleton data, and NTU RGB+D - a worldwide benchmark among the large-scale datasets. Therefore, these datasets cover different dataset scales as well as the quality of skeleton data. To overcome the limitations of the skeleton data, the dissertation presents tech- niques in different approaches. First, as joints have different levels of engagement in each action, techniques for selecting joints that play an important role in human actions are proposed, including both Preset joint subset selection and automatic joint subset selection.
Two frameworks are evaluated to show the performance of using a subset of joints for action representation. The first framework employs Dynamic Time Warping (DTW) and Fourier Temporal Pyramid (FTP), while the second one applies Covari- ance Descriptors extracted on both joint position and joint velocity. Experimental results show that joint subsect selection helps improve action recognition performance on datasets with noise in skeleton data. However, HAR based on hand-designed features could not exploit the inherent graph structure of the human skeleton.
Recent Graph Convolution Networks (GCNs) are studied to handle these issues. Among GCN models, Attention-enhanced Adaptive Convolutional Network (AAGCN) is used as the baseline model. AAGCN achieves state-of-the-art performance on large-scale datasets such as NTU-RGBD and Kinetics. However, AAGCN employs only joint information.
Therefore, a Feature Fusion (FF) module is proposed in this dissertation. The new model is named FF-AAGCN. The performance of FF-AAGCN is evaluated on the large-scale dataset NTU-RGBD and CMDFALL. The evaluation results show that the proposed method is robust to noise iii and invariant to the skeleton translation.
Particularly, FF-AAGCN achieves remark- able results on challenging datasets. Finally, as the computing capacity of edge devices is limited, a lightweight deep learning model is expected for application deployment. A lightweight GCN architecture is proposed to show that the complexity of GCN archi- tecture can still be reduced depending on the dataset’s characteristics. The proposed lightweight model is suitable for application development on edge devices.
Hanoi, March 08, 2022 Ph. Student iv CONTENTS DECLARATION OF AUTHORSHIP. x LIST OF TABLES. xiii LIST OF FIGURES.
An overview on action recognition. Data modalities for action recognition. Skeleton data collection. Data collection from motion capture systems.
Data collection from RGB+D sensors. Data collection from pose estimation. Skeleton-based action recognition methods. Handcraft-based methods.
Joint-based action recognition. Body part-based action recognition. Deep learning-based methods. Convolutional Neural Networks.
Recurrent Neural Networks. Research on action recognition in Vietnam. Conclusion of the chapter. JOINT SUBSET SELECTION FOR SKELETON-BASED HUMAN ACTION RECOGNITION.
Preset Joint Subset Selection. Spatial-Temporal Representation. Dynamic Time Warping. Fourier Temporal Pyramid.
Automatic Joint Subset Selection. Joint weight assignment. Most informative joint selection. Human action recognition based on MIJ joints.
Preset Joint Subset Selection. Automatic Joint Subset Selection. Conclusion of the chapter. FEATURE FUSION FOR THE GRAPH CONVOLUTIONAL NETWORK.
Related work on Graph Convolutional Networks. Conclusion of the chapter. THE PROPOSED LIGHTWEIGHT GRAPH CONVOLU- TIONAL NETWORK. Related work on Lightweight Graph Convolutional Networks.
Conclusion of the chapter. 97 CONCLUSION AND FUTURE WORKS. 102 vii ABBREVIATIONS No. Abbreviation Meaning 1 2D Two-Dimensional 2 3D Three-Dimensional 3 AAGCN Attention-enhanced Adaptive Graph Convolutional Network 4 AMIJ Adaptive number of Most Informative Joints 5 AGCN Adaptive Graph Convolutional Network 6 AS Action Set 7 AS-GCN Actional-Structural Graph Convolutional Network 8 BN Batch Normalization 9 BPL Body Part Location 10 CAM Channel Attention Module 11 CCTV Close-Circuit Television 12 CNN Convolutional Neural Network 13 CovMIJ Covariance Descriptor on Most Informative Joints 14 CPU Central Processing Unit 15 CS Cross-Subject 16 CV Cross-View 17 DFT Discrete Fourier Transform 18 DTW Dynamic Time Warping 19 FC Fully Connected 20 FF Feature Fusion 21 FLOP Floating Point OPeration 22 FMIJ Fixed number of Most Informative Joints 23 fps f rames per second 24 FTP Fourier Temporal Pyramid 25 GCN Graph Convolutional Network 26 GCNN Graph-based Convolutional Neural Network 27 GPU Graphical Processing Unit 28 GRU Gated Recurrent Unit 29 HAR Human Action Recognition 30 HCI Human-Computer Interaction viii 31 HMM Hidden Markov Model 32 MRF Markov Random Field 33 JA Joint Angle 34 JP Joint Position 35 JSS Joint Subset Selection 36 LARP Lie Algebra Relative Pair 37 LSTM Long-Short Term Memory 38 MIJ Most Informative Joint 39 Mocap Motion capture System 40 MRF Markov Random Field 41 MTLN Multi-Task Learning Network 42 OL OverLapping 43 RA-GCN Richly Activated Graph Convolutional Network 44 ReLU Rectified Linear Unit 45 ResNet Residual Neural Network 46 RJP Relative Joint Position 47 RNN Recurrent Neural Network 48 RVM Relevance Vector Machine 49 SAM Spatial Attention Module 50 SDK Software Development Kit 51 SE Special Euclidean group 52 SO Special Orthogonal group 53 ST-GCN Spatial-Temporal Graph Convolutional Network 54 STC Spatial-Temporal-Channel Attention Module 55 SVM Support Vector Machine 56 t-SNE t-Distributed Stochastic Neighbor Embedding 57 TAM Temporal Attention Module 58 TCD Temporal Covariance Descriptor 59 TCN Temporal Convolutional Network 60 UAV Unmanned Aerial Vehicle 61 VFDT Very Fast Decision Trees ix SYMBOLS No.
Symbol Meaning 1 A The adjacency matrix of the graph 2 Ãk Normalized adjacency matrix 3 C The number of action classes 4 COV (Sp ) Temporal covariance descriptor of joint positions 5 COV (Sv ) Temporal covariance descriptor of joint velocities 6 COVsample (Sp ) The sample covariance descriptor of joint positions 7 COVsample (Sv ) The sample covariance descriptor of joint velocities 8 D The cost function in Dynamic Time Warping 9 D The degree matrix of the graph 10 ES The set of intra-skeleton edges 11 ET The set of inter-frame edges 12 E The set of graph edges 13 F The feature vector 14 fin Input feature 15 fout Output feature 16 G The graph 17 L The number of layers in the temporal hierarchy 18 Ks Spatial kernel size 19 M The number of most informative joints 20 N The number of joints in the skeleton 21 N bc The number of samples in the cth action class 22 pi (t) Joint position of the ith joint at the tth frame 23 ReLU The Rectified Linear Unit activation function 24 Sp (t) The coordinate vector of informative joints at the tth frame 25 Sv (t) The velocity vector of all informative joints at the tth frame 26 Sp Mean of Sp (t) 27 Softmax The Softmax activation function 28 T The length of a skeleton sequence 29 Tmax The maximum length of skeleton sequences 30 tanh The hyperbolic tangent activation function x 31 wij The weight of the ith joint for the j th sample in an action class (1 ≤ j ≤ N bc ) 32 wi The weight of the ith joint 33 Wk Matrix of trainable weights 34 vti The node corresponding to the ith joint at time frame tth 35 V (t) Joint velocity at time frame tth 36 V The set of vertexes of the graph 37 xi (t) Coordinate of the ith joint along the x-axis at the tth frame 38 yi (t) Coordinate of the ith joint along the y-axis at the tth frame 39 zi (t) Coordinate of the ith joint along the z-axis at the tth frame xi LIST OF TABLES 1.1 Datasets with different data modalities: Skeleton (S), Depth (D), Accel- eration (Ac).2 List of actions in MSR-Action3D.3 List of actions in CMDFALL.4 List of actions in NTU RGB+D.1 Accuracy (%) comparison on MSR-Action3D.2 Ablation study on MSR-Action3D by accuracy (%).3 Computational time (ms) of Preset JSS on MSR-Action3D.4 Performance evaluation for Preset JSS on CMDFALL.5 The accuracy (%) obtained by the proposed method with different num- bers of layers and features for MSR-Action3D and CMDFALL.6 Accuracy (%) comparison of MIJ methods with existing methods on MSR-Action3D.7 Performance evaluation for FMIJ/AMIJ on CMDFALL.8 Computational time (ms) of FMIJ/AMIJ on MSR-Action3D.9 Comparison between Preset JSS using Covariance Descriptors and AMIJ.2 Input and output data shapes of FF-AAGCN on MICA-Action3D.3 Software package version information on Ubuntu 18.4 Ablation study on CMDFALL.5 Performance of FF-AAGCN on CMDFALL using velocity with different frame offsets. Performance scores are in percentage.6 Performance evaluation on CMDFALL with Precision, Recall, and F1 scores (%).7 Comparison of computation time between FF-AAGCN and the baseline on CMDFALL. Testing time is calculated per sample.8 Ablation study by accuracy (%) on NTU RGB+D.9 Performance evaluation by accuracy (%) on NTU RGB+D.10 Comparison of training/testing time between FF-AAGCN and the base- line on NTU RGB+D with cross-subject (CS) benchmark. Training time is calculated in hours (h).
Testing time per sample is calculated in milli- seconds (ms).11 Comparison on training/testing time between FF-AAGCN and the base- line on NTU RGB+D with cross-view (CV) benchmark. Training time is calculated in hours (h).