Nghiên cứu về kỹ thuật học sâu cho nhận diện hành động con người từ dữ liệu xương

Nghiên cứu về kỹ thuật học sâu trong việc đại diện và nhận diện hành động con người từ dữ liệu xương. Tìm hiểu ứng dụng và tiềm năng.

Chuyên ngành

Computer Engineering

Người đăng

Ẩn danh

Thể loại

doctoral dissertation

2022

130
3
0

Phí lưu trữ

35 Point

Mục lục chi tiết

DECLARATION OF AUTHORSHIP

ACKNOWLEDGEMENT

ABSTRACT

CONTENTS

1. CHƯƠNG 1: AN OVERVIEW ON ACTION RECOGNITION

1.1. Data modalities for action recognition

1.2. Skeleton data collection

1.3. Data collection from motion capture systems

1.4. Data collection from RGB+D sensors

1.5. Data collection from pose estimation

1.6. Skeleton-based action recognition methods

1.6.1. Handcraft-based methods

1.6.2. Joint-based action recognition

1.6.3. Body part-based action recognition

1.6.4. Deep learning-based methods

1.6.4.1. Convolutional Neural Networks
1.6.4.2. Recurrent Neural Networks

1.7. Research on action recognition in Vietnam

1.8. Conclusion of the chapter

2. CHƯƠNG 2: JOINT SUBSET SELECTION FOR SKELETON-BASED HUMAN ACTION RECOGNITION

2.1. Preset Joint Subset Selection

2.2. Spatial-Temporal Representation

2.3. Dynamic Time Warping

2.4. Fourier Temporal Pyramid

2.5. Automatic Joint Subset Selection

2.6. Joint weight assignment. Most informative joint selection

2.7. Human action recognition based on MIJ joints

2.7.1. Preset Joint Subset Selection

2.7.2. Automatic Joint Subset Selection

2.8. Conclusion of the chapter

3. CHƯƠNG 3: FEATURE FUSION FOR THE GRAPH CONVOLUTIONAL NETWORK

3.1. Related work on Graph Convolutional Networks

3.2. Conclusion of the chapter

4. CHƯƠNG 4: THE PROPOSED LIGHTWEIGHT GRAPH CONVOLUTIONAL NETWORK

4.1. Related work on Lightweight Graph Convolutional Networks

4.2. Conclusion of the chapter

CONCLUSION AND FUTURE WORKS

ABBREVIATIONS

SYMBOLS

LIST OF TABLES

LIST OF FIGURES

Tóm tắt

I. Tổng quan về nghiên cứu kỹ thuật học sâu trong nhận diện hành động con người

Nghiên cứu về nhận diện hành động con người (HAR) đang thu hút sự chú ý của cộng đồng nghiên cứu nhờ vào ứng dụng rộng rãi của nó. Các kỹ thuật học sâu, đặc biệt là mạng nơ-ron, đã được áp dụng để cải thiện độ chính xác trong việc nhận diện hành động từ dữ liệu xương. Dữ liệu xương cung cấp thông tin chi tiết về chuyển động của cơ thể, giúp cho việc phân tích hành động trở nên hiệu quả hơn.

1.1. Định nghĩa và tầm quan trọng của nhận diện hành động

Nhận diện hành động là quá trình xác định và phân loại các hành động của con người từ dữ liệu thu thập được. Điều này có ứng dụng trong nhiều lĩnh vực như giám sát an ninh, tương tác giữa người và máy, và thực tế ảo.

1.2. Các loại dữ liệu sử dụng trong nhận diện hành động

Các loại dữ liệu thường được sử dụng trong HAR bao gồm dữ liệu hình ảnh, dữ liệu chiều sâu và dữ liệu xương. Dữ liệu xương, với khả năng cung cấp thông tin về vị trí và chuyển động của các khớp, là một trong những lựa chọn phổ biến nhất.

II. Thách thức trong nhận diện hành động từ dữ liệu xương

Mặc dù dữ liệu xương mang lại nhiều lợi ích, nhưng vẫn tồn tại một số thách thức trong việc nhận diện hành động. Các vấn đề như sai số trong ước lượng tư thế, nhiễu trong dữ liệu xương và sự không đầy đủ do che khuất là những yếu tố cần được giải quyết.

2.1. Sai số trong ước lượng tư thế

Sai số trong ước lượng tư thế có thể dẫn đến việc nhận diện hành động không chính xác. Điều này thường xảy ra khi các khớp bị che khuất hoặc không được phát hiện đúng cách.

2.2. Nhiễu trong dữ liệu xương

Nhiễu trong dữ liệu xương có thể làm giảm độ chính xác của các mô hình học sâu. Việc xử lý và làm sạch dữ liệu là rất quan trọng để cải thiện hiệu suất nhận diện.

III. Phương pháp cải thiện nhận diện hành động từ dữ liệu xương

Để cải thiện hiệu suất nhận diện hành động, nhiều phương pháp đã được đề xuất. Các phương pháp này bao gồm việc lựa chọn tập hợp khớp quan trọng và áp dụng các mô hình học sâu như mạng nơ-ron tích chậpmạng nơ-ron đồ thị.

3.1. Lựa chọn tập hợp khớp quan trọng

Việc lựa chọn các khớp quan trọng giúp giảm thiểu nhiễu và cải thiện độ chính xác của mô hình. Các phương pháp như lựa chọn tự động và lựa chọn theo quy định đã được áp dụng.

3.2. Sử dụng mạng nơ ron đồ thị

Mạng nơ-ron đồ thị (GCN) cho phép khai thác cấu trúc đồ thị của dữ liệu xương, giúp cải thiện khả năng nhận diện hành động trong các tình huống phức tạp.

IV. Ứng dụng thực tiễn của nhận diện hành động từ dữ liệu xương

Nhận diện hành động từ dữ liệu xương có nhiều ứng dụng thực tiễn, từ giám sát an ninh đến tương tác giữa người và máy. Các nghiên cứu đã chỉ ra rằng HAR có thể cải thiện đáng kể trải nghiệm người dùng trong các ứng dụng thực tế ảo và trò chơi điện tử.

4.1. Giám sát an ninh

HAR có thể được sử dụng để phát hiện các hành động bất thường trong video giám sát, giúp nâng cao an ninh cho các khu vực công cộng.

4.2. Tương tác giữa người và máy

Các hệ thống tương tác giữa người và máy có thể sử dụng HAR để nhận diện các hành động của người dùng, từ đó cải thiện khả năng phản hồi và tương tác.

V. Kết luận và tương lai của nghiên cứu nhận diện hành động

Nghiên cứu về nhận diện hành động từ dữ liệu xương đang phát triển mạnh mẽ. Các kỹ thuật học sâu tiếp tục được cải thiện, hứa hẹn mang lại những kết quả tốt hơn trong tương lai. Việc phát triển các mô hình nhẹ và hiệu quả cho các thiết bị biên cũng là một xu hướng quan trọng.

5.1. Xu hướng phát triển mô hình nhẹ

Mô hình nhẹ giúp giảm thiểu yêu cầu về tài nguyên tính toán, cho phép triển khai trên các thiết bị di động và IoT.

5.2. Tương lai của nhận diện hành động

Với sự phát triển của công nghệ, nhận diện hành động sẽ ngày càng trở nên chính xác và nhanh chóng hơn, mở ra nhiều cơ hội ứng dụng mới trong cuộc sống hàng ngày.

23/07/2025
Luận văn thạc sĩ a study on deep learning techniques for human action representation and recognition with skeleton data

Trích đoạn nội dung tài liệu

MINISTRY OF EDUCATION AND TRAINING HANOI UNIVERSITY OF SCIENCE AND TECHNOLOGY PHAM DINH TAN A STUDY ON DEEP LEARNING TECHNIQUES FOR HUMAN ACTION REPRESENTATION AND RECOGNITION WITH SKELETON DATA DOCTORAL DISSERTATION IN COMPUTER ENGINEERING Hanoi−2022 MINISTRY OF EDUCATION AND TRAINING HANOI UNIVERSITY OF SCIENCE AND TECHNOLOGY PHAM DINH TAN A STUDY ON DEEP LEARNING TECHNIQUES FOR HUMAN ACTION REPRESENTATION AND RECOGNITION WITH SKELETON DATA Major: Computer Engineering Code: 9480106 DOCTORAL DISSERTATION IN COMPUTER ENGINEERING SUPERVISORS: 1. Le Thi Lan Hanoi−2022 DECLARATION OF AUTHORSHIP I, Pham Dinh Tan, declare that the dissertation titled "A study on deep learning techniques for human action representation and recognition with skeleton data" has been entirely composed by myself. I assure some points as follows:  This work was done wholly or mainly while in candidature for a Ph. research degree at Hanoi University of Science and Technology.

 The work has not been submitted for any other degree or qualifications at Hanoi University of Science and Technology or any other institution.  Appropriate acknowledgment has been given within this dissertation, where ref- erence has been made to the published work of others.  The dissertation submitted is my own, except where work in the collaboration has been included. The collaborative contributions have been indicated.

Hanoi, March 08, 2022 Ph. Le Thi Lan i ACKNOWLEDGEMENT This dissertation is composed during my Ph. at the Computer Vision Department, MICA Institute, Hanoi University of Science and Technology. I am grateful to all people who contribute in different ways to my Ph.

First, I would like to express sincere thanks to my supervisors Assoc. Vu Hai and Assoc. Le Thi Lan for their guidance and support. I would like to thank all MICA members for their help during my Ph.

My sincere thank to Dr. Nguyen Viet Son, Assoc. Dao Trung Kien, and Assoc. Tran Thi Thanh Hai for giving me a lot of support and valuable advice.

Many thanks to Dr. Nguyen Thuy Binh, Nguyen Hong Quan, Hoang Van Nam, Nguyen Tien Nam, and Pham Quang Tien for their support. I would like to thank colleagues at Hanoi University of Mining and Geology for all support during my Ph. Special thanks to my family for understanding my hours glued to the computer screen.

Hanoi, March 08, 2022 Ph. Student ii ABSTRACT Human action recognition (HAR) from color and depth sensors (RGB-D), especially derived information such as skeleton data, is receiving the research community’s at- tention due to its wide range of applications. HAR has many practical applications such as abnormal event detection from camera surveillance, gaming, human-machine interaction, elderly monitoring, and virtual/augmented reality. In addition to the ad- vantages in fast computation, low storage, and immutability with human appearance, skeleton data have shortcomings.

The shortcomings include pose estimation errors, skeleton noise in complex actions, and incompleteness due to occlusion. Moreover, action recognition remains challenging due to the diversity of human actions, intra- class variations, and inter-class similarities. The dissertation focuses on methods to improve the performances of action recognition using the skeleton data. The proposed methods are evaluated using public skeleton datasets collected by RGB-D sensors.

Es- pecially, they consist of MSR-Action3D/MICA-Action3D - datasets with high-quality skeleton data, CMDFALL - a challenging dataset with noise in skeleton data, and NTU RGB+D - a worldwide benchmark among the large-scale datasets. Therefore, these datasets cover different dataset scales as well as the quality of skeleton data. To overcome the limitations of the skeleton data, the dissertation presents tech- niques in different approaches. First, as joints have different levels of engagement in each action, techniques for selecting joints that play an important role in human actions are proposed, including both Preset joint subset selection and automatic joint subset selection.

Two frameworks are evaluated to show the performance of using a subset of joints for action representation. The first framework employs Dynamic Time Warping (DTW) and Fourier Temporal Pyramid (FTP), while the second one applies Covari- ance Descriptors extracted on both joint position and joint velocity. Experimental results show that joint subsect selection helps improve action recognition performance on datasets with noise in skeleton data. However, HAR based on hand-designed features could not exploit the inherent graph structure of the human skeleton.

Recent Graph Convolution Networks (GCNs) are studied to handle these issues. Among GCN models, Attention-enhanced Adaptive Convolutional Network (AAGCN) is used as the baseline model. AAGCN achieves state-of-the-art performance on large-scale datasets such as NTU-RGBD and Kinetics. However, AAGCN employs only joint information.

Therefore, a Feature Fusion (FF) module is proposed in this dissertation. The new model is named FF-AAGCN. The performance of FF-AAGCN is evaluated on the large-scale dataset NTU-RGBD and CMDFALL. The evaluation results show that the proposed method is robust to noise iii and invariant to the skeleton translation.

Particularly, FF-AAGCN achieves remark- able results on challenging datasets. Finally, as the computing capacity of edge devices is limited, a lightweight deep learning model is expected for application deployment. A lightweight GCN architecture is proposed to show that the complexity of GCN archi- tecture can still be reduced depending on the dataset’s characteristics. The proposed lightweight model is suitable for application development on edge devices.

Hanoi, March 08, 2022 Ph. Student iv CONTENTS DECLARATION OF AUTHORSHIP. x LIST OF TABLES. xiii LIST OF FIGURES.

An overview on action recognition. Data modalities for action recognition. Skeleton data collection. Data collection from motion capture systems.

Data collection from RGB+D sensors. Data collection from pose estimation. Skeleton-based action recognition methods. Handcraft-based methods.

Joint-based action recognition. Body part-based action recognition. Deep learning-based methods. Convolutional Neural Networks.

Recurrent Neural Networks. Research on action recognition in Vietnam. Conclusion of the chapter. JOINT SUBSET SELECTION FOR SKELETON-BASED HUMAN ACTION RECOGNITION.

Preset Joint Subset Selection. Spatial-Temporal Representation. Dynamic Time Warping. Fourier Temporal Pyramid.

Automatic Joint Subset Selection. Joint weight assignment. Most informative joint selection. Human action recognition based on MIJ joints.

Preset Joint Subset Selection. Automatic Joint Subset Selection. Conclusion of the chapter. FEATURE FUSION FOR THE GRAPH CONVOLUTIONAL NETWORK.

Related work on Graph Convolutional Networks. Conclusion of the chapter. THE PROPOSED LIGHTWEIGHT GRAPH CONVOLU- TIONAL NETWORK. Related work on Lightweight Graph Convolutional Networks.

Conclusion of the chapter. 97 CONCLUSION AND FUTURE WORKS. 102 vii ABBREVIATIONS No. Abbreviation Meaning 1 2D Two-Dimensional 2 3D Three-Dimensional 3 AAGCN Attention-enhanced Adaptive Graph Convolutional Network 4 AMIJ Adaptive number of Most Informative Joints 5 AGCN Adaptive Graph Convolutional Network 6 AS Action Set 7 AS-GCN Actional-Structural Graph Convolutional Network 8 BN Batch Normalization 9 BPL Body Part Location 10 CAM Channel Attention Module 11 CCTV Close-Circuit Television 12 CNN Convolutional Neural Network 13 CovMIJ Covariance Descriptor on Most Informative Joints 14 CPU Central Processing Unit 15 CS Cross-Subject 16 CV Cross-View 17 DFT Discrete Fourier Transform 18 DTW Dynamic Time Warping 19 FC Fully Connected 20 FF Feature Fusion 21 FLOP Floating Point OPeration 22 FMIJ Fixed number of Most Informative Joints 23 fps f rames per second 24 FTP Fourier Temporal Pyramid 25 GCN Graph Convolutional Network 26 GCNN Graph-based Convolutional Neural Network 27 GPU Graphical Processing Unit 28 GRU Gated Recurrent Unit 29 HAR Human Action Recognition 30 HCI Human-Computer Interaction viii 31 HMM Hidden Markov Model 32 MRF Markov Random Field 33 JA Joint Angle 34 JP Joint Position 35 JSS Joint Subset Selection 36 LARP Lie Algebra Relative Pair 37 LSTM Long-Short Term Memory 38 MIJ Most Informative Joint 39 Mocap Motion capture System 40 MRF Markov Random Field 41 MTLN Multi-Task Learning Network 42 OL OverLapping 43 RA-GCN Richly Activated Graph Convolutional Network 44 ReLU Rectified Linear Unit 45 ResNet Residual Neural Network 46 RJP Relative Joint Position 47 RNN Recurrent Neural Network 48 RVM Relevance Vector Machine 49 SAM Spatial Attention Module 50 SDK Software Development Kit 51 SE Special Euclidean group 52 SO Special Orthogonal group 53 ST-GCN Spatial-Temporal Graph Convolutional Network 54 STC Spatial-Temporal-Channel Attention Module 55 SVM Support Vector Machine 56 t-SNE t-Distributed Stochastic Neighbor Embedding 57 TAM Temporal Attention Module 58 TCD Temporal Covariance Descriptor 59 TCN Temporal Convolutional Network 60 UAV Unmanned Aerial Vehicle 61 VFDT Very Fast Decision Trees ix SYMBOLS No.

Symbol Meaning 1 A The adjacency matrix of the graph 2 Ãk Normalized adjacency matrix 3 C The number of action classes 4 COV (Sp ) Temporal covariance descriptor of joint positions 5 COV (Sv ) Temporal covariance descriptor of joint velocities 6 COVsample (Sp ) The sample covariance descriptor of joint positions 7 COVsample (Sv ) The sample covariance descriptor of joint velocities 8 D The cost function in Dynamic Time Warping 9 D The degree matrix of the graph 10 ES The set of intra-skeleton edges 11 ET The set of inter-frame edges 12 E The set of graph edges 13 F The feature vector 14 fin Input feature 15 fout Output feature 16 G The graph 17 L The number of layers in the temporal hierarchy 18 Ks Spatial kernel size 19 M The number of most informative joints 20 N The number of joints in the skeleton 21 N bc The number of samples in the cth action class 22 pi (t) Joint position of the ith joint at the tth frame 23 ReLU The Rectified Linear Unit activation function 24 Sp (t) The coordinate vector of informative joints at the tth frame 25 Sv (t) The velocity vector of all informative joints at the tth frame 26 Sp Mean of Sp (t) 27 Softmax The Softmax activation function 28 T The length of a skeleton sequence 29 Tmax The maximum length of skeleton sequences 30 tanh The hyperbolic tangent activation function x 31 wij The weight of the ith joint for the j th sample in an action class (1 ≤ j ≤ N bc ) 32 wi The weight of the ith joint 33 Wk Matrix of trainable weights 34 vti The node corresponding to the ith joint at time frame tth 35 V (t) Joint velocity at time frame tth 36 V The set of vertexes of the graph 37 xi (t) Coordinate of the ith joint along the x-axis at the tth frame 38 yi (t) Coordinate of the ith joint along the y-axis at the tth frame 39 zi (t) Coordinate of the ith joint along the z-axis at the tth frame xi LIST OF TABLES 1.1 Datasets with different data modalities: Skeleton (S), Depth (D), Accel- eration (Ac).2 List of actions in MSR-Action3D.3 List of actions in CMDFALL.4 List of actions in NTU RGB+D.1 Accuracy (%) comparison on MSR-Action3D.2 Ablation study on MSR-Action3D by accuracy (%).3 Computational time (ms) of Preset JSS on MSR-Action3D.4 Performance evaluation for Preset JSS on CMDFALL.5 The accuracy (%) obtained by the proposed method with different num- bers of layers and features for MSR-Action3D and CMDFALL.6 Accuracy (%) comparison of MIJ methods with existing methods on MSR-Action3D.7 Performance evaluation for FMIJ/AMIJ on CMDFALL.8 Computational time (ms) of FMIJ/AMIJ on MSR-Action3D.9 Comparison between Preset JSS using Covariance Descriptors and AMIJ.2 Input and output data shapes of FF-AAGCN on MICA-Action3D.3 Software package version information on Ubuntu 18.4 Ablation study on CMDFALL.5 Performance of FF-AAGCN on CMDFALL using velocity with different frame offsets. Performance scores are in percentage.6 Performance evaluation on CMDFALL with Precision, Recall, and F1 scores (%).7 Comparison of computation time between FF-AAGCN and the baseline on CMDFALL. Testing time is calculated per sample.8 Ablation study by accuracy (%) on NTU RGB+D.9 Performance evaluation by accuracy (%) on NTU RGB+D.10 Comparison of training/testing time between FF-AAGCN and the base- line on NTU RGB+D with cross-subject (CS) benchmark. Training time is calculated in hours (h).

Testing time per sample is calculated in milli- seconds (ms).11 Comparison on training/testing time between FF-AAGCN and the base- line on NTU RGB+D with cross-view (CV) benchmark. Training time is calculated in hours (h).

Nội dung được bảo vệ bản quyền — Tải xuống đầy đủ