Luận án tiến sĩ nghiên cứu các kỹ thuật học sâu trong biểu diễn và nhận dạng hoạt động của người từ dữ liệu khung xương

Luận án tiến sĩ nghiên cứu các kỹ thuật học sâu trong biểu diễn và nhận dạng hoạt động của người từ dữ liệu khung xương, mang lại ứng dụng thực tiễn cao.

Chuyên ngành

Kỹ thuật máy tính

Người đăng

Ẩn danh

Thể loại

Luận án tiến sĩ

2022

130
4
0

Phí lưu trữ

35 Point

Mục lục chi tiết

DECLARATION OF AUTHORSHIP

ACKNOWLEDGEMENT

ABSTRACT

1. CHƯƠNG 1: AN OVERVIEW ON ACTION RECOGNITION

1.1. Data modalities for action recognition

1.2. Skeleton data collection

1.3. Data collection from motion capture systems

1.4. Data collection from RGB+D sensors

1.5. Data collection from pose estimation

1.6. Skeleton-based action recognition methods

1.6.1. Handcraft-based methods

1.6.2. Joint-based action recognition

1.6.3. Body part-based action recognition

1.7. Deep learning-based methods

1.7.1. Convolutional Neural Networks

1.7.2. Recurrent Neural Networks

1.8. Research on action recognition in Vietnam

1.9. Conclusion of the chapter

2. CHƯƠNG 2: JOINT SUBSET SELECTION FOR SKELETON-BASED HUMAN ACTION RECOGNITION

2.1. Preset Joint Subset Selection

2.2. Spatial-Temporal Representation

2.3. Dynamic Time Warping

2.4. Fourier Temporal Pyramid

2.5. Automatic Joint Subset Selection

2.6. Joint weight assignment

2.7. Most informative joint selection

2.8. Human action recognition based on MIJ joints

2.8.1. Preset Joint Subset Selection

2.8.2. Automatic Joint Subset Selection

2.9. Conclusion of the chapter

3. CHƯƠNG 3: FEATURE FUSION FOR THE GRAPH CONVOLUTIONAL NETWORK

3.1. Related work on Graph Convolutional Networks

3.2. Conclusion of the chapter

4. CHƯƠNG 4: THE PROPOSED LIGHTWEIGHT GRAPH CONVOLUTIONAL NETWORK

4.1. Related work on Lightweight Graph Convolutional Networks

4.2. Conclusion of the chapter

5. CHƯƠNG 5: CONCLUSION AND FUTURE WORKS

ABBREVIATIONS

SYMBOLS

LIST OF TABLES

LIST OF FIGURES

Tóm tắt

I. Tổng quan về nghiên cứu kỹ thuật học sâu trong nhận dạng hoạt động

Nghiên cứu về nhận dạng hoạt động của con người từ dữ liệu khung xương đang thu hút sự chú ý của cộng đồng khoa học. Kỹ thuật học sâu, đặc biệt là mạng nơ-ron, đã chứng minh hiệu quả trong việc phân tích và nhận diện các hành động phức tạp. Dữ liệu khung xương cung cấp thông tin chính xác về vị trí và chuyển động của các khớp, giúp cải thiện độ chính xác trong nhận dạng hoạt động.

1.1. Tầm quan trọng của nhận dạng hoạt động trong cuộc sống

Nhận dạng hoạt động có nhiều ứng dụng thực tiễn như giám sát an ninh, tương tác giữa người và máy, và theo dõi người cao tuổi. Việc áp dụng machine learning trong lĩnh vực này giúp nâng cao hiệu quả và độ chính xác.

1.2. Các thách thức trong nhận dạng hoạt động từ dữ liệu khung xương

Mặc dù có nhiều lợi ích, nhưng việc nhận dạng hoạt động từ dữ liệu khung xương cũng gặp phải nhiều thách thức như sai số trong ước lượng tư thế, nhiễu trong dữ liệu khung xương, và sự không đầy đủ do che khuất.

II. Phương pháp học sâu trong nhận dạng hoạt động từ dữ liệu khung xương

Các phương pháp học sâu như mạng nơ-ron tích chập (CNN) và mạng nơ-ron hồi tiếp (RNN) đã được áp dụng để cải thiện khả năng nhận dạng hoạt động. Những phương pháp này cho phép khai thác cấu trúc đồ thị của khung xương, từ đó nâng cao độ chính xác trong việc phân tích hành động.

2.1. Mạng nơ ron tích chập trong nhận dạng hoạt động

Mạng nơ-ron tích chập giúp nhận diện các đặc trưng không gian trong dữ liệu khung xương, từ đó cải thiện khả năng phân loại hành động. Các nghiên cứu đã chỉ ra rằng CNN có thể đạt được độ chính xác cao trong các bài toán nhận dạng phức tạp.

2.2. Mạng nơ ron hồi tiếp và ứng dụng của nó

Mạng nơ-ron hồi tiếp, đặc biệt là LSTM, cho phép xử lý thông tin theo chuỗi thời gian, giúp nhận diện các hành động liên tiếp trong một khoảng thời gian nhất định. Điều này rất quan trọng trong việc phân tích hành động của con người.

III. Kỹ thuật lựa chọn tập hợp khớp cho nhận dạng hoạt động

Kỹ thuật lựa chọn tập hợp khớp là một trong những phương pháp quan trọng để cải thiện hiệu suất nhận dạng hoạt động. Việc lựa chọn các khớp có vai trò quan trọng trong hành động giúp giảm thiểu nhiễu và tăng độ chính xác của mô hình.

3.1. Lựa chọn khớp cố định và tự động

Có hai phương pháp chính để lựa chọn khớp: lựa chọn khớp cố định và lựa chọn khớp tự động. Phương pháp tự động cho phép xác định các khớp quan trọng dựa trên dữ liệu thực tế, từ đó tối ưu hóa quá trình nhận dạng.

3.2. Đánh giá hiệu suất của các phương pháp lựa chọn khớp

Các nghiên cứu đã chỉ ra rằng việc lựa chọn khớp phù hợp có thể cải thiện đáng kể hiệu suất nhận dạng hoạt động, đặc biệt là trong các tập dữ liệu có nhiễu.

IV. Ứng dụng thực tiễn của nhận dạng hoạt động từ dữ liệu khung xương

Nhận dạng hoạt động từ dữ liệu khung xương có nhiều ứng dụng trong các lĩnh vực như giám sát an ninh, tương tác người-máy, và thực tế ảo. Các ứng dụng này không chỉ giúp nâng cao trải nghiệm người dùng mà còn cải thiện hiệu quả trong các hệ thống giám sát.

4.1. Giám sát an ninh và phát hiện sự kiện bất thường

Hệ thống giám sát sử dụng nhận dạng hoạt động có thể phát hiện các hành động bất thường, từ đó cảnh báo kịp thời cho người quản lý. Điều này giúp nâng cao an ninh trong các khu vực công cộng.

4.2. Tương tác giữa người và máy

Nhận dạng hoạt động cũng được áp dụng trong các hệ thống tương tác giữa người và máy, giúp cải thiện khả năng giao tiếp và tương tác tự nhiên hơn giữa con người và thiết bị.

V. Kết luận và tương lai của nghiên cứu nhận dạng hoạt động

Nghiên cứu về nhận dạng hoạt động từ dữ liệu khung xương đang phát triển mạnh mẽ với nhiều ứng dụng tiềm năng. Tương lai của lĩnh vực này hứa hẹn sẽ mang lại nhiều cải tiến trong công nghệ nhận dạng và ứng dụng thực tiễn.

5.1. Xu hướng phát triển trong nghiên cứu

Các xu hướng nghiên cứu hiện tại đang tập trung vào việc cải thiện độ chính xác và giảm thiểu độ phức tạp của mô hình. Việc phát triển các mô hình nhẹ hơn sẽ giúp ứng dụng dễ dàng hơn trên các thiết bị di động.

5.2. Tầm nhìn tương lai cho nhận dạng hoạt động

Tương lai của nhận dạng hoạt động sẽ tiếp tục mở rộng với sự phát triển của công nghệ computer visionmachine learning, hứa hẹn mang lại những giải pháp mới cho các bài toán phức tạp trong cuộc sống hàng ngày.

18/07/2025
Luận án tiến sĩ nghiên cứu các kỹ thuật học sâu trong biểu diễn và nhận dạng hoạt động của người từ dữ liệu khung xương

Trích đoạn nội dung tài liệu

MINISTRY OF EDUCATION AND TRAINING HANOI UNIVERSITY OF SCIENCE AND TECHNOLOGY PHAM DINH TAN A STUDY ON DEEP LEARNING TECHNIQUES FOR HUMAN ACTION REPRESENTATION AND RECOGNITION WITH SKELETON DATA DOCTORAL DISSERTATION IN COMPUTER ENGINEERING Hanoi−2022 LUAN VAN CHAT LUONG download : add luanvanchat@agmail.com MINISTRY OF EDUCATION AND TRAINING HANOI UNIVERSITY OF SCIENCE AND TECHNOLOGY PHAM DINH TAN A STUDY ON DEEP LEARNING TECHNIQUES FOR HUMAN ACTION REPRESENTATION AND RECOGNITION WITH SKELETON DATA Major: Computer Engineering Code: 9480106 DOCTORAL DISSERTATION IN COMPUTER ENGINEERING SUPERVISORS: 1. Le Thi Lan Hanoi−2022 LUAN VAN CHAT LUONG download : add luanvanchat@agmail.com DECLARATION OF AUTHORSHIP I, Pham Dinh Tan, declare that the dissertation titled "A study on deep learning techniques for human action representation and recognition with skeleton data" has been entirely composed by myself. I assure some points as follows:  This work was done wholly or mainly while in candidature for a Ph. research degree at Hanoi University of Science and Technology.

 The work has not been submitted for any other degree or qualifications at Hanoi University of Science and Technology or any other institution.  Appropriate acknowledgment has been given within this dissertation, where ref- erence has been made to the published work of others.  The dissertation submitted is my own, except where work in the collaboration has been included. The collaborative contributions have been indicated.

Hanoi, March 08, 2022 Ph. Le Thi Lan i LUAN VAN CHAT LUONG download : add luanvanchat@agmail.com ACKNOWLEDGEMENT This dissertation is composed during my Ph. at the Computer Vision Department, MICA Institute, Hanoi University of Science and Technology. I am grateful to all people who contribute in different ways to my Ph.

First, I would like to express sincere thanks to my supervisors Assoc. Vu Hai and Assoc. Le Thi Lan for their guidance and support. I would like to thank all MICA members for their help during my Ph.

My sincere thank to Dr. Nguyen Viet Son, Assoc. Dao Trung Kien, and Assoc. Tran Thi Thanh Hai for giving me a lot of support and valuable advice.

Many thanks to Dr. Nguyen Thuy Binh, Nguyen Hong Quan, Hoang Van Nam, Nguyen Tien Nam, and Pham Quang Tien for their support. I would like to thank colleagues at Hanoi University of Mining and Geology for all support during my Ph. Special thanks to my family for understanding my hours glued to the computer screen.

Hanoi, March 08, 2022 Ph. Student ii LUAN VAN CHAT LUONG download : add luanvanchat@agmail.com ABSTRACT Human action recognition (HAR) from color and depth sensors (RGB-D), especially derived information such as skeleton data, is receiving the research community’s at- tention due to its wide range of applications. HAR has many practical applications such as abnormal event detection from camera surveillance, gaming, human-machine interaction, elderly monitoring, and virtual/augmented reality. In addition to the ad- vantages in fast computation, low storage, and immutability with human appearance, skeleton data have shortcomings.

The shortcomings include pose estimation errors, skeleton noise in complex actions, and incompleteness due to occlusion. Moreover, action recognition remains challenging due to the diversity of human actions, intra- class variations, and inter-class similarities. The dissertation focuses on methods to improve the performances of action recognition using the skeleton data. The proposed methods are evaluated using public skeleton datasets collected by RGB-D sensors.

Es- pecially, they consist of MSR-Action3D/MICA-Action3D - datasets with high-quality skeleton data, CMDFALL - a challenging dataset with noise in skeleton data, and NTU RGB+D - a worldwide benchmark among the large-scale datasets. Therefore, these datasets cover different dataset scales as well as the quality of skeleton data. To overcome the limitations of the skeleton data, the dissertation presents tech- niques in different approaches. First, as joints have different levels of engagement in each action, techniques for selecting joints that play an important role in human actions are proposed, including both Preset joint subset selection and automatic joint subset selection.

Two frameworks are evaluated to show the performance of using a subset of joints for action representation. The first framework employs Dynamic Time Warping (DTW) and Fourier Temporal Pyramid (FTP), while the second one applies Covari- ance Descriptors extracted on both joint position and joint velocity. Experimental results show that joint subsect selection helps improve action recognition performance on datasets with noise in skeleton data. However, HAR based on hand-designed features could not exploit the inherent graph structure of the human skeleton.

Recent Graph Convolution Networks (GCNs) are studied to handle these issues. Among GCN models, Attention-enhanced Adaptive Convolutional Network (AAGCN) is used as the baseline model. AAGCN achieves state-of-the-art performance on large-scale datasets such as NTU-RGBD and Kinetics. However, AAGCN employs only joint information.

Therefore, a Feature Fusion (FF) module is proposed in this dissertation. The new model is named FF-AAGCN. The performance of FF-AAGCN is evaluated on the large-scale dataset NTU-RGBD and CMDFALL. The evaluation results show that the proposed method is robust to noise iii LUAN VAN CHAT LUONG download : add luanvanchat@agmail.com and invariant to the skeleton translation.

Particularly, FF-AAGCN achieves remark- able results on challenging datasets. Finally, as the computing capacity of edge devices is limited, a lightweight deep learning model is expected for application deployment. A lightweight GCN architecture is proposed to show that the complexity of GCN archi- tecture can still be reduced depending on the dataset’s characteristics. The proposed lightweight model is suitable for application development on edge devices.

Hanoi, March 08, 2022 Ph. Student iv LUAN VAN CHAT LUONG download : add luanvanchat@agmail.com CONTENTS DECLARATION OF AUTHORSHIP. x LIST OF TABLES. xiii LIST OF FIGURES.

An overview on action recognition. Data modalities for action recognition. Skeleton data collection. Data collection from motion capture systems.

Data collection from RGB+D sensors. Data collection from pose estimation. Skeleton-based action recognition methods. Handcraft-based methods.

Joint-based action recognition. Body part-based action recognition. 24 v LUAN VAN CHAT LUONG download : add luanvanchat@agmail. Deep learning-based methods.

Convolutional Neural Networks. Recurrent Neural Networks. Research on action recognition in Vietnam. Conclusion of the chapter.

JOINT SUBSET SELECTION FOR SKELETON-BASED HUMAN ACTION RECOGNITION. Preset Joint Subset Selection. Spatial-Temporal Representation. Dynamic Time Warping.

Fourier Temporal Pyramid. Automatic Joint Subset Selection. Joint weight assignment. Most informative joint selection.

Human action recognition based on MIJ joints. Preset Joint Subset Selection. Automatic Joint Subset Selection. Conclusion of the chapter.

FEATURE FUSION FOR THE GRAPH CONVOLUTIONAL NETWORK. Related work on Graph Convolutional Networks. Conclusion of the chapter. THE PROPOSED LIGHTWEIGHT GRAPH CONVOLU- TIONAL NETWORK.

Related work on Lightweight Graph Convolutional Networks. 86 vi LUAN VAN CHAT LUONG download : add luanvanchat@agmail. Conclusion of the chapter. 97 CONCLUSION AND FUTURE WORKS.

102 vii LUAN VAN CHAT LUONG download : add luanvanchat@agmail.com ABBREVIATIONS No. Abbreviation Meaning 1 2D Two-Dimensional 2 3D Three-Dimensional 3 AAGCN Attention-enhanced Adaptive Graph Convolutional Network 4 AMIJ Adaptive number of Most Informative Joints 5 AGCN Adaptive Graph Convolutional Network 6 AS Action Set 7 AS-GCN Actional-Structural Graph Convolutional Network 8 BN Batch Normalization 9 BPL Body Part Location 10 CAM Channel Attention Module 11 CCTV Close-Circuit Television 12 CNN Convolutional Neural Network 13 CovMIJ Covariance Descriptor on Most Informative Joints 14 CPU Central Processing Unit 15 CS Cross-Subject 16 CV Cross-View 17 DFT Discrete Fourier Transform 18 DTW Dynamic Time Warping 19 FC Fully Connected 20 FF Feature Fusion 21 FLOP Floating Point OPeration 22 FMIJ Fixed number of Most Informative Joints 23 fps f rames per second 24 FTP Fourier Temporal Pyramid 25 GCN Graph Convolutional Network 26 GCNN Graph-based Convolutional Neural Network 27 GPU Graphical Processing Unit 28 GRU Gated Recurrent Unit 29 HAR Human Action Recognition 30 HCI Human-Computer Interaction viii LUAN VAN CHAT LUONG download : add luanvanchat@agmail.com 31 HMM Hidden Markov Model 32 MRF Markov Random Field 33 JA Joint Angle 34 JP Joint Position 35 JSS Joint Subset Selection 36 LARP Lie Algebra Relative Pair 37 LSTM Long-Short Term Memory 38 MIJ Most Informative Joint 39 Mocap Motion capture System 40 MRF Markov Random Field 41 MTLN Multi-Task Learning Network 42 OL OverLapping 43 RA-GCN Richly Activated Graph Convolutional Network 44 ReLU Rectified Linear Unit 45 ResNet Residual Neural Network 46 RJP Relative Joint Position 47 RNN Recurrent Neural Network 48 RVM Relevance Vector Machine 49 SAM Spatial Attention Module 50 SDK Software Development Kit 51 SE Special Euclidean group 52 SO Special Orthogonal group 53 ST-GCN Spatial-Temporal Graph Convolutional Network 54 STC Spatial-Temporal-Channel Attention Module 55 SVM Support Vector Machine 56 t-SNE t-Distributed Stochastic Neighbor Embedding 57 TAM Temporal Attention Module 58 TCD Temporal Covariance Descriptor 59 TCN Temporal Convolutional Network 60 UAV Unmanned Aerial Vehicle 61 VFDT Very Fast Decision Trees ix LUAN VAN CHAT LUONG download : add luanvanchat@agmail.com SYMBOLS No. Symbol Meaning 1 A The adjacency matrix of the graph 2 Ãk Normalized adjacency matrix 3 C The number of action classes 4 COV (Sp ) Temporal covariance descriptor of joint positions 5 COV (Sv ) Temporal covariance descriptor of joint velocities 6 COVsample (Sp ) The sample covariance descriptor of joint positions 7 COVsample (Sv ) The sample covariance descriptor of joint velocities 8 D The cost function in Dynamic Time Warping 9 D The degree matrix of the graph 10 ES The set of intra-skeleton edges 11 ET The set of inter-frame edges 12 E The set of graph edges 13 F The feature vector 14 fin Input feature 15 fout Output feature 16 G The graph 17 L The number of layers in the temporal hierarchy 18 Ks Spatial kernel size 19 M The number of most informative joints 20 N The number of joints in the skeleton 21 N bc The number of samples in the cth action class 22 pi (t) Joint position of the ith joint at the tth frame 23 ReLU The Rectified Linear Unit activation function 24 Sp (t) The coordinate vector of informative joints at the tth frame 25 Sv (t) The velocity vector of all informative joints at the tth frame 26 Sp Mean of Sp (t) 27 Softmax The Softmax activation function 28 T The length of a skeleton sequence 29 Tmax The maximum length of skeleton sequences 30 tanh The hyperbolic tangent activation function x LUAN VAN CHAT LUONG download : add luanvanchat@agmail.com 31 wij The weight of the ith joint for the j th sample in an action class (1 ≤ j ≤ N bc ) 32 wi The weight of the ith joint 33 Wk Matrix of trainable weights 34 vti The node corresponding to the ith joint at time frame tth 35 V (t) Joint velocity at time frame tth 36 V The set of vertexes of the graph 37 xi (t) Coordinate of the ith joint along the x-axis at the tth frame 38 yi (t) Coordinate of the ith joint along the y-axis at the tth frame 39 zi (t) Coordinate of the ith joint along the z-axis at the tth frame xi LUAN VAN CHAT LUONG download : add luanvanchat@agmail.com LIST OF TABLES 1.1 Datasets with different data modalities: Skeleton (S), Depth (D), Accel- eration (Ac).2 List of actions in MSR-Action3D.3 List of actions in CMDFALL.4 List of actions in NTU RGB+D.1 Accuracy (%) comparison on MSR-Action3D.2 Ablation study on MSR-Action3D by accuracy (%).3 Computational time (ms) of Preset JSS on MSR-Action3D.4 Performance evaluation for Preset JSS on CMDFALL.5 The accuracy (%) obtained by the proposed method with different num- bers of layers and features for MSR-Action3D and CMDFALL.6 Accuracy (%) comparison of MIJ methods with existing methods on MSR-Action3D.7 Performance evaluation for FMIJ/AMIJ on CMDFALL.8 Computational time (ms) of FMIJ/AMIJ on MSR-Action3D.9 Comparison between Preset JSS using Covariance Descriptors and AMIJ.2 Input and output data shapes of FF-AAGCN on MICA-Action3D.3 Software package version information on Ubuntu 18.4 Ablation study on CMDFALL.5 Performance of FF-AAGCN on CMDFALL using velocity with different frame offsets. Performance scores are in percentage.

Nội dung được bảo vệ bản quyền — Tải xuống đầy đủ