VIETNAM NATIONAL UNIVERSITY — HO CHI MINH CAMPUS UNIVERSITY OF INFORMATION AND TECHNOLOGY FACULTY OF COMPUTER SCIENCE Object’s vibration frequency estimation using video magnification and applications PHAN THANH NHAN - 19521944 INTRUCTOR : DR NGUYEN VINH TIẾP DANH SÁCH HỘI ĐÒNG BẢO VỆ KHÓA LUẬN Hội đồng cham khóa luận tốt nghiệp, thành lập theo Quyết định số. của Hiệu trưởng Trường Đại học Công nghệ Thông tin. Lê Minh Hưng. Ths Nguyễn Thị Ngọc Diễm.
ĐẠI HỌC QUỐC GIA TP. HỒ CHÍ MINH CỘNG HÒA XÃ HỘI CHỦ NGHĨA VIỆT NAM TRƯỜNG ĐẠI HỌC Độc Lập - Tự Do - Hạnh Phúc CÔNG NGHỆ THÔNG TIN ĐĂNG KÝ ĐÈ TÀI KHÓA LUẬN TÓT NGHIỆP Tên đề tài: Ước lượng tần số chuyển động của vật thể dựa trên khuếch đại video và ứng dụng Tên đề tài tiếng Anh: Object’s vibration frequency estimation using video magnification and applications Ngôn ngữ thực hiện: Tiếng Anh Cán bộ hướng dẫn: TS. Nguyễn Vinh Tiệp Thời gian thực hiện:Từ ngày 5/09/2022 đến ngày 24/12/2022. Sinh viên thực hiện: Phan Thành Nhân - 19521944 Lớp:KHCL2019.vn Điện thoại:0918095450 Nội dung đề tài:(Mô tả chi tiết mục tiêu, phạm vi, đối tượng, phương pháp thực hiện, kết quả mong đợi của đề tài) Giới thiệu đề tài : “If you want to find the secrets of the universe, think in terms of energy, frequency and vibration.” Nikola Tesla Nhà bác hoc Nikola Tesla từng có một câu nói : “nếu bạn muốn tim thấy được bí mat của vũ trụ , hãy suy nghĩvề năng lượng , tần số và sự rung động “ Như chúng ta đều biết, vạn vật trong vũ trụ này đều được hình thành từ các hạt vật chất , và chúng đều sẽ xảy ra giao động khi gặp đúng tần số cộng hưởng của nó.
Từ những con ốc bị lỏng trong máy móc công nghiệp, cho đến các động mạch đập dưới da đều liên tục sinh ra các giao động có tần số ổn định. Tuy nhiên đó là những chuyển động mà mắt thường con người không thể nhìn thấy được, bởi sự giao động của chúng quá nhỏ. Tuy nhiên, các chuyển động này vẫn có thể được phần nào ghi lại bởi các máy ảnh điện tử và được khuếch đại bởi các thuật toán cũng như các mô hình học sâu. Từ đây có thể giúp chúng ta quan sát được các hiện tượng mà trước đó không thể quan sát được.
Đồng thời ước lượng được tần số giao động của vật thể và biến chúng thành những thông tin hữu ích và trực quan hơn cho con người có thể tham khảo. Đề tài này nhắm đến việc tạo ra một ứng dụng với input là một video có bao gồm một đối tượng đang giao động, output sẽ là một video đã khuếch đại các chuyển động nói trên và tần số giao động ước lượng được. Từ đó có thể sử dụng tần số ước lượng được cho các ứng dụng ví dụ như : do nhịp tim cũng như phục hồi lại những âm thanh từ đó. Với những nghiên cứu sơ bộ, em nhận thấy đã có rất nhiều hướng tiếp cận cho bài toán frequency estimation, tuy nhiên trong phạm vi bài toán này, em xin tìm hiểu dạng bài toán video magnification với hướng tiếp cận đến từ các nhà nghiên cứu thuộc đại học MIT với các bài toán nổi bật như Learning-based Video Motion Magnification khuếch đại chuyển động trong video giúp quá trình ước lượng tần số của chuyển động được dễ dàng hơn.
Mục tiêu đề tài : - Hiéu được dạng bài toán khuếch đại chuyên động cũng như bài toán frequency estimation. - Ứng dụng phương pháp vào các input tự thu thập được trong thực tế. - Minh họa trực quan phương pháp và ứng dụng thực tiễn trong nhiều lĩnh vực, ví dụ như công nghiệp và y khoa. Nội dung nghiên cứu của đề tài : Nội dung 1 : tìm hiểu về quy trình của phương pháp.
- Khảo sát và tổng hợp tài liệu liên quan đến các công nghệ cũng như các kỹ thuật được sử dụng trong các bài báo : Motion manipulation - motion magnification, deep convolutional neural network ,visual acoustics. Qua đó tổng quát hóa quy trình của phương pháp. - Chay thử các model và dataset được cung cấp sẵn và đánh giá. -_ Tỉnh chỉnh cũng như update các code lỗi thời thành một phiên bản phù hợp hơn.
- Kết qua dự kiến : báo cáo kết quả chạy thử va tng hợp tai liệu kỹ thuật chỉ tiết về phương pháp. Nội dung 2 : Xây dựng ứng dụng minh họa và cái test case thực tế. - _ thực hiện quay các video về các hiện tượng thực tế như máy móc hoặc mạch đập trên tay. -_ Biên tập các đoạn video để phù hợp nhất với ứng dụng.
- Xây dựng demo UI + Kết qua dự kiến : ứng dung demo và đánh giá hiệu năng Tài liệu tham khảo - Eulerian Video Magnification for Revealing Subtle Changes in the World , MIT CSAIL http://people.edu/mrub/papers/vidmag.pd f - The Visual Microphone: Passive Recovery of Sound from Video MIT CSAIL http://people.edu/mrub/papers/VisualMic_SIGGRAPH2014.pdf - Learning-based Video Motion Magnification MIT CSAIL, Cambridge, MA, USA http://people.edu/mrub/vidmag/papers/deepmag.pdf Kế hoạch thực hiện: +Giai đoạn 1 (5/9/2022 đến 15/10/2022 ) : Nghiên cứu bài báo,tìm hiểu công nghệ được sử dụng, chạy mô hình và dataset bài toán sử dụng , ghi chép lại các thông số chạy model + Giai đoạn 2 (15/10/2022 đến 20/11/2022 ) : Sử dụng bộ data khác để thử nghiệm model được chạy ở giai đoạn 1, ghi chép các thông số và đánh giá kết quả, phân tích và chỉ ra các điểm ảnh hướng tới kết quả dự đoán. Để từ đây tạo thành các video test case phù hợp nhất. + Giai đoạn 3 (20/11/2022 đến 20/12/2022) : xây dựng ứng dụng minh họa và tìm hướng cải thiện kết quả dự đoán. + Giai đoạn 4 (20/12/2022 đến khi báo cáo ): Tìm hướng cải thiện kết quả dự đoán, viết file bao cao , chuẩn bị slide bảo vệ khỏa luận.
Xác nhận của CBHD TP.năm 2022 (Ký tên và ghi rõ họ tên) Sinh viên (Ký tên và ghi rõ họ tên) ACKNOWLEDGEMENTS | would like to begin this thesis by acknowledging all of whom | owe this ac- complishment. We can not be the place we are today without all the unwavering support and guidance that was given to us. To my instructor Dr.Nguyen Vinh Tiep, his guide has become a shining bea- con for my journey. His vast wealth of knowledge has given me countless ideas and provided us with state-of-the-art equipment for my research, pointing out my inaccuracies and giving us priceless instructions.
| would like to show my utmost gratitude to Dr.Tiep for his contributions. To all the teachers and professors in the computer science department, | would like to send my appreciation for all of the clear directions that were given when | was on my way to completing this thesis. To our friends at the MMLab, who has been nothing but the greatest group of friend anyone could ever ask for. Without all the help and feedback | received my thesis could not be as complete as it is today.
furthermore, | would like to acknowledge my family to whom | owe everything. My achievement could never come to fruition without all the love and sacrifices that my family has made. Contents 1 Abstract 2 Introduction 2.1 Applicability in medical.2 Applicability in industry.3 Applicability in astrology.4 Sound retrieval from still video .4 Challenges and solutions. Random noise Ringing artifact.5 the main goal.3 Eulerian Video Magnification [Wu et al.4 Phase-Based Video Motion Processing.
30 Contents 3 4 Related Works 36 4.1 Learning-based Video Motion Magnification .1 Introdution to Learning-based Motion Magnification. 39 Deep Convolutional Neural Network Architecture .2 Synthetic Training Dataset .3 pixel Intensity signal tracking .3 Object’s vibration frequency estimation using video magnifi- cation and applications. 48 6 Results and Evaluations 51 6. 56 List of Figures 2.1 Input and output using video magnification + frequency estimation .2 Example: the pulse frequency output using video magnification + frequency estimation .3 this the result of a video magnification + frequency estimation that was proven to be reliable comparing to a real medical machine) .4 Example of a machine that can be track using camera.5 Example of the sky full with stars .6 A mock setup that can be used to test this application .7 Example of ringing artifact (a) is a sharp image (b) Is a picture with ringing artifact.8 The set up that was used in the thesis to collect data.1 Mô tả sample set (training set) 21 3.2 Mô tả sample set (training set) c.3 Learned regions of support allow features (a) and (b) to reliably track the leaf and background, respectively, despite partial occlusions.
For feature (b) on the stationary background, the plots show the x (left) and y (right) coordinates of the track both with (red) and without (blue) a learned region of support for appearance comparisons. The track using a learned region of support is constant, as desired for feature point on the stationary background.4 The input and output frame that show the deform after magnification 25 List of Figures 5 3.5 An example of using our Eulerian Video Magnification framework for visualizing the human pulse. (b) The same four frames with the subject’s pulse signal amplified. (c) A vertical scan line from the input (top) and output (bottom) videos plotted over time shows how our method amplifies the periodic color variation.
In the input sequence the signal is imperceptible, but in the magnified sequence the variation is clear. The complete sequence is available in the supplemental video.6 Overview of the Eulerian video magnification framework .7 relationship between temporal processing and motion magnification .8 the diagram of the phase-based approach manipulates motion .9 A big world of small motions. Representative frames from videos in which we amplify imperceptible motions. The full sequences and results are available in the supplemental video.10 Comparison of result on a common sequence 34 3.11 The main differences between the linear approximation of Wu et al.
[2012] and our approach for motion magnification. The representation size is given as a factor of the original frame size, where k represents the number of orientation bands and n represents the number of filters per octave for each orientation. 35 41 Detailed diagram for each part denotes a convolutional layer of c channels, k x k kernel size, and stride s.2 Overview of the architecture. Our network consists of 3 main parts: the encoder, the manipulator, and the decoder.
During training, the inputs to the network are two video frames, (Xa, Xb), with a magni- fication factor a, and the output is the magnified frame Y*. 40 43 Picture from MS COCO dataset as the background. 41 44 Picture from segmented objects from the PASCAL VOC dataset. 42 45 Applying our network in 2-frame settings.
We compare our network applied in dynamic mode to acceleration magnification. Because is based on the complex steerable pyramid, their result suffers from ringing artifacts and blurring.6 Example of the result of intensity signal tracking. 44 List of Figures 6 4.7 Example of the result of intensity signal tracking on human face.1 Testing dataset that was collected .2 How the testing dataset was collected .3 the testing data that was collected are sharp and usable .4 the testing data that was collected are sharp and usable 47 5.5 showing the detail get lost after the process .6 Our improvement by adding stoping and anchor point.