VIETNAM NATIONAL UNIVERSITY HO CHI MINH CITY HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY NGUYỄN THỊ TY RESEARCH AND DEVELOP SOLUTIONS TO TRAFFIC DATA COLLECTION BASED ON VOICE TECHNIQUES Major: Computer Science Major code: 8480101 MASTER THESIS HO CHI MINH CITY, July 2023 VIETNAM NATIONAL UNIVERSITY HO CHI MINH CITY HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY NGUYỄN THỊ TY RESEARCH AND DEVELOP SOLUTIONS TO TRAFFIC DATA COLLECTION BASED ON VOICE TECHNIQUES Major: Computer Science Major code: 8480101 MASTER THESIS HO CHI MINH CITY, July 2023 THIS THESIS IS COMPLETED AT HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY – VNU-HCM Supervisor: Assoc. Trần Minh Quang Examiner 1: Assoc. Nguyễn Văn Vũ Examiner 2: Assoc. Nguyễn Tuấn Đăng This master’s thesis is defended at HCM City University of Technology, VNU- HCM City on July 11, 2023 The board of the Master’s Thesis Defense Council includes: (Please write down the full name and academic rank of each member of the Master Thesis Defense Coun- cil) 1.
Lê Hồng Trang 2. Phan Trọng Nhân 3. Nguyễn Tuấn Đăng 5. Trần Minh Quang Approval of the Chairman of Master’s Thesis Committee and Dean of Faculty of Computer Science and Engineering after the thesis is corrected (If any).
CHAIRMAN OF THESIS COMMITTEE DEAN OF FACULTY OF COMPUTER SCIENCE AND ENGINEERING i VIETNAM NATIONAL UNIVERSITY - HO CHI MINH CITY SOCIALIST REPUBLIC OF VIETNAM HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY Independence – Freedom - Happiness THE TASK SHEET OF MASTER’S THESIS Full name: NGUYỄN THỊ TY Student code: 2171072 Date of birth: 22/11/1996 Place of birth: Binh Dinh Province Major: Computer Science Major code: 8480101 I. THESIS TITLE: Research and develop solutions to traffic data collection based on voice tech- niques (Nghiên cứu và phát triển các giải pháp thu thập dữ liệu giao thông dựa trên các kỹ thuật giọng nói). TASKS AND CONTENTS: • Task 1: Traffic Data Collection and Processing. The first task involves collecting comprehensive traffic data.
Extensive re- search will be conducted to identify reliable data sources, followed by the implementation of appropriate data collection techniques. Subsequently, ex- periments will be carried out to determine the most effective data processing methods. The aim is to enhance data quality and optimize processing effi- ciency for further analysis. • Task 2: Research and Experimentation for Automatic Speech Recognition Model Development.
In this phase, the focus will be on researching and experimenting with vari- ous architectures to develop high-performance automatic speech recognition models. Different techniques will be explored to achieve accurate speech-to- text conversion. The goal is to identify the best-performing model that meets the project’s requirements. ii • Task 3: Automatic Speech Recognition Model Evaluation and Future Work.
Once the automatic speech recognition models are developed, a comprehen- sive evaluation process will be undertaken. The achieved results will be ana- lyzed using appropriate metrics and techniques to assess their performance. Strengths and weaknesses of each model will be identified. Based on this analysis, recommendations for future work will be provided, outlining po- tential enhancements or modifications to the automatic speech recognition models.
THESIS START DAY: 06/02/2023. THESIS COMPLETION DAY: 09/06/2023. TRẦN MINH QUANG. Ho Chi Minh City, June 9, 2023 SUPERVISOR CHAIR OF PROGRAM COMMITTEE (Full name and signature) (Full name and signature) DEAN OF FACULTY OF COMPUTER SCIENCE AND ENGINEERING (Full name and signature) iii ACKNOWLEDGMENTS I would like to extend my sincere gratitude to the individuals who have provided invaluable support and assistance throughout my research journey.
I would like to express my formal appreciation to Assoc. Trần Minh Quang for his exceptional guidance, expertise, and unwavering support. His mentorship has been instrumen- tal in helping me navigate the necessary steps to complete this thesis. Whenever I encountered difficulties or felt lost, Assoc.
Quang provided invaluable advice that steered me back in the correct direction. His suggestion to process the data to enhance its quality was a significant contribution to my research. Furthermore, his assistance in establishing contact with esteemed researchers working on topics sim- ilar to mine and facilitating connections with individuals who could provide server support for training large models, such as automatic speech recognition models, has been immensely valuable. I would like to express my profound gratitude to the esteemed researchers, Mr.
Nguyễn Gia Huy and Mr. Nguyễn Tiến Thành, for their generous contributions in sharing their profound insights and knowledge. Their willingness to address my in- quiries regarding the Urban Traffic Estimation System, collected data, and existing issues has significantly enriched my comprehension of the subject matter. Further- more, I am sincerely thankful to my sisters, Ms.
Nguyễn Thị Nghĩa and Ms. Nguyễn Thị Hiển, as well as Lương Duy Hưng and Vũ Anh Nhi, for their invaluable support in meticulously creating precise transcripts for the audio files. Furthermore, I would like to express my deep appreciation to Mr. Tăng Quốc Thái for his diligent efforts in meticulously collecting and securely storing the traffic reports from VOH 95.
Additionally, I am profoundly grateful to Mr. Mai Tấn Hà, who graciously provided me with access to a server for the training of automatic speech recognition models. His generosity and support have been instrumental in en- abling the successful execution of the model training process. I would also like to extend my formal gratitude to Dr.
Lê Thành Sách and Mr. Nguyễn Hoàng Minh from the Data Science Laboratory at Ho Chi Minh City University of Technology (HC- iv MUT) for their kind approval in granting me the opportunity to utilize an independent server for automatic speech recognition model training. Their trust and support from the Data Science Laboratory have been pivotal in facilitating the smooth progress of my research. In addition, I am sincerely thankful for the invaluable support rendered by my friends, Mr.
Nguyễn Tấn Sang and Mr. Huỳnh Ngọc Thiện, in working with the server that has limited permissions. Their expertise and assistance have been in- dispensable in effectively navigating the constraints imposed by the server limitations. Lastly, I would like to express my heartfelt gratitude to my boss, co-workers, friends, and family for their unwavering emotional support and understanding during the challenging times that I encountered throughout this research endeavor.
Their encouragement and belief in my abilities have been instrumental in my success. Once again, I am deeply grateful to all of the individuals mentioned above for their significant contributions and support, without which this thesis would not have been possible. v ABSTRACT This thesis addresses two fundamental challenges within the domain of the cur- rent intelligent traffic system, specifically the Urban Traffic Estimation (UTraffic) System. The first challenge pertains to the insufficiency of data that meets the req- uisite standards for training the automatic speech recognition (ASR) model that will be deployed in the UTraffic system.
The current dataset predominantly consists of synthesized data, resulting in a bias towards recognizing synthesized traffic speech reports while struggling to accurately transcribe real-life traffic speech reports im- ported by UTraffic users. The second challenge involves the accuracy of the ASR model deployed in the current UTraffic system, particularly in transcribing real-life traffic speech reports into text. To address these challenges, this research proposes several approaches. Firstly, an alternative traffic data source is identified to reduce the reliance on synthesized data and mitigate the bias.
Secondly, a pipeline incorporating audio processing tech- niques such as sampling rate conversion and speech enhancement is designed to ef- fectively process the dataset, with the ultimate objective of improving ASR model performance. Thirdly, advanced and suitable ASR architectures are experimented with using the processed dataset to identify the most optimal model for deployment within the UTraffic system. Significant achievements have been obtained through this research. Firstly, a new dataset of superior quality compared to the previous one has been developed.
Con- tinuous data collection from the alternative traffic data source can further enhance this dataset, making it a valuable resource for future research endeavors aiming to im- prove the ASR model deployed in the UTraffic system. Additionally, notable progress has been made in improving the accuracy of the ASR model compared to the results achieved by the current architecture of the UTraffic system’s ASR model. vi TÓM TẮT LUẬN VĂN Luận văn này giải quyết hai thách thức cơ bản trong lĩnh vực hệ thống giao thông thông minh hiện tại, cụ thể là Hệ Thống Dự Báo Tình Trạng Giao Thông Đô Thị (UTraffic). Thách thức đầu tiên liên quan đến sự thiếu hụt dữ liệu đáp ứng tiêu chuẩn cần thiết cho việc huấn luyện mô hình nhận dạng giọng nói tự động (ASR), sẽ được triển khai trong hệ thống UTraffic.
Bộ dữ liệu hiện tại chủ yếu bao gồm dữ liệu tổng hợp, dẫn đến sự thiên vị cho việc nhận dạng các báo cáo giao thông tạo từ giọng nói tổng hợp, trong khi gặp khó khăn trong việc chuyển các báo cáo giao thông ở dạng giọng nói được cung cấp bởi người dùng UTraffic sang văn bản chính xác. Thách thức thứ hai liên quan đến độ chính xác của mô hình ASR triển khai trong hệ thống UTraffic hiện tại. Để giải quyết những thách thức này, nghiên cứu này đề xuất một số phương pháp. Thứ nhất, xác định nguồn dữ liệu giao thông thay thế để giảm thiểu sự phụ thuộc vào dữ liệu tổng hợp.
Thứ hai, thiết kế luồng xử lý thích hợp, trong đó kết hợp các kỹ thuật xử lý âm thanh như chuyển đổi tỉ lệ lấy mẫu và tăng cường giọng nói để xử lý hiệu quả bộ dữ liệu đang có, với mục tiêu cuối cùng là cải thiện hiệu suất mô hình ASR. Thứ ba, thử nghiệm bộ dữ liệu đã được xử lý trên các kiến trúc ASR tiên tiến để xác định được mô hình tối ưu nhất cho việc triển khai trong hệ thống UTraffic. Nghiên cứu này đã đạt được thành tựu đáng kể. Thứ nhất, chúng ta hình thành được một bộ dữ liệu mới có chất lượng vượt trội hơn so với bộ dữ liệu ban đầu.
Việc tiếp tục thu thập dữ liệu từ nguồn thay thế có thể nâng cao hơn nữa chất lượng của bộ dữ liệu hiện có, biến nó thành nguồn tài nguyên quý giá cho những nỗ lực nghiên cứu cải thiện hiệu suất mô hình ASR triển khai trong hệ thống UTraffic trong tương lai. Ngoài ra, so với các kết quả đạt được bởi mô hình ASR hiện tại trong hệ thống UTraffic, chúng ta đã đạt được những tiến bộ đáng kể, đặc biệt trong việc cải thiện độ nhận dạng giọng nói chính xác. vii DECLARATION I, Nguyễn Thị Ty, solemnly declare that this thesis titled "Research and develop solutions to traffic data collection based on voice techniques" is the result of my own work, conducted under the supervision of Assoc. Trần Minh Quang.
I af- firm that all the information presented in this thesis is based on my own knowledge, research, and understanding, acquired through extensive study and investigation. I further declare that any external assistance, whether in the form of data, ideas, or references, has been duly acknowledged and properly cited in accordance with the established academic conventions. I have provided appropriate references and citations for all the sources and materials used in this thesis, giving credit to the original authors and their contributions. I acknowledge that this thesis is intended to fulfill the demands of society and to contribute to the existing body of knowledge in the field.
It represents the culmination of my efforts, dedication, and commitment to advancing knowledge and understand- ing in this area. I hereby affirm that this thesis is an authentic and original piece of work, and I take full responsibility for its content. I understand the consequences of any act of plagiarism or academic dishonesty, and I assure that this thesis has been prepared with utmost integrity and honesty.