MINISTRY OF EDUCATION AND TRAINING HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY AND EDUCATION HCMUIIE GRADUATION THESIS MAJOR: DATA ENGINEERING A VIETNAMESE SPEECH CORPUS FOR TEXT TO SPEECH INSTRUCTOR: HOANG VAN DUNG STUDENT: HUYNH NGUYEN TIN NGUYEN MINH TIEN SKLO13687 Ho Chi Minh city, 7/2024 TRƯỜNG ĐẠI HỌC SƯ PHẠM KỸ THUẬT TP. HÒ CHÍ MINH KHOA CÔNG NGHỆ THÔNG TIN BO MON KY THUAT DU LIEU ( HCMUTE HUỲNH NGUYÊN TÍN - 20133094 NGUYEN MINH TIEN - 20133093 DE TAI: A VIETNAMESE SPEECH CORPUS FOR TEXT TO SPEECH KHOA LUAN TOT NGHIEP KY SU KTDL GIAO VIEN HUONG DAN PGS. HOANG VAN DUNG KHOA 2020-2024 TP. Hà Chí Minh, Tháng 07 năm 2024 TRƯỜNG ĐẠI HỌC SƯ PHẠM KỸ THUẬT TP.
HÒ CHÍ MINH KHOA CÔNG NGHỆ THÔNG TIN BO MON KY THUAT DU LIEU ( HCMUTE HUỲNH NGUYÊN TÍN - 20133094 NGUYEN MINH TIEN - 20133093 DE TAI: A VIETNAMESE SPEECH CORPUS FOR TEXT TO SPEECH KHOA LUAN TOT NGHIEP KY SU KTDL GIAO VIEN HUONG DAN PGS. HOANG VAN DUNG KHOA 2020-2024 TP. Hà Chí Minh, Tháng 07 năm 2024 ĐH SƯ PHẠM KỸ THUẬT TP.HCM XÃ HỘI CHÚ NGHĨA VIỆT NAM KHOA CNTT Độc lập —- Tự do - Hạnh Phúc 2k 2 i 2 2 2k ok slesle ke ok sk sk sÉ PHIEU NHAN XET CUA GIAO VIEN HUONG DAN Họ và tên Sinh viên 1: Huỳnh Nguyễn Tín MSSV 1: 20133094 Ho va tén Sinh vién 2: Nguyén Minh Tién MSSV 2: 20133093 Ngành: Kỹ Thuật Dữ Liệu Tên đề tài: A Vietnamese speech corpus for text to speech Họ và tên Giáo viên hướng dẫn: PGS. Hoàng Văn Dũng NHẬN XÉT 1.
Về nội dung đề tài & khối lượng thực hiện: 4. Đề nghị cho bảo vệ hay không? 5. Đánh giá loại: 6. Hồ Chí Minh, ngày tháng 7 năm 2024 Giáo viên hướng dẫn (Ký & ghi rõ họ tên) DH SU PHAM KY THUAT TP.HCM XA HOI CHU NGHIA VIET NAM KHOA CNTT Độc lập —- Tự do - Hạnh Phúc FOR OE OK OK OK OK slesle ke ok sk sk sÉ PHIEU NHAN XET CUA GIAO VIEN PHAN BIEN Họ và tên Sinh viên 1: Huỳnh Nguyễn Tín MSSV 1: 20133094 Ho va tén Sinh vién 2: Nguyén Minh Tién MSSV 2: 20133093 Ngành: Kỹ Thuật Dữ Liệu Tên đề tài: A Vietnamese speech corpus for text to speech Họ và tên Giáo viên phản biện: TS.
Nguyễn Thành Sơn NHẬN XÉT 1. Về nội dung đề tài & khối lượng thực hiện: 4. Đề nghị cho bảo vệ hay không? 5. Đánh giá loại: 6.
Hồ Chí Minh, ngày tháng 7 năm 2024 Giáo viên phản biện (Ký & ghi rõ họ tên) Trường ĐH Sư Phạm Kỹ Thuật TP.HCM Khoa : Công Nghệ Thông Tin ĐÈ CƯƠNG LUẬN VĂN TÓT NGHIỆP Họ và Tên SV thực hiện 1 : Huỳnh Nguyễn Tín Mã Số SV: 20133094 Họ và Tên SV thực hiện 2 : Nguyễn Minh Tiến Mã Số SV: 20133093 Thời gian làm luận văn : Từ : 03/2024 Đến : 07/2024 Chuyên ngành : Kỹ Thuật Dữ Liệu Tên luận văn : A Vietnamese speech corpus for text to speech GV hướng dẫn : PGS. Hoàng Văn Dũng Nhiệm vụ của luận văn: 1. Nghiên cứu cách lấy dữ liệu speech có sẵn từ Internet 2. Nghiên cứu và xây dựng text normalizer cho tiếng Việt 3.
Xây dựng pipeline đê xây dựng thành speech corpus 4. Thực nghiệm Cấu trúc khóa luận CHAPTER 1: INTRODUCTION 1. Objectives of the Thesis 1. Related Works CHAPTER 2: LITERATURE REVIEW 2.
Existing Vietnamese Speech Corpora 2. Export and Publish 2. Vietnamese Text Normalization 2. Part-of-speech Tagging 2.
RDRsegmenter for word segmentation 2. VnMarMoT CHAPTER 3: METHODOLOGY 3. Addressing Non-Standard Words 3. Detection and Normalization of Non-Standard Words 3.
Addressing Out-of-Vocabulary Words 3. Bridging the Gap 3. Post processing CHAPTER 4: CORPUS ANALYSIS 4. Fine-tuning a Text-to-Speech Model CONCLUSION REFERENCES Kế hoạch thực hiện STT Thời gian Công việc Ghi - chú 1 04/03/2024 - 10/03/2024 | o_ Tìm hiệu các công trình nghiên cứu trước đó o_ Phân tích, đánh giá tổng quan đề tài 2 11/03/2024 - 31/03/2024 | o_ Tìm hiểu, đánh giá các công cụ để cào dữ liệu Xây dựng pipeline cào dữ liệu Thực hiện cào dữ liệu 3 01/04/2024 — 14/04/2024 Nghiên cứu về các công cụ chuẩn hóa văn bản trước đó Nghiên cứu âm học, âm tiết của tiếng Việt Nghiên cứu các mô hình bổ trợ như POSTagger,.
4 1504/2024 - 19/05/2024 Xây dựng các từ điển cho các dấu câu, từ viết tắt, phiên âm Xây dựng pipeline để chuẩn hóa dữ liệu Thực hiện kiểm thử 5 20/05/2024 - 02/06/2024 Nghiên cứu về Forced Aligner và áp dụng vào corpus 6 03/06/2024 — 16/06/2024 Xây dung pipeline cho text- normalization va forced alignment 7 17/06/2024 — 07/2024 Thực hiện đánh giá, phân tích corpus Kiểm thử đữ liệu Viết báo cáo Ý kiến của giáo viên hướng dẫn (ký và ghỉ rõ họ tên) Ngày tháng năm 2024 Người viết đề cương ABSTRACT Tin Nguyen Huynh, Tien Minh Nguyen: A Vietnamese speech corpus for text to speech (Under the direction of Van-Dung Hoang) Vietnamese is an under-resourced language and expectation for the large-scale Vietnamese speech corpus is increasing day by day as a result of the evolution of language models. In this work, we present a cost-free and low-dependency pipeline for processing and curating a Vietnamese speech corpus from publicly available videos on YouTube. For low- resource languages, obtaining sufficient quality data is the main bottleneck in developing large language models. Vietnamese is one such language, and while public corpora exist, they consist of manually annotated recordings of different speakers.
While this approach is guaranteed to produce high-quality and error-proof data, it is expensive, time-consuming, and unrealistic for students or independent speech researchers. Motivated by this, our objective is to tackle the problem at its root by proposing a scalable processing pipeline to curate a large Vietnamese speech corpus. We leverage free and public Vietnamese news videos on YouTube that contain closed captions as the source of our data. We also explore various aspects of Vietnamese text normalization and propose a text normalizer that supports a novel method for transliterating foreign words.
ACKNOWLEDGEMENTS First, we would like to extend our sincere gratitude to our thesis advisor, Assoc. Van-Dung Hoang. Your exceptional academic expertise and dedication have been instrumental in the development of this thesis. We would also like to express our heartfelt appreciation to our thesis committee member, Dr.
Thanh-Son Nguyen. Thank you for your time and effort in evaluating our work and providing thought-provoking questions. During our final years, our academic journeys have piqued our interest in natural language processing. Researching such a fast-growing field not only pushed us beyond our comfort zones but also introduced us to the vast world of new possibilities and experiences.
This would not have ever been possible had it not been for the university’s motivation and encouragement in pursuing knowledge and pushing boundaries. Thank you! Table of Contents CHAPTER 1: INTRODUCTIONN. 11121111 11T TH TH TH HH 1 1. Objectives of the THheSiS .-- - Ăn Họ HT Ert 1 1.- E3 123112311211 9112 21T TH TH HH HT 2 CHAPTER 2: LITERA TURE REV IEW.
Exisng Vietnamese Speech COTpDOTa.--- «+ s3 3x vn ng ng nghe 3 2. c5 5 231321111231 11 11v 3v TH ng ng 3 2.-- - c1 SH nh ng no HT 4 2.c ST HH HH HH TH ng ng 6 2. Export and Publish .--- -‹-- «c1 1211112111 vn vn vn ng HH ng 6 P895: on ốc ốốốẽốố. Vietnamese Text NormalizafIOn .- + «c2 3113121113211 9111 vn ng ng nghe 7 2.
Part-of-speech Tagging .-sc Ăn TH ng nhi 9 2. RDRsegmenter for word segmenfafiON.v ng ng tre. VnMarMoT CHAPTER 3: METHODOLOGY.--- -- «+ + TH HH ng.-- «5c +3 ng ng TH ng. c1 nọ KH Họ HE 21 3.
Addressing Non-Standard WOrs.--c + 3n ng HH ngư. Ăn KH nọ TT HE 24 3. Detection and Normalization of Non-Standard WOTCS. Addressing Out-of-Vocabulary WOFdS .--- «c3 33 vn vn ng ng rưyy 34 3.
Bridging the Gap.Ă SH HH HH.--- c3 1311111 ng ng tế.-- «nh ng HT HH TT ng 39 3.-- -- + s1 KH nọ KH HT 44 3.-- - - 311111291011 KH Họ Họ ko kg 45 3.---- -- «+ c1 v3 vn HH ng ng. POS{ DTOC€SSITB.- TQ ng HH ng ng 60 CHAPTER 4: CORPUS ANALLYSIS. cà SH HH ng ng tr 62 9 oi nh. Fine-tuning a Text-to-Speech Model.
CONCLUSION REFERENCES LIST OF TABLES Table 1. Vietnamese speech COTDOTA OV€TVICW. óc 1S TT ng nh nhe 2 Table 2. Old and modern methods for placing tone 1marks.
Examples of key-value 1n in the 5-syllable context dictionary. Short description of RDR segm€TI€T .- + 6S Sx xe, 12 Table 5. Sample of tag, plain text and the explanatiO. Speed comparison between different resamplers in downsampling.
Categorization of Vietnamese non-standard WOrds. Accuracy and speed comparison between part-of-speech taggers. Part-of-speech tag set of the MarMoT-based tagØØeT. SemiotiCc cÏaSS€S D€T {Ag.
(11 1T T HH TH TH HT HH ngư 32 Table 11. Classification rule per semIOfiC CÏ4SS .-ó- (6 5+ 1S Sk*k key 33 Table 12. ARPA pronunC1afIOTIS.- 5 + 6< 11 1 1E 1111 1 1111 TH TH ng ngư, 38 Table 14. Behavior of Gorman”s syl|la1ÍTT.---- - + + S1 ng ngư, 41 Table 15.
Comparison between Gorman”s and our syllab1fÍier. Transliterated ARPA vowel phones (left) and consonant phones (right). Transliterated two-phone nuCÌ€1.-¿-¿- - 6 + +11 12k HT ngư, 49 Table 18. Transliterated three-phone nucÌe.
Nucleus correction of illicit rimes after transliterafion. Quantity of audios 1n each range of alignmenf SCOF€. LIST OF FIGURES Front-end of recording application 00. ccc - ¿c1 1S TT HT HH rên 5 Pipeline architecture of VnCorelNLP.------ - - + 1S vn HH rên 9 A sample of SCRDR tree for POS 'Tagging.---- - 6 + xe, 10 A diagram of another approach to construct an SCRDR tree.---- 11 Content layout of a Sub RIp ẨiÌ€.
- -- 6 + S1, 17 Content of a converted SubRip Ẩile.- -- - 6 11k Hi, 18 Directory structure of the data collection fOlder.---- - «+ «+s+sc++s+scsx+ 19 Comparison of anti-aliasing filters of different resamplers in upsampling. 22 Subset oŸ possible semIOtIC CÏSS€S. - cà 1n HH gi, 26 Components OŸ a syÏÏabÏ€.---- - 6 + 1S 1211121 HH TH TH nghiệt 39 Pseudo-code for Gorman”s syllabification algorithm. ¿5 «5< £ss2 41 Modifications to Gorman”s syllabification algorithm.
eee 43 Content of the annotated CMŨ diCfIOTATV.-- (5+ S*SEsekrerskekrerske 44 Forced alignment. 55 Forced alignment on the word leVeÌL. - - 6 ¿+ + *kE*E*EkEeEekrkrerekrkrke 58 Trimming of non-participating words and silenC€s .---- - 5-55 s2 59 The distribution of gender of the raw COFPUS. cece cesses ects teeeneteeseeees 62 The distribution of alignment score of the raw COTPUS.
5-5 «55s 52 62 The distribution of audio duratlons 1n our clean COTPUS. Context In speech-related studies like speech analysis and speech synthesis, the emergence of high-quality speech corpora is increasing day by day. The most popular language, such as English has many studies related to speech, especially the plenty of research in speech corpora because of its universality. On the other hand, linguistics studies in Vietnamese are challenges for every research project in every field because of the ambiguity and lack of existing-based linguistics studies for the following studies in the future.
Vietnamese is spoken by approximately 100 million people (2024), but a very limited amount of speech samples are available. Building a speech corpus is the most essential part of every task related to speech, like TTS and speech recognition, etc. This step required massive effort, resources, and considerable time to converge.